Black Forest Labs' open-weights multimodal flow transformer for in-context image generation and editing, available for non-commercial use with character consistency and style transfer capabilities
Settled by community votes across 6 shared challenges, with an AI judge weighing in on each.
FLUX.1 Kontext [dev]
#54 of 62 in Text-to-Image
Imagen 3.0 Generate 002
#36 of 62 in Text-to-Image
Where the votes landed
FLUX.1 Kontext [dev]
0%
win rate
Ties
0%
Imagen 3.0 Generate 002
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + High resolution food photography with excellent texture and color.
- + Uses bold and modern sans-serif fonts as requested.
- + Clean, asymmetrical grid layout that feels contemporary.
- − Nonsensical text with significant rendering glitches and overlapping characters.
- − Sections for specific food items are not clearly organized or readable.
- − Food photos, while high quality, are somewhat repetitive in composition.
Imagen 3.0 Generate 002
- + Excellent adherence to the 'grid' layout and 'sections' requirement.
- + Better text rendering for headers like 'PIZZA' and 'APPETIEES'.
- + Very clear and professional menu structure with prices and item names.
- − Food photography is slightly lower in quality compared to the first model.
- − Formatting of small body text becomes blurry and illegible.
- − Repeated use of pizza images across different sections diminishes variety.
Verdict: Imagen 3.0 Generate 002 creates a more functional and realistic menu layout that successfully incorporates all requested sections (Appetizers, Pizza, Mains) in a clean grid. While FLUX.1 Kontext [dev] offers superior individual image quality for the food, it fails to organize the layout into a coherent menu, resulting in significant text corruption and poor sectioning.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Excellent typography rendering with almost perfect spelling
- + Vibrant colors with a strong, clean commercial feel
- + Good use of the fiery background elements
- − Failed the 'exploded burger' requirement, showing a fully assembled burger instead
- − Typography layout is a bit crowded and overlaps background elements poorly
- − Typo in the secondary message: 'LNHLY' instead of 'ONLY'
Imagen 3.0 Generate 002
- + Perfectly followed the 'exploded burger' layout with suspended components
- + High level of photorealistic detail on individual ingredients like the patty and lettuce
- + Superior composition with a more dynamic and professional commercial look
- − Text rendering is weaker, with gibberish present in the starburst
- − Main title text is slightly less prominent than requested
Verdict: While FLUX.1 Kontext [dev] produced very clear text, it failed the primary creative requirement of an 'exploded' burger, showing a static assembled product instead. Imagen 3.0 Generate 002 captured the difficult 'exploded' layout perfectly with high photorealism, making it the more successful image despite some minor gibberish text in the starburst.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Excellent chalk texture and realistic handwriting variation.
- + Strong composition that feels like a natural café environment.
- − Significant spelling errors throughout the text (e.g., 'Mashroom', 'Risoktso', 'Octpus').
- − The date formatting is garbled and difficult to read.
Imagen 3.0 Generate 002
- + Text rendering is very clear and more accurate than Model A.
- + Great chalk dust/smudge details on the background of the board.
- + Captures a consistent cursive handwriting style throughout.
- − Includes some duplicate lines of text and repetitive menu entries.
- − Contains several spelling errors in the menu items (e.g., 'Risottto', 'Octpsuin', 'Choouip').
Verdict: FLUX.1 Kontext [dev] captures the requested chalk texture and natural variations in handwriting much better than its competitor, but fails significantly on spelling and text legibility. Imagen 3.0 Generate 002 is much clearer and handles the cursive script well, but it suffers from repetitive text blocks and several typos. FLUX.1 Kontext [dev] is preferred for its superior aesthetic realism, despite the spelling issues.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Successfully followed the difficult spatial instruction of placing the horse on top of the astronaut.
- + Maintains clear, high-resolution textures on the space suit and horse fur.
- + Captures a surreal, clean aesthetic.
- − The composition is a bit stiff with the horse appearing to grow out of the astronaut's back rather than 'riding' him.
- − The background is relatively sparse compared to the cinematic request.
Imagen 3.0 Generate 002
- + Beautiful cinematic lighting and background nebulae.
- + High level of detail in the horse's mane and the astronaut's gear.
- − Completely failed the negative constraint/spatial instruction, showing the astronaut on top of the horse.
- − The horse's front legs have anatomical inconsistencies (multiple joints/strange hoof angles).
Verdict: This comparison highlights a classic test of prompt adherence vs. aesthetic quality. FLUX.1 Kontext [dev] is the winner because it successfully followed the specific and difficult instruction to have the horse on top of the astronaut, whereas Imagen 3.0 defaulted to a standard (inverse) trope despite the explicit prompt.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Excellent text rendering and alignment
- + Very clean, bold graphic style
- + Perfectly centered composition
- − The flag icon is stylized incorrectly
- − The 'isometric' perspective is a bit shallow
- − The sushi piece is very abstract and lacks realistic textures
Imagen 3.0 Generate 002
- + Perfect adherence to the isometric perspective
- + Accurate Japanese flag icon
- + Beautiful detailed 3D modeling of various sushi types
- − Text is smaller and less impactful than requested
- − Text is not perfectly top-center
- − Includes extra garnish like chopsticks and wasabi not explicitly requested
Verdict: FLUX.1 Kontext [dev] delivers a much stronger graphic design look with bold, clear text that dominates the center, but the 'flag' is an abstract shape. Imagen 3.0 Generate 002 creates a superior 3D isometric diorama with better textures and a correct flag, though it is slightly less 'ultra-clean' due to the extra props. Imagen 3.0 is preferred for its superior adherence to the 'isometric' and 'miniature 3D cartoon scene' aspects of the prompt.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
FLUX.1 Kontext [dev]
- + Excellent sense of motion and action
- + Beautiful golden lighting and backlighting
- − Failed to include the red fox kit and baby bunny
- − Repeated the 'kitten' subject multiple times instead of diverse animals
- − Butterflies appear somewhat flat and artificial
Imagen 3.0 Generate 002
- + Included all four requested animals (dog, cat, bunny, fox)
- + Superior texture in the fur and floral details
- + Captured the 'tumbling' and 'dew sparkles' elements very effectively
- − Static composition compared to the 'chasing' prompt
- − Slight anatomical confusion under the puppy's tail
Verdict: Imagen 3.0 Generate 002 is the clear winner as it successfully followed the complex prompt to include four specific, different animals, whereas FLUX.1 Kontext [dev] only rendered puppies and kittens. Imagen also provided much higher detail in the 'dew sparkles' and fur textures, resulting in a more 'hyper-photorealistic' look.
Explore each model
Google's Imagen 3.0 text-to-image generation model, producing high-quality images with improved detail and lighting