Black Forest Labs' 12-billion parameter multimodal flow transformer for in-context image generation and editing with character consistency, typography handling, and commercial-ready quality
Settled by community votes across 6 shared challenges, with an AI judge weighing in on each.
FLUX.1 Kontext [pro]
#41 of 62 in Text-to-Image
Imagen 3.0 Generate 002
#36 of 62 in Text-to-Image
Where the votes landed
FLUX.1 Kontext [pro]
0%
win rate
Ties
0%
Imagen 3.0 Generate 002
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
FLUX.1 Kontext [pro]
- + Excellent text legibility and font hierarchy
- + Clean minimalist layout that follows professional graphic design principles
- + Perfect adherence to requested sections (Appetizers/Pizza/Mains)
- − Menu items are nonsensical/gibberish words
- − Food photo content is slightly inconsistent with the specific section headers
Imagen 3.0 Generate 002
- + Large variety of high-quality food photography
- + Strict adherence to the grid layout requirement
- + Professional 'lookbook' style aesthetic
- − Text is largely unreadable and repetitive across sections
- − Overwhelming amount of repetitive images makes it less functional as a menu
- − Misspelled section headers like 'APPETIEES' and 'MANS'
Verdict: FLUX.1 Kontext [pro] produces a much more functional and realistic menu design with clear typography and professional spacing, effectively representing the 'casual dining' vibe. While Imagen 3.0 Generate 002 creates an attractive grid, its failure to produce readable text and the repetitive nature of its content makes it less useful as a design template.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
FLUX.1 Kontext [pro]
- + Excellent typography with a glowing fire effect that naturally integrates into the scene.
- + High-quality photorealistic food textures, especially the melted cheese and seared patty.
- + Great atmospheric lighting with realistic embers and smoke.
- − The burger is less 'exploded' than Model B, appearing more like it is just slightly floating.
- − Redundancy in text, displaying the price twice which wasn't requested.
Imagen 3.0 Generate 002
- + Perfectly captures the 'exploded' view with clearly separated suspended components.
- + Includes the specific starburst element requested for the price.
- + Dynamic composition with sauce droplets and elements flying through the air.
- − The food textures look slightly more CGI/plastic and less photorealistic than Model A.
- − Failure on text rendering for the secondary content inside the starburst (gibberish).
Verdict: FLUX.1 Kontext [pro] creates a much more appetizing and professional-looking advertisement with superior lighting and texture detail, though it misses the 'starburst' and full 'exploded' effect. Imagen 3.0 Generate 002 follows the structural layout requirements of the prompt more accurately, but suffers from lower photorealism and garbled text within the price badge.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
FLUX.1 Kontext [pro]
- + Excellent text accuracy with only one minor spelling error at the bottom.
- + Very realistic chalk texture and natural handwriting style.
- + Clean and consistent composition that feels authentic to a café.
- − The title is in print-style handwriting rather than the requested 'elegant cursive'.
- − Distracting spelling artifacts in the bottom disclaimer text ('ous gluten tree').
Imagen 3.0 Generate 002
- + Successfully used a cursive/script style for the menu items.
- + Good variety in chalk colors and smudging details on the chalkboard surface.
- − Significant spelling errors and hallucinations in the menu items (e.g., 'Risottto', 'Octpsuin & ebbs').
- − Repeated lines of text which break the realism of a handwritten menu.
- − Title text has a digital, outlined appearance rather than a natural chalk feel.
Verdict: FLUX.1 Kontext [pro] is the clear winner as it produces legible, mostly accurate text with a convincing chalk texture. While it failed to use cursive for the title, Imagen 3.0 Generate 002 suffered from significant spelling failures and repetitive text blocks that ruined the prompt adherence.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
FLUX.1 Kontext [pro]
- + Excellent adherence to the complex spatial instruction of 'horse on top'.
- + High cinematic quality with realistic space lighting and detailed textures.
- + Creative surrealism by merging the two subjects into a paradoxical composition.
- − Anatomical glitch where the astronaut's legs appear to be horse hooves.
- − The smaller astronaut on top of the horse creates a confusing recursive loop.
Imagen 3.0 Generate 002
- + Beautiful cosmic background with vibrant nebulae.
- + Crisp lighting and sharp details on the astronaut's suit.
- − Completely failed the negative instruction; the astronaut is on the horse, not vice versa.
- − Generic interpretation of the prompt without the requested surreal reversal.
Verdict: FLUX.1 Kontext [pro] successfully followed the difficult instruction to place the horse on top of the astronaut, creating a truly surreal and cinematic image. Imagen 3.0 Generate 002 ignored the specific 'not vice versa' constraint, producing a high-quality but standard image of an astronaut riding a horse. Despite a few anatomical glitches in the background figures, FLUX.1 Kontext [pro] is the clear winner for prompt adherence.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
FLUX.1 Kontext [pro]
- + Excellent typography style that matches the 3D aesthetic
- + High-quality PBR textures on the salmon and rice grains
- + Perfect central composition and lighting
- − The flag icon is stylized to the point of being a non-standard shape
Imagen 3.0 Generate 002
- + Accurate isometric 45-degree angle
- + Clean rendering with no artifacts
- + Proper flag icon rendering
- − Text is not center-aligned as requested
- − Texture details are a bit flat and more plastic-like than refined
- − Multiple sushi pieces deviate from the 'minimal garnish and plate' instruction compared to Image A
Verdict: FLUX.1 Kontext [pro] followed the layout instructions more closely, particularly the requirement for the text to be top-center and the overall composition to be centered. While Imagen 3.0 captures the isometric angle well, FLUX uses superior PBR materials that give the miniature a more premium, tactile look.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
FLUX.1 Kontext [pro]
- + Excellent golden sunrise lighting with visible god rays.
- + Vibrant color palette and cohesive aesthetic.
- + Accurately depicts all four species requested in the prompt.
- − The bunny has cat-like facial features and whiskers, lacking distinct rabbit characteristics.
- − The animals appear somewhat static and posed rather than 'playfully chasing and tumbling'.
- − The depth of field is very shallow, blurring the background into a generic glow.
Imagen 3.0 Generate 002
- + Successfully captures the dynamic 'tumbling' action requested in the prompt.
- + Superior anatomical accuracy for the bunny, kitten, and fox.
- + Beautiful rendering of dew sparkles on the grass and flowers.
- − The lighting is a bit harsher compared to Model A's soft golden hour.
- − The butterflies are less integrated into the scene's movement.
- − Less emphasis on 'big expressive eyes' compared to the slightly stylized look of Model A.
Verdict: Imagen 3.0 Generate 002 is the winner because it better follows the behavioral instructions, showing the animals actually tumbling and interacting in the grass. While FLUX.1 Kontext [pro] has a very beautiful lighting style, it failed to differentiate the bunny's face from the kitten's face, whereas Imagen 3.0 Generate 002 maintained distinct species characteristics.
Explore each model
Google's Imagen 3.0 text-to-image generation model, producing high-quality images with improved detail and lighting