Black Forest Labs' premium multimodal flow transformer with greatly improved prompt adherence and typography generation for in-context image generation and editing without compromise on speed
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
FLUX.1 Kontext [max]
#23 of 62 in Text-to-Image
Imagen 4.0 Generate 001
#55 of 62 in Text-to-Image
Where the votes landed
FLUX.1 Kontext [max]
0%
win rate
Ties
0%
Imagen 4.0 Generate 001
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent depiction of window lighting and caustic reflections on the table.
- + High realism in the glass texture and refraction of the background plant.
- + The cube has a believable hollow structure that contains the sphere naturally.
- − The sphere texture is slightly grainy/glittery rather than a smooth sphere.
- − The book title text is nonsensical gibberish.
Imagen 4.0 Generate 001
- + Perfectly smooth blue sphere with realistic shading.
- + Solid composition with clear visibility of all requested elements.
- + The red book looks very tactile and high-quality.
- − The glass cube appears solid rather than hollow, making the sphere look like it is embedded in acrylic.
- − The plant is visible beside the cube rather than being clearly visible through it as requested.
- − Lighting is a bit flat compared to the dynamic shadows and highlights in the other model.
Verdict: FLUX.1 Kontext is the clear winner as it successfully renders the glass cube as a hollow container, creating beautiful and realistic light refractions from the window and the plant behind it. While Imagen 4.0 produces a very clean image, it fails the spatial logic of the prompt by making the cube look like a solid block with a sphere suspended inside, and it misses the specific request for the plant to be visible through the glass.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Expertly captures the light rain and atmospheric motion blur requested in the prompt.
- + Excellent shallow depth of field and realistic background bokeh.
- + The man's skin texture and posture feel authentic and candid.
- − The man's hands appear somewhat fused with the bicycle chain.
- − The rain streaks are a bit uniform and overlay-like in the foreground.
Imagen 4.0 Generate 001
- + Very high detail in the facial features and skin texture.
- + Good color vibrance with the red bike and city lights.
- + Realistic rendering of wet surfaces and clothing textures.
- − Fails to include the 'motion blur from passing cars' requested.
- − The bicycle chain and derailleur mechanics are visually nonsensical.
- − The framing feels too staged and centered for a 'candid' street photo.
Verdict: FLUX.1 Kontext [max] delivered a much more accurate interpretation of the prompt, successfully incorporating the motion blur and the requested 'imperfect framing' for a true candid feel. While Imagen 4.0 had impressive facial detail, it failed to include the motion blur and produced a more generic, centered composition with anatomical and mechanical errors in the hands and bicycle.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Exceptional photorealistic skin texture and lifelike eyes
- + Atmospheric use of shallow depth of field and bokeh sparks
- + Natural integration of warm torchlight reflections
- − The 'small beads' in the hair are minimally visible
- − Missing the leather straps and cloth underlayer mentioned in the prompt
Imagen 4.0 Generate 001
- + Perfect adherence to specific details like the beads in hair and leather straps
- + Exquisite engraving detail on the plate armor
- + Clear depiction of 'faint scars and dirt' as requested
- − Composition feels slightly more 'CG' and less like a photograph compared to A
- − The torch flame on the left edge is slightly distracting for a 'close portrait'
Verdict: FLUX.1 Kontext [max] produces a much more lifelike and cinematic portrait with incredible eye and skin detail, making it the stronger image visually. However, Imagen 4.0 Generate 001 followed the complex prompt requirements significantly better, including the specific requested textures for leather, cloth, and hair beads that FLUX overlooked.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent photo rendering with appetizing, consistent food items
- + Clean, professional document layout that looks like a real-world flyer
- + Good use of white space and hierarchy
- − Mixes serif and sans-serif fonts despite the 'bold sans-serif' prompt
- − Food variety is lacking as almost every photo is a pizza
- − Missing the specific section headers requested in the prompt
Imagen 4.0 Generate 001
- + Strong adherence to the 'sections' prompt, clearly labeling Appetizers, Pizza, and Mains
- + Modern, vibrant geometric accents that create a cohesive graphic style
- + Balanced 3x3 grid composition
- − Text is somewhat garbled compared to the cleaner fonts in Model A
- − The food photography is less consistent in lighting and style across the grid
Verdict: FLUX.1 Kontext creates a more realistic and professional-looking physical menu with high-quality food photography, though it fails to include the specific section categories. Imagen 4.0 Generate 001 follows the prompt's structural requirements much better, organizing the content into the requested sections with a modern, colorful grid, making it the better interpretation of the design brief despite slightly weaker text rendering.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent photorealistic texture on the meat and veggies
- + High-quality rendering of the fiery, glowing text
- + Dramatic ground-level embers and smoke effects
- − The burger is mostly assembled rather than 'exploded' as requested
- − Misinterpreted the starburst requirement for the price
Imagen 4.0 Generate 001
- + Perfectly captures the 'exploded' concept with vertically suspended layers
- + Accurately follows all layout constraints including the starburst for the price
- + Clean, professional advertisement composition
- − Lighting on the burger is a bit flat and less realistic than its competitor
- − Price text inside the starburst lacks the specific fiery/glowing effect requested
Verdict: FLUX.1 Kontext produces a more photorealistic image with impressive textures and lighting, but it fails to capture the 'exploded' layout requested in the prompt. Imagen 4.0 delivers much better prompt adherence by correctly deconstructing the burger layers and including the price starburst, making it the more effective advertisement overall.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent typography with a realistic chalk texture and natural smudging
- + Flawless text rendering of the requested menu items
- + Clean, professional composition that accurately reflects a café setting
- − The title is in print-style block letters rather than the 'elegant cursive' requested
Imagen 4.0 Generate 001
- + Natural wood grain texture on the chalkboard frame
- + Includes the specific menu items requested with correct pricing
- − Suffers from severe prompt leakage where instructions to the model are written out as text on the board
- − Numerous spelling errors in the extra text (e.g., 'Tittle', 'Berbs', 'veriations')
- − The writing style looks more like a digital marker than realistic chalk on a board
Verdict: FLUX.1 Kontext [max] produced a high-quality, professional image that perfectly captures the texture and aesthetics of a real chalkboard. While it missed the 'cursive' instruction for the title, it far outperformed Imagen 4.0 Generate 001, which mistakenly rendered parts of the prompt instructions as visible text and included multiple spelling errors.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Exceptional realism in textures, particularly the horse's coat and harness.
- + Cleaner composition with a cinematic lighting profile.
- + Superior rendering of fine details like the stitching on the saddle pad.
- − Followed the common interpretation of an astronaut riding a horse, ignoring the specific instruction 'horse on top, not vice versa'.
Imagen 4.0 Generate 001
- + Creative use of color and 'surreal' elements like the galaxy patterns on the horse.
- + Highly detailed background with orbital trails and nebulae.
- + Strong visual storytelling through the reflection in the helmet.
- − Failed to follow the contradictory logic instruction 'horse on top, not vice versa'.
- − The anatomy of the horse's front legs and hooves is slightly messy compared to Model A.
Verdict: Both models failed the negative constraint and logic trap in the prompt ('horse on top'), instead opting for the more logical 'astronaut riding a horse'. FLUX.1 Kontext [max] produced a much higher quality, more photorealistic image with better lighting, while Imagen 4.0 leaned into a more fantastical, digital-art aesthetic that felt less grounded.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent fur texture and photographic lighting.
- + Composition creates a more cinematic and intimate atmosphere.
- + The woman's expression is subtle and realistic.
- − The passenger is holding a phone to her ear instead of looking at it.
- − Only one paw is visible on the steering wheel.
- − The passenger is sitting in the front or middle, not clearly the back seat.
Imagen 4.0 Generate 001
- + Perfect adherence to specific instructions like 'both front paws on the steering wheel'.
- + Captures the 'inside a taxi' perspective effectively showing both front and back seats.
- + The passenger's interaction with the phone matches the prompt exactly.
- − The capybara's claws are slightly too long and sharp, looking somewhat unnatural.
- − The lighting is a bit flat compared to Model A.
- − The dashboard in the foreground is very plain.
Verdict: Imagen 4.0 Generate 001 followed the complex scene instructions much more accurately, correctly placing the passenger in the back seat looking at her phone and putting both paws on the wheel. While FLUX.1 Kontext [max] has a more realistic and cinematic photographic quality, it missed several specific positional and behavioral details requested in the prompt.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent typography style that matches the gothic aesthetic
- + Superior atmospheric lighting on the jack-o-lantern
- + Text is centered and perfectly legible
- − Repetitive layout on the location section at the bottom
- − The scroll banner is less distinct than requested
Imagen 4.0 Generate 001
- + Features a very clear scroll banner as requested
- + Detailed thorny border and visible twisted trees
- + The parchment element is integrated into the side design
- − The lighting is more cartoonish than cinematic
- − A large vertical scroll-pole on the left feels slightly out of place
- − The text character kerning on the scroll banner is a bit uneven
Verdict: FLUX.1 Kontext wins for its superior atmospheric and cinematic quality, producing an invitation that feels more authentic and formal. While Imagen 3 includes all requested elements like the scroll banner more literally, its overall art style is more vector-like and lacks the moody, vintage parchment depth of FLUX.1.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent adherence to all text requirements including 'JAPAN' and 'SUSHI'.
- + Perfectly captures the 'cartoon scene' and 'isometric miniature' aesthetic.
- + Clean, solid blue background as requested.
- − Missed the small flag icon.
- − The chopsticks are slightly floating/emerging from the base side.
Imagen 4.0 Generate 001
- + Higher realism in the sushi textures (PBR materials).
- + Good variety of sushi types on the plate.
- + Accurate miniature scale appearance.
- − Completely failed to include any of the requested text.
- − Background is grey/white instead of solid light blue.
- − Lacks the 'cartoon scene' stylization requested.
Verdict: FLUX.1 Kontext [max] followed the prompt instructions much more comprehensively, including the specific text and the cartoonish isometric style. Imagen 4.0 prioritized hyper-realistic rendering but ignored the background color requirements and all text-related instructions.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent soft lighting and atmosphere
- + Cute, unified character design
- + Strong interpretation of a 'lush wildflower meadow'
- − Static composition; animals are sitting rather than chasing/tumbling
- − The rabbit has slightly unnatural facial proportions
Imagen 4.0 Generate 001
- + Perfect dynamic movement showing animals tumbling and playing
- + Highly detailed dew drops and grass textures
- + Superior adherence to the specific action of the prompt
- − The fox looks significantly larger and more mature than a 'kit'
- − Digital-art look rather than 'hyper-photorealistic'
Verdict: FLUX.1 Kontext [max] creates a more cohesive and aesthetically pleasing photograph with beautiful golden lighting, but it misses the dynamic actions requested. Imagen 4.0 captures the 'tumbling' and 'chasing' aspect perfectly with a much more energetic composition, though it feels slightly more like a 3D illustration than a realistic photo. Imagen 4.0 is the winner for its superior prompt adherence and attention to small details like dew drops.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Perfectly captures the vintage aesthetic with woodblock-style texture and shading.
- + Excellent typography with authentic period-appropriate serifs.
- + Higher quality textured background that adds to the retro feel.
- − The 'E' in Caffè uses a circumflex-style accent rather than a grave accent requested by the spelling.
- − Composition is a bit vertically crowded compared to the clean layout of Model B.
Imagen 4.0 Generate 001
- + Very clean, modern vector execution that perfectly follows the 'minimalist' prompt.
- + Excellent spacing and visual balance between elements.
- + Correct accent usage on the 'È' of Caffè.
- − The banner for 'Est. 1720' is more of a simple ribbon and lacks the classic detail of Model A.
- − Less texture and 'vintage' character than requested, feeling more like a modern tech logo.
Verdict: FLUX.1 Kontext [max] delivers a superior vintage aesthetic with rich textures and classic typography that feels authentic to the year 1720. Imagen 4.0 provides a cleaner, more modern minimalist interpretation with better spacing, but fails to capture the 'retro' and 'textured' prompt requirements as convincingly as FLUX.1 Kontext [max].
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Features a more complex, artistic illustration of the lunar surface.
- + Captures the requested navy and muted red palette well.
- + Includes more detailed depictions of the rocket and lunar lander.
- − Numerous spelling errors in text (e.g., 'Lunear Modulle', 'Laarth', 'Unony').
- − Logical inconsistency with Saturn and multiple moons appearing in an Apollo 11 infographic.
- − Failed to follow the requested sequential step-by-step numbering.
Imagen 4.0 Generate 001
- + Excellent text rendering with no spelling errors.
- + Clearer, more logical infographic layout that follows the requested 'Steps' sequence.
- + Highly consistent flat-vector iconography and clean lines.
- − Included a duplicate, smaller rocket at the bottom left for no apparent reason.
- − Missing the final stages (Descent and Landing) requested in the prompt.
- − The 'Launch' section shows an Earth orbit icon rather than a unique launch-specific icon.
Verdict: Imagen 4.0 Generate 001 is the superior choice because it adheres to the requested vector infographic style and produces professional, legible text, despite missing the final two steps. FLUX.1 Kontext [max] creates more intricate illustrations but suffers from significant spelling errors and nonsensical elements like Saturn being included in an Earth-Moon trajectory chart.
Explore each model
Google's latest Imagen 4.0 text-to-image generation model with significantly better text rendering and overall image quality