Black Forest Labs' premium multimodal flow transformer with greatly improved prompt adherence and typography generation for in-context image generation and editing without compromise on speed
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
FLUX.1 Kontext [max]
#23 of 62 in Text-to-Image
Imagen 4.0 Ultra Generate 001
#34 of 62 in Text-to-Image
Where the votes landed
FLUX.1 Kontext [max]
100.0%
win rate
Ties
0.0%
Imagen 4.0 Ultra Generate 001
0.0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent handling of light refraction and shadows on the table surface
- + The glass cube has realistic thin-walled construction suitable for a container
- + Accurate placement of all requested elements including the plant behind the glass
- − The text on the book spine is garbled and unreadable
- − The blue sphere has a slightly rough, glittery texture rather than being a smooth sphere
Imagen 4.0 Ultra Generate 001
- + Superb rendering of the book with clear, legible text on the spine
- + Clean, high-quality aesthetic with a solid optical glass cube effect
- + Stronger overall composition and color saturation
- − The blue sphere appears to be floating rather than resting inside/on the base of the cube
- − The refraction of the plant through the glass is less complex than Model A
Verdict: Both models followed the prompt instructions perfectly regarding the spatial relationships of the objects. Model A (FLUX.1 Kontext) produced more realistic lighting and refraction patterns on the wood, but Model B (Imagen 4.0 Ultra) delivered a much cleaner image with perfectly legible text and a more premium photographic feel.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent depiction of rain with visible streaks and atmospheric 'wet' feel
- + Superior reflections on the pavement including neon signage colors
- + Stronger adherence to the 'imperfect framing' and 50mm lens aesthetics
- − The man's skin texture appears slightly smoothed compared to the prompt's request for 'natural textures'
- − The bicycle chain and wheel mechanics have some structural AI artifacts
Imagen 4.0 Ultra Generate 001
- + Exceptional skin texture and facial detail with high realism
- + The man's pose and character feel more authentically 'Japanese' and elderly
- + High-quality rendering of water droplets on the jacket fabric
- − The bike's orientation and relationship to the wall feel physically awkward
- − The 'rain' looks more like static mist rather than a rainy day on the street
- − Lacks the motion blur from cars requested in the prompt
Verdict: FLUX.1 Kontext [max] captures the overall mood and technical requirements of the prompt much better, specifically the rain, reflections, and motion blur. However, Imagen 4.0 Ultra Generate 001 provides much more realistic and detailed skin textures and character rendering, despite failing to include the requested car motion and convincing rain effects. FLUX.1 feels like a more coherent 'candid street photo' while Imagen 4.0 feels like a higher-quality portrait that missed some environment cues.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Exceptional skin texture with lifelike pores and subtle dirt
- + Hyper-realistic eye rendering with complex reflections
- + Beautifully integrated golden-hour lighting and bokeh
- − Missed the request for beads in the hair braids
- − The portrait is so close that many requested elements like leather straps are barely visible
Imagen 4.0 Ultra Generate 001
- + Strong adherence to all prompt details including beads, leather straps, and armor engravements
- + Effective use of warm torchlight to define the armor's shape
- + Balanced composition that captures the 'battle-worn' aesthetic with clear scars
- − Skin texture appears slightly more synthetic and smoothed compared to Model A
- − The torch in the background is a bit distracting and lacks realistic integration with the depth of field
Verdict: FLUX.1 Kontext [max] produces a significantly more lifelike and cinematic portrait with superior skin and eye details, but it ignores the specific request for hair beads. Imagen 4.0 Ultra follows the prompt's instructions much more closely, including the beads, straps, and battle scars, though the overall image feels more like a high-end digital render rather than a photograph.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent professional menu layout that feels realistic.
- + Good use of hierarchy with clear headers and item descriptions.
- + High-quality, appetizing food photography that fits the restaurant theme.
- − Failed to include the specific sections requested (Appetizers, Pizza, Mains), mostly showing pizza.
- − The text, while well-placed, is mostly illegible gibberish.
Imagen 4.0 Ultra Generate 001
- + Strictly followed the prompt by including specific sections for Appetizers, Pizza, and Mains.
- + Perfectly executed the grid layout requested.
- + Text rendering for titles and prices is very clear and legible.
- − The layout is a bit repetitive with all photos being the same size in a simple grid.
- − The spacing at the bottom of the page is uneven.
Verdict: While FLUX.1 Kontext creates a more 'ready-to-use' professional design aesthetic, Imagen 4.0 Ultra is the superior choice for prompt adherence. Imagen 4.0 correctly implemented all three requested food categories and clear text, whereas FLUX.1 predominantly featured pizza and used illegible text.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent photorealistic texture on the meat patty.
- + Includes stylized embers and a fiery ground for environmental depth.
- + Clear and legible typography for the main title.
- − Failed the 'exploded' instruction as the core burger is still stacked.
- − Price is not in a starburst as requested.
- − The '€6.99' price uses a comma instead of a period, which may not be intended.
Imagen 4.0 Ultra Generate 001
- + Perfectly executed the 'exploded' burger concept with vertically separated components.
- + Successfully included the price in a fiery starburst element as requested.
- + The text 'MAGIC BURGER' has a more integrated fiery, glowing effect.
- − The sesame seeds on the bun appear a bit repetitive/uniform.
- − Slightly less gritty/realistic texture on the meat compared to the competitor.
Verdict: Imagen 4.0 Ultra is the clear winner as it followed every detail of the complex prompt, specifically the 'exploded' layout and the starburst requirement for the price. While FLUX.1 Kontext produced a high-quality image, it failed several key compositional instructions, providing a mostly intact burger instead.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent chalk texture with realistic smudges and dust on the board.
- + Perfect spelling and completion of the truncated prompt text.
- + Strong handwriting style consistency that feels authentic to a café.
- − The title is in print-style block letters rather than the 'elegant cursive' requested.
- − The lighting is a bit flat compared to the dramatic lighting in the alternative.
Imagen 4.0 Ultra Generate 001
- + Successfully incorporated a more slanted, script-like style for the title.
- + Pleasant atmospheric lighting and composition with the brick background.
- + High text legibility and clean alignment.
- − The text looks more like a digital chalk font than actual hand-drawn chalk.
- − Missing the gritty texture and grain found in a real chalkboard environment.
- − Minor grammatical change in the footer from the source text.
Verdict: FLUX.1 Kontext [max] wins on realism, providing a chalkboard that looks like it was actually written on and wiped, with perfect adherence to the text content. While Imagen 4.0 Ultra attempted the cursive title more successfully, its text rendering looks too much like a digital overlay rather than natural chalk.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent photorealism in the textures of the horse hair and spacesuit.
- + Dynamic sense of motion with the horse's gallop pose in zero gravity.
- + Cinematic lighting consistent with a distant sun source.
- − Failed the paradoxical negative constraint 'horse on top, not vice versa'.
- − The horse lacks any life-support or space gear, reducing the 'highly detailed' immersion in a space setting.
Imagen 4.0 Ultra Generate 001
- + Creative addition of life-support gear for the horse, including a visor and illuminated hooves.
- + Strong composition with a detailed asteroid base and background nebulae.
- + Good adherence to the surreal and cinematic style keywords.
- − Failed the paradoxical negative constraint 'horse on top, not vice versa'.
- − Lower resolution in the horse's fur texture compared to Model A.
Verdict: Both models failed the complex prompt instruction to have the horse on top of the astronaut, instead defaulting to the standard image of an astronaut riding a horse. FLUX.1 Kontext [max] has superior technical quality and realism, while Imagen 4.0 Ultra Generate 001 shows more creativity in the character design by giving the horse its own space-faring equipment.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent photorealism with shallow depth of field and natural lighting.
- + Text rendering on the taxi cap is subtle and realistic.
- + Capybara fur texture is highly detailed and believable.
- − The passenger appears to be on a phone call rather than just looking at the screen as requested.
- − The capybara's paws are partially obscured by the frame and wheel.
Imagen 4.0 Ultra Generate 001
- + Strong composition that clearly shows both characters and the interior layout.
- + Accurately depicts the capybara with 'both front paws on the steering wheel'.
- + Captures the bored, normal expression of the businesswoman well.
- − The anatomy of the capybara's paws is strange, appearing more like sharp bird-like talons.
- − The overall lighting and texture look more like a high-quality 3D render than a photorealistic image.
- − The cap is floating slightly above the capybara's head.
Verdict: FLUX.1 Kontext [max] produces a significantly more photorealistic image with cinematic lighting and natural textures, though it misses some framing details for the paws. Imagen 4.0 Ultra Generate 001 provides a better composition for the prompt's specific actions, but the texture work and anatomical anomalies on the paws make it look artificial compared to FLUX.1.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent typography style that matches the gothic aesthetic
- + Atmospheric, moody lighting and a rich parchment texture
- + High level of detail in the thorny border and webs
- − Repetitive layout of text lines at the bottom
- − Replaced the central scroll banner with a flattened banner at the bottom
Imagen 4.0 Ultra Generate 001
- + Perfect text accuracy and clean layout of event details
- + Highly detailed thorns and spiderweb border that stands out
- + Matches the 'scroll banner' requirement better than Image A
- − The lighting is a bit too bright/saturated for a 'dark parchment' vintage look
- − The composition feels slightly more like a digital illustration than a vintage poster
Verdict: Both models followed the prompt instructions very well. FLUX.1 Kontext [max] produced a more authentic 'vintage gothic' atmosphere with superior texture, though it struggled with the specific text layout at the bottom. Imagen 4.0 Ultra provided perfect text rendering and a better interpretation of the scroll banner, making it more functional as an actual invitation.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent typography with a cohesive 3D cartoon style
- + Soft, beautiful PBR lighting and textures
- + Simple, elegant composition that feels like a polished game asset
- − Missed the small flag icon requested in the prompt
- − The diorama base is a bit plain compared to the sushi plate
Imagen 4.0 Ultra Generate 001
- + Includes all prompted elements including the flag icon
- + Higher variety of sushi types and garnishes
- + Perfectly executed isometric perspective and diorama layering
- − The text is plain and lacks the 3D 'cartoon' feel of the rest of the image
- − Some sushi items look slightly floating or disconnected from the plate
Verdict: Both models followed the prompt well, but they excelled in different areas. FLUX.1 Kontext prioritized the '3D cartoon' aesthetic with stylized typography and soft textures, whereas Imagen 4.0 Ultra captured more specific details like the flag icon and a more complex variety of sushi. Imagen 4.0 Ultra is the winner for its comprehensive inclusion of all prompt elements and superior miniature diorama presentation.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent adherence to the 'golden sunrise' and 'god rays' lighting request
- + Highly realistic fur textures and soft lighting that feels cohesive
- + Natural integration of the animals within the meadow
- − The animals are sitting still rather than 'playfully chasing and tumbling'
- − The bunny has a slightly unnatural, human-like facial expression
Imagen 4.0 Ultra Generate 001
- + Successfully captures the action of 'playfully chasing' and 'tumbling'
- + Dynamic composition with varied butterfly designs
- + Includes the requested 'dew sparkles' on the grass blades
- − Looks more like a digital illustration or greeting card than 'hyper-photorealistic'
- − The lighting is oversaturated and lacks the natural depth of Image A
- − The fox paws and proportions look slightly stylized/cartoonish
Verdict: FLUX.1 Kontext [max] creates a much more photorealistic and atmospheric image with beautiful lighting, though it fails to capture the dynamic action requested in the prompt. Imagen 4.0 Ultra Generate 001 follows the 'playful chasing' instruction much better but falls short on the realism requirement, resulting in a more illustrative aesthetic. FLUX.1 is the likely winner for its superior visual quality and technical execution of the requested environment.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Excellent typography with a bold, professional layout
- + Strong subtle texture that matches the vintage aesthetic
- + Perfect adherence to all prompt elements including the banner and cloche
- − The accent mark on 'CAFFÈ' is slightly stylized in a way that looks like a hat
Imagen 4.0 Ultra Generate 001
- + Clean vector lines and elegant minimalist style
- + Good color palette adherence to warm brown and cream
- + Balanced layout with clear separation of elements
- − Incorrect accent orientation on 'CAFFÈ' (grave used instead of acute/circumflex, though both images struggle with precise Italian grammar)
- − The 'Est. 1720' text is slightly off-center within the banner
Verdict: FLUX.1 Kontext [max] produced a superior, more cohesive logo that feels like a real vintage brand, particularly due to the beautiful texture and integrated typography. Imagen 4.0 Ultra Generate 001 is clean and technically sound, but much more basic in its execution and lacks the character found in the first image.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
FLUX.1 Kontext [max]
- + Captures the requested flat-vector art style and NASA-inspired palette effectively.
- + Visualizes a cohesive lunar mission scene with a clear launch vehicle and landing module.
- + High clarity and resolution in the vector elements.
- − Labels include significant spelling errors and logic failures (e.g., labeling Earth as 'Moon').
- − Does not follow the 6-step linear infographic structure requested in the prompt.
- − Includes a random planet with rings (Saturn/Uranus) that was not part of the mission or prompt.
Imagen 4.0 Ultra Generate 001
- + Follows the specific step-by-step infographic layout requested in the prompt.
- + Accurately applies the requested color palette throughout the design.
- + Excellent text legibility for headers despite nonsense filler text.
- − Uses nonsensical 'lorem ipsum' style labels for all steps.
- − The icons are abstract and often fail to represent the specific objects requested (e.g., the Saturn V or the Moon).
- − The central graphic elements are somewhat cluttered and visually confusing.
Verdict: Both models struggled to fully realize the technical requirements of an infographic. FLUX.1 Kontext [max] produced a much more visually appealing and recognizable illustration but failed significantly on factual accuracy and following the requested 6-step structure. Imagen 4.0 Ultra followed the layout instructions much better, but its icons were less representative of the Apollo mission. FLUX is slightly preferred for its aesthetic quality, but both failed at providing a functional infographic.
Explore each model
Google's Imagen 4.0 Ultra model offering the highest fidelity and resolution for professional-grade image generation