Head to head
Esc

Models · slot A

to navigate to pick

FLUX.1 [schnell] Black Forest Labs Imagen 4.0 Generate 001 Google

Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.

FLUX.1 [schnell]

18.7 arena score

#48 of 62 in Text-to-Image

Skill signature · Text-to-Image

Imagen 4.0 Generate 001

17.1 arena score

#54 of 62 in Text-to-Image

Vote tally

Where the votes landed

FLUX.1 [schnell]

0%

win rate

Ties

0%

Imagen 4.0 Generate 001

0%

win rate

Shared challenges 13

Challenge by challenge

The strongest take from each model on every shared challenge, with the AI judge's read.

Geometric Composition

Text-to-Image

“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”

FLUX.1 [schnell]
Imagen 4.0 Generate 001

AI Judge Analysis

FLUX.1 [schnell]

  • + Includes an extra blue sphere on top of the book which adds visual interest
  • + Very high clarity in the glass reflections and wood grain
  • + The plant is very well integrated into the background and partially visible through the glass
  • Failed the spatial instruction by adding a second blue sphere outside of the cube
  • The lighting feels slightly more surgical and less atmospheric than requested

Imagen 4.0 Generate 001

  • + Strictly followed the quantity of items in the prompt with one sphere and one book
  • + Excellent soft window lighting that feels natural and directional from the left
  • + High physical realism in the texture of the red book's cover
  • The sphere appears to be floating mid-air inside the cube without resting on the bottom
  • The glass cube looks more like a mirror on some faces, obscuring parts of the plant behind it

Verdict: Imagen 4.0 followed the prompt more accurately by including only one blue sphere as requested, whereas FLUX.1 added a second sphere on top of the book. While FLUX.1 has slightly higher rendering detail in the wood and plant, Imagen 4.0 captures the requested 'soft window light' much more effectively and adheres better to the composition instructions.

Candid Street Photography

Text-to-Image

“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”

FLUX.1 [schnell]
Imagen 4.0 Generate 001

AI Judge Analysis

FLUX.1 [schnell]

  • + Features excellent environmental reflections on the wet pavement.
  • + The background cars and architecture feel authentic to a Japanese urban setting.
  • + Good balance of colors with the red bicycle standing out against the muted street.
  • Fails to show actual 'repairing' as the man is just holding the handlebars.
  • The background cars are stationary and in-focus, missing the requested motion blur.
  • The man's hands have structural issues where they grip the handlebars.

Imagen 4.0 Generate 001

  • + Excellent adherence to the 'repairing' action, showing him working on the chain with a tool.
  • + Superior skin texture and fine details like raindrops on the jacket.
  • + Better composition that feels like a 'candid' street photo with the requested shallow depth of field.
  • The transition between the man's head and the background has some minor masking artifacts.
  • The bicycle anatomy near the pedals and chain is slightly confused.
  • Less noticeable motion blur on the background vehicles than requested.

Verdict: Imagen 4.0 Generate 001 is the clear winner for its superior interpretation of the 'repairing' action and its impressive detail in textures, such as the raindrops and skin. FLUX.1 [schnell] captures a better overall city atmosphere but fails the core prompt by having the subject simply stand with the bike rather than fix it.

Fantasy Warrior

Text-to-Image

“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”

FLUX.1 [schnell]
Imagen 4.0 Generate 001

AI Judge Analysis

FLUX.1 [schnell]

  • + Extremely lifelike eyes with realistic moisture and reflection
  • + Hyper-detailed skin texture and convincing faint scars
  • + Intense, dramatic lighting and facial expression
  • Missed the request for 'beads' in the hair
  • The armor engraving is very dark and difficult to see compared to Model B

Imagen 4.0 Generate 001

  • + Excellent adherence to all prompt elements including beads and decorative engraving
  • + Clear, high-quality texture on leather straps and metal
  • + Effective use of bokeh sparks and warm torchlight
  • Skin texture looks slightly more 'digital' or smoothed compared to Model A
  • The torch flame is very close to the face but the highlights on the skin are somewhat muted

Verdict: FLUX.1 [schnell] produces a much more emotionally intense and photorealistic face, but it ignores the specific request for beads in the hair. Imagen 4.0 follows the prompt instructions more comprehensively, including the intricate armor engravings and the specific hair accessories, making it a better technical match for the request.

Modern Clean Menu

Text-to-Image

“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”

FLUX.1 [schnell]
Imagen 4.0 Generate 001

AI Judge Analysis

FLUX.1 [schnell]

  • + Strong minimalist aesthetic with clean white space
  • + Good use of bold sans-serif headers
  • + Clear hierarchy for Appetizers, Pizza, and Mains sections
  • Repetitive food photos that don't match the specific section labels
  • Body text is illegible gibberish
  • Image resolution within the grid is low and blurry

Imagen 4.0 Generate 001

  • + High-quality, vibrant food photography that looks appetizing
  • + Effective use of grid layout and colorful geometric accents
  • + Legible prices and distinct menu section headers
  • Text characters are scrambled and incoherent
  • Some minor alignment issues with the geometric frames

Verdict: Imagen 4.0 significantly outperforms FLUX.1 [schnell] in visual quality and adherence to the 'colorful' and 'vibrant' request. While both models fail to produce meaningful English text, Imagen 4.0 creates a much more professional layout with distinct food items, whereas FLUX.1 [schnell] uses repetitive, lower-resolution imagery that does not match the category labels.

Magic Burger Explosion: Fiery Photorealism Challenge

Text-to-Image

“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”

FLUX.1 [schnell]
Imagen 4.0 Generate 001

AI Judge Analysis

FLUX.1 [schnell]

  • + Strong bokeh effect in the background with realistic fire elements.
  • + High-quality texture on the bun and cheese melt.
  • + Vibrant lighting that creates a sense of heat.
  • Failed to render the full text correctly, missing the 'M' in 'MAGIC'.
  • The starburst and price are cluttered and contain repeated text errors.
  • The burger is mostly intact rather than 'exploded' as requested.

Imagen 4.0 Generate 001

  • + Excellent adherence to the 'exploded' component requirement, showing clear separation of layers.
  • + Perfect text rendering for all requested strings including the fiery glow effect.
  • + Creative and clean integration of the starburst price tag.
  • The overall layout is slightly more vertical, which might feel less 'dynamic' than a radial explosion.
  • The lighting on the lettuce and onion rings feels slightly flatter compared to the bun.

Verdict: Imagen 4.0 significantly outperformed FLUX.1 [schnell] by accurately following all prompt instructions, particularly regarding the 'exploded' anatomy of the burger and the specific text strings. FLUX.1 failed several text elements, missing the first letter of the title and creating a cluttered price starburst, while also failing to properly separate the burger components in mid-air.

Chalkboard Menu

Text-to-Image

“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”

FLUX.1 [schnell]
Imagen 4.0 Generate 001

AI Judge Analysis

FLUX.1 [schnell]

  • + The handwriting style looks very authentic to a real chalkboard.
  • + The composition feels natural for a café setting with background depth.
  • Several spelling errors in the menu items, such as 'Taffle' and 'Octtoopus'.
  • It fails to render the specific handwriting styles requested, like the elegant cursive for the title.

Imagen 4.0 Generate 001

  • + Successfully renders specific menu items like 'Truffle Mushroom Risotto' with near-perfect spelling.
  • + Captures the contrast between the title font and the item font more effectively.
  • Includes instructional prompt text literally on the board, such as 'Tittle' and 'Footer'.
  • The layout is cluttered and contains repetitive, nonsensical sentences derived from the prompt.

Verdict: While FLUX.1 [schnell] creates a more believable and aesthetically pleasing café scene, it struggles significantly with spelling and reading comprehension. Imagen 4.0 has superior spelling for the main items, but it mistakenly includes prompt instructions as text on the board and creates a messy layout. FLUX.1 [schnell] is likely the winner for users wanting a realistic image, though Imagen 4.0 is better for raw text accuracy.

The Reversed Rodeo

Text-to-Image

“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”

FLUX.1 [schnell]
Imagen 4.0 Generate 001

AI Judge Analysis

FLUX.1 [schnell]

  • + Successfully followed the difficult spatial instruction of having the horse on top.
  • + Creates a surreal and cinematic atmosphere with warm lighting.
  • + The composition feels unique and matches the 'surreal' prompt.
  • The horse has two heads/torsos fused together, which is a major anatomical artifact.
  • The astronaut's legs and lower body are awkwardly cropped or missing.

Imagen 4.0 Generate 001

  • + Excellent visual quality and sharp details on the astronaut suit and horse hair.
  • + Beautiful color palette with vibrant celestial effects.
  • + High anatomical accuracy for both the horse and human.
  • Completely failed the negative constraint to have the horse on top of the astronaut.

Verdict: FLUX.1 [schnell] is the only model that attempted the specific 'horse on top' spatial arrangement, though it suffered from significant anatomical errors like a double-headed horse. Imagen 4.0 Generate 001 produced a much higher quality image with better lighting and details, but it ignored the core instruction to reverse the typical horse-riding roles. Because FLUX.1 [schnell] specifically followed the surreal prompt logic while Imagen 4.0 failed that primary instruction, FLUX.1 [schnell] is the winner for prompt adherence.

The Capybara Taxi Driver

Text-to-Image

“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”

FLUX.1 [schnell]
Imagen 4.0 Generate 001

AI Judge Analysis

FLUX.1 [schnell]

  • + Excellent fur texture on the capybara
  • + More cinematic lighting and depth of field
  • + Accurate depiction of a bored expression on the human passenger
  • The capybara's paws are not both on the steering wheel
  • Image composition feels slightly crowded due to the close-up angle

Imagen 4.0 Generate 001

  • + Perfect adherence to the instruction for both front paws on the wheel
  • + Clear and accurate rendering of the 'TAXI' sign and professional driver cap
  • + Well-defined composition showing both front and back seats clearly
  • The interior plastic and dashboard have a slightly 'CGI' smooth look
  • The capybara's facial anatomy looks a bit stiff compared to Model A

Verdict: While FLUX.1 [schnell] (Model A) creates a more photorealistic and moody atmosphere with better lighting, Imagen 4.0 (Model B) followed the specific prompt instructions more accurately, particularly regarding the positioning of the capybara's paws on the steering wheel and the inclusion of the exterior taxi sign. Model B's clearer composition makes the narrative of the scene easier to read at a glance.

The Halloween Invitation

Text-to-Image

“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”

FLUX.1 [schnell]
Imagen 4.0 Generate 001

AI Judge Analysis

FLUX.1 [schnell]

  • + Strong moody atmosphere and cinematic lighting.
  • + Good placement of the central jack-o-lantern.
  • + Captures the vintage gothic aesthetic well.
  • Several severe text hallucinations and spelling errors in the details section.
  • Text formatting is cluttered and repetitive.
  • Fails to include a small scroll banner as requested.

Imagen 4.0 Generate 001

  • + Exceptional text rendering with perfect spelling of all requested phrases.
  • + Highly detailed border featuring the requested webs and thorns.
  • + Clean, professional layout that feels like a completed graphic design product.
  • The parchment/scroll element on the left looks a bit clunky and cuts off the border.
  • Lighting is slightly more illustrated than 'cinematic' dark parchment.

Verdict: Imagen 4.0 Generate 001 is the clear winner due to its superior text rendering and adherence to layout instructions, correctly spelling all event details and the scroll banner. FLUX.1 [schnell] struggles significantly with the text, producing multiple typos and repetitive lines that make the invitation unusable.

Isometric Miniature Diorama Scenes

Text-to-Image

“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”

FLUX.1 [schnell]
Imagen 4.0 Generate 001

AI Judge Analysis

FLUX.1 [schnell]

  • + Excellent adherence to the isometric miniature 3D cartoon style requested.
  • + Followed all text requirements including 'JAPAN' and the flag icon.
  • + Clean, solid blue background and diorama base meet all layout specifications.
  • Failed to include the word 'SUSHI' below the 'JAPAN' text.
  • The sushi piece has an odd red graphic on top that looks a bit messy.

Imagen 4.0 Generate 001

  • + High visual quality with realistic textures on the salmon and roe.
  • + Good interpretation of the 'diorama base' with a cylindrical pedestal.
  • Completely ignored all text requirements ('JAPAN', 'SUSHI', and flag).
  • Failed to use the requested solid light blue background.
  • Does not follow the 45 degree top-down isometric view as strictly as requested.

Verdict: FLUX.1 [schnell] followed the stylistic and composition instructions much more closely, successfully rendering the isometric viewpoint, the diorama base, and the specific text/icon elements. While Imagen 4.0 produced a more realistic-looking set of sushi, it failed almost every specific negative and affirmative constraint regarding the background, text, and overall layout.

Adorable Baby Animals in Sunny Meadow

Text-to-Image

“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”

FLUX.1 [schnell]
Imagen 4.0 Generate 001

AI Judge Analysis

FLUX.1 [schnell]

  • + Excellent soft lighting and bokeh effect.
  • + High level of detail in the fur texture.
  • + Cohesive and warm color palette.
  • Failed to include a rabbit, instead showing two cat-like creatures.
  • Anatomy of the middle-right animal is ambiguous and lacks distinct species traits.
  • Less playful 'tumbling' action compared to the other model.

Imagen 4.0 Generate 001

  • + Successfully included all four requested animals: puppy, kitten, bunny, and fox.
  • + Captures the 'tumbling' and 'chasing' action perfectly with dynamic poses.
  • + Includes requested 'god rays' and 'dew sparkles' explicitly.
  • The lighting looks somewhat artificial and illustrative rather than photorealistic.
  • The composition feels slightly cluttered with the density of flowers.
  • Perspective on the fox's hind legs is slightly awkward.

Verdict: While FLUX.1 [schnell] offers superior textures and a more convincing photorealistic lighting style, it failed the prompt's core requirement by omitting the baby bunny. Imagen 4.0 successfully rendered all four distinct animals and captured the energetic, playful atmosphere described in the prompt, making it the more accurate representation despite its slightly more digital look.

Vintage Cafe Logo

Text-to-Image

“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”

FLUX.1 [schnell]
Imagen 4.0 Generate 001

AI Judge Analysis

FLUX.1 [schnell]

  • + Elegant vector emblem style
  • + Good use of texture and vintage coloring
  • Serious spelling errors: 'Framilan' instead of 'Florian' and '7720' instead of '1720'
  • Missing the steam element from the prompt

Imagen 4.0 Generate 001

  • + Perfect text rendering of the name and date
  • + Includes the steam element and the banner as requested
  • + Clean, minimalist composition perfectly suited for a logo
  • The line weight on the steam is slightly disconnected from the body of the logo

Verdict: While FLUX.1 [schnell] creates a nice vintage aesthetic, it fails significantly on the prompt's specific text requirements, misspelling the name as 'Framilan' and the date as '7720'. Imagen 4.0 Generate 001 provides a much better interpretation by following all prompt instructions, including the steam and specific text, resulting in a cleaner and more professional minimalist logo.

Apollo 11: Journey to Tranquility

Text-to-Image

“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”

FLUX.1 [schnell]
Imagen 4.0 Generate 001

AI Judge Analysis

FLUX.1 [schnell]

  • + Successfully follows the color palette requested.
  • + Maintains a clean, abstract infographic layout.
  • Text is entirely illegible gibberish.
  • The central rocket icon is malformed with two nose cones.
  • Fails to clearly represent the specific 6-step sequence requested.

Imagen 4.0 Generate 001

  • + Excellent text rendering and typography for all labels.
  • + Includes a highly detailed and accurate Saturn V rocket vector.
  • + Clearly follows the requested iconography for orbits and trajectories.
  • Fails to include steps 5 and 6 (Descent and Landing) as specified in the prompt.
  • The 'Translunar' label is placed next to a small moon icon rather than the trajectory arc.

Verdict: Imagen 4.0 Generate 001 is the clear winner due to its superior text legibility and high-quality vector illustrations, even though it missed the final two steps of the sequence. FLUX.1 [schnell] produced a more abstract design, but it was severely undermined by nonsensical text and a broken rocket icon.

Next steps

Explore each model