Head to head
Esc

Models · slot A

to navigate to pick

FLUX.1 Kontext [max] Black Forest Labs Imagen 3.0 Generate 002 Google

Settled by community votes across 6 shared challenges, with an AI judge weighing in on each.

FLUX.1 Kontext [max]

23.7 arena score

#23 of 62 in Text-to-Image

Skill signature · Text-to-Image

Imagen 3.0 Generate 002

21.0 arena score

#36 of 62 in Text-to-Image

Vote tally

Where the votes landed

FLUX.1 Kontext [max]

0%

win rate

Ties

0%

Imagen 3.0 Generate 002

0%

win rate

Shared challenges 6

Challenge by challenge

The strongest take from each model on every shared challenge, with the AI judge's read.

Modern Clean Menu

Text-to-Image

“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”

FLUX.1 Kontext [max]
Imagen 3.0 Generate 002

AI Judge Analysis

FLUX.1 Kontext [max]

  • + Clean, professional white space usage
  • + Consistent photography style of the food
  • + Elegant, high-end casual dining feel
  • Text is largely gibberish
  • Food variety is limited mostly to pizza despite different category labels
  • Visual hierarchy is slightly cluttered by long lists of garbled text

Imagen 3.0 Generate 002

  • + Stronger adherence to the 'grid' prompt
  • + Clear section headers like 'APPETIEES' and 'MAINS'
  • + High-contrast, vibrant food photography with more variety in dishes
  • Text rendering contains several spelling errors
  • The layout is quite dense with less 'minimalist' white space compared to A
  • Repetitive section titles (multiple 'MAINS' and 'PIZZA' sections)

Verdict: FLUX.1 Kontext [max] creates a more cohesive and professional-looking physical menu with a sophisticated minimalist aesthetic, while Imagen 3.0 Generate 002 follows the 'grid' instruction much more literally and provides better food variety. FLUX.1 Kontext [max] is the preferred choice as its layout feels more authentic to a real-world dining establishment despite the garbled text.

Magic Burger Explosion: Fiery Photorealism Challenge

Text-to-Image

“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”

FLUX.1 Kontext [max]
Imagen 3.0 Generate 002

AI Judge Analysis

FLUX.1 Kontext [max]

  • + Excellent text rendering and typography for the primary and secondary messages.
  • + High photorealism in the textures of the meat and vegetables.
  • + Strong adherence to the price and fiery glow requirement.
  • The 'exploded' effect is weak, with the burger remaining mostly assembled.
  • The starburst element for the price is missing, replaced by a floating sun-like icon.

Imagen 3.0 Generate 002

  • + Perfectly captures the 'exploded' view with clear separation of all components.
  • + Includes the starburst element for the price as requested.
  • + Dynamic composition with a strong sense of motion.
  • Garbled text inside the starburst above the price.
  • The lighting on the ingredients feels slightly more digital and less photorealistic than Image A.
  • The 'fiery' glow on the text is less pronounced.

Verdict: FLUX.1 Kontext [max] produces a more polished advertisement with superior text rendering and food photorealism, but fails to deliver the 'exploded' composition requested. Imagen 3.0 Generate 002 follows the structural instructions much better, providing a true exploded view and the starburst element, though it suffers from some text artifacts in the fine print.

Chalkboard Menu

Text-to-Image

“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”

FLUX.1 Kontext [max]
Imagen 3.0 Generate 002

AI Judge Analysis

FLUX.1 Kontext [max]

  • + Perfect spelling for all requested menu items and numbers.
  • + Highly realistic chalk texture with smudge marks on the board.
  • + Successful completion of the truncated prompt for 'Brown Butter Chocolate Chip Cookies'.
  • The title uses a print style rather than the requested 'elegant cursive' handwriting.
  • The background café environment is slightly blurred/messy.

Imagen 3.0 Generate 002

  • + Attempts more stylistic variety in the lettering, including the requested cursive elements.
  • + The framing of the chalkboard is clean and well-centered.
  • Numerous catastrophic spelling errors like 'Risotttog', 'Octpsuin', and 'Choouip'.
  • Repeated text lines for items and prices, creating a redundant and cluttered layout.
  • The text looks more like a digital chalk-font overlay than organic handwriting.

Verdict: FLUX.1 Kontext [max] is the clear winner as it produced perfectly spelled, legible text that accurately completed the truncated prompt. While it ignored the 'cursive' instruction for the title, its overall realism and accuracy far surpass Imagen 3.0 Generate 002, which suffered from severe spelling hallucinations and repetitive lines of text.

The Reversed Rodeo

Text-to-Image

“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”

FLUX.1 Kontext [max]
Imagen 3.0 Generate 002

AI Judge Analysis

FLUX.1 Kontext [max]

  • + Excellent photographic lighting and crisp texture on the horse's coat.
  • + Clear, high-detail rendering of the astronaut suit and visor reflections.
  • + Strong composition with the planet in the background providing depth.
  • Failed the negative constraint; the astronaut is riding the horse, not vice versa.
  • The harness and reins are physically inconsistent and overlap weirdly with the horse's neck.

Imagen 3.0 Generate 002

  • + Beautiful, colorful nebulae in the background enhance the surreal space theme.
  • + Realistic proportions and dynamic posing of the horse in 0g.
  • + Good use of space-specific equipment details on the astronaut suit.
  • Failed the negative constraint; the astronaut is riding the horse instead of the horse on top.
  • The horse's front legs look slightly distorted and unnatural where they meet the chest.

Verdict: Both models failed the specific prompt instruction to place the 'horse on top' of the astronaut, choosing instead to generate the classic 'astronaut riding a horse' trope. FLUX.1 Kontext [max] produced a sharper image with better lighting, while Imagen 3.0 Generate 002 had a more vibrant and visually interesting space background; FLUX.1 is slightly preferred for its superior clarity and realistic textures.

Isometric Miniature Diorama Scenes

Text-to-Image

“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”

FLUX.1 Kontext [max]
Imagen 3.0 Generate 002

AI Judge Analysis

FLUX.1 Kontext [max]

  • + Excellent typography with a playful 3D cartoon style.
  • + Beautiful lighting and shading on the fish textures.
  • + Very clean and centered composition.
  • Completely missed the requested flag icon.
  • The text is brown rather than a crisp white/neutral, which slightly clashes with the background.
  • The diorama base is very simple, lacks the 'raised' legs look often associated with sushi platters.

Imagen 3.0 Generate 002

  • + Successfully included the Japanese flag icon as requested.
  • + Perfect adherence to the 45-degree isometric perspective.
  • + More variety in sushi types (ikura, shrimp, nigiri) creating a better diorama feel.
  • The text is aligned to the top-left rather than top-center as requested.
  • The lighting is a bit flat compared to the more dynamic shadows in the other model.
  • The chopsticks are floating/passing through the chopstick rest instead of sitting on it.

Verdict: Both models performed well, but Image B (Imagen 3.0) is the winner for following more specific prompt details like the flag icon and the isometric perspective. While Image A (FLUX.1 Kontext) has more appealing lighting and superior text rendering, it missed one of the core objects in the prompt and chose a top-left alignment over center alignment.

Adorable Baby Animals in Sunny Meadow

Text-to-Image

“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”

FLUX.1 Kontext [max]
Imagen 3.0 Generate 002

AI Judge Analysis

FLUX.1 Kontext [max]

  • + Excellent depiction of golden hour lighting and god rays as requested.
  • + The animal species are well-defined and look very 'fluffy' and soft.
  • + Captures the whimsical and wholesome vibe perfectly.
  • The bunny's face is overly stylized and borders on looking like a stuffed toy rather than a photorealistic animal.
  • The composition is a bit static with the animals just sitting in a row.

Imagen 3.0 Generate 002

  • + Successfully captures the 'tumbling' and 'playful' aspect of the prompt with the bunny on its back.
  • + Superior photorealism in the fur textures and facial features of all four animals.
  • + Better interaction between the subjects and the environment.
  • The 'god rays' are less distinct compared to the other model.
  • The kitten's eye expression is a bit wide and slightly unnatural.

Verdict: While both models followed the prompt well, Imagen 3.0 Generate 002 is the winner due to its superior photorealism and dynamic composition that actually shows the animals 'tumbling' together. FLUX.1 Kontext [max] has beautiful lighting, but the bunny looks significantly more artificial and the pose is more generic.

Next steps

Explore each model