Head to head
Esc

Models · slot A

to navigate to pick

FLUX.1 [dev] Black Forest Labs Grok Imagine Image xAI

Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.

FLUX.1 [dev]

24.6 arena score

#16 of 62 in Text-to-Image

Skill signature · Text-to-Image

Grok Imagine Image

23.4 arena score

#26 of 62 in Text-to-Image

Vote tally

Where the votes landed

FLUX.1 [dev]

0%

win rate

Ties

0%

Grok Imagine Image

0%

win rate

Shared challenges 13

Challenge by challenge

The strongest take from each model on every shared challenge, with the AI judge's read.

Geometric Composition

Text-to-Image

“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”

FLUX.1 [dev]
Grok Imagine Image

AI Judge Analysis

FLUX.1 [dev]

  • + Excellent photorealism with shallow depth of field.
  • + Accurate spatial reasoning for object placement.
  • + High-quality textures, particularly on the red book cover and wooden table.
  • The sphere is quite large, filling much of the cube rather than being 'small'.
  • The plant's visibility through the cube is somewhat obscured by internal reflections.

Grok Imagine Image

  • + Better adherence to the 'small' scale of the blue sphere.
  • + The plant is clearly visible through the glass cube as requested.
  • + Realistic natural lighting on the wooden surface.
  • The sphere appears to be floating unnaturally without support or resting on the bottom.
  • The book is slightly misaligned with the top of the cube.

Verdict: Both models followed the complex spatial instructions well. FLUX.1 [dev] produced a more aesthetically pleasing and polished image with superior textures, though Grok Imagine Image followed the scale of the 'small sphere' more accurately and showcased the plant through the glass more clearly. FLUX.1 [dev] is the winner due to higher visual fidelity and more realistic object integration.

Candid Street Photography

Text-to-Image

“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”

FLUX.1 [dev]
Grok Imagine Image

AI Judge Analysis

FLUX.1 [dev]

  • + Excellent skin texture and facial detail
  • + Highly realistic cinematic lighting and rain effects
  • + Perfect adherence to the 50mm shallow depth of field request
  • The subject is holding the handlebars rather than performing a repair action
  • The background car looks slightly static despite the motion blur request

Grok Imagine Image

  • + Successfully captured the 'imperfect framing' and 'candid' feel
  • + Good implementation of motion blur on the passing vehicle
  • + The posture accurately suggests he is actually repairing the bike
  • The red bicycle has significant structural AI artifacts (broken frame near the seat)
  • The subject's face is obscured and lacks the 'natural skin texture' requested
  • Lower overall image resolution and clarity compared to the competitor

Verdict: FLUX.1 [dev] produced a stunning, high-fidelity image that perfectly matches the aesthetic tone and technical camera specs, though it missed the specific 'motion blur' on the car. Grok Imagine Image captured the candid 'street photography' vibe and motion blur better, but failed on technical execution with a broken bicycle geometry and lack of facial detail. FLUX.1 [dev] is the clear winner for its superior realism and adherence to visual quality prompts.

Fantasy Warrior

Text-to-Image

“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”

FLUX.1 [dev]
Grok Imagine Image

AI Judge Analysis

FLUX.1 [dev]

  • + Excellent photographic quality and skin texture rendering.
  • + Beautiful soft lighting and bokeh effect.
  • + Lifelike eyes with striking clarity and detail.
  • Failed to include the requested engraving on the plate armor.
  • Missed the 'battle-worn' requirement, as the subject looks pristine with freckles instead of scars.
  • Did not include beads in the hair braids.

Grok Imagine Image

  • + Successfully included all prompt elements including engraved armor, hair beads, and scars.
  • + Strong adherence to the 'battle-worn' aesthetic with dirt and visible facial marks.
  • + Detailed textures on leather straps and the cloth underlayer.
  • The torch light sources are a bit distracting and slightly bloom-heavy.
  • The hair braids merge somewhat awkwardly into the armor plates.

Verdict: While FLUX.1 [dev] produced a stunningly high-quality photographic portrait, it failed to follow almost all of the specific thematic details such as the engravings, hair beads, and scars. Grok Imagine followed the prompt with high precision, accurately depicting a battle-worn character with all the requested technical and aesthetic details, making it the superior choice for this specific challenge.

Modern Clean Menu

Text-to-Image

“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”

FLUX.1 [dev]
Grok Imagine Image

AI Judge Analysis

FLUX.1 [dev]

  • + Clean, professional aesthetic with excellent white space utilization.
  • + Very high-quality food photography with realistic lighting and shadows.
  • + Excellent adherence to a modern minimalist style suitable for a high-end casual dining spot.
  • Failed to include a dedicated 'Pizza' section heading as requested.
  • Food photos are clustered in corners rather than in a defined grid layout.
  • Large amount of illegible gibberish text in the menu body.

Grok Imagine Image

  • + Follows the grid layout prompt more accurately than Model A.
  • + Includes all requested sections: Appetizers, Pizza, and Mains with bold headers.
  • + Text rendering for headers is much clearer and legible.
  • Layout feels cluttered and lacks the 'minimalist' feel requested in the prompt.
  • Contains several duplicate menu items (e.g., multiple 'Grilled Salmon' and 'Steak Frites' entries).
  • Food images have a slightly artificial, overly processed appearance compared to Model A.

Verdict: Grok Imagine followed the functional requirements of the prompt more closely by including the specific 'Pizza' header and attempting a grid-based layout for the food. However, FLUX.1 [dev] produced a much more visually appealing and professional-looking design that better captured the 'modern minimalist' aesthetic, despite missing one header and ignoring the grid instruction. Grok's output suffered from repetitive entries and a cluttered composition.

Magic Burger Explosion: Fiery Photorealism Challenge

Text-to-Image

“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”

FLUX.1 [dev]
Grok Imagine Image

AI Judge Analysis

FLUX.1 [dev]

  • + Clean photorealistic textures on the burger components
  • + Solid rendering of the price and limited time text
  • Completely failed to include the primary title 'MAGIC BURGER' as requested
  • Composition is static and vertical rather than 'dynamic' and 'exploded'
  • Missing the starburst element for the price

Grok Imagine Image

  • + Successfully integrated all requested text elements with the specified fiery effect
  • + Dynamic composition with splashes of sauce and floating ingredients that conveys motion
  • + Perfectly followed the instruction for a price starburst and fiery background
  • Texture on the bun is slightly less realistic compared to Model A
  • The lettuce leaves look a bit repetitive in shape

Verdict: Grok Imagine Image followed the prompt requirements almost perfectly, including all requested text and the specific 'starburst' for the price, while maintaining a high-energy composition. FLUX.1 [dev] failed on several key instructions, most notably missing the main title 'MAGIC BURGER' and failing to create a dynamic 'exploded' layout.

Chalkboard Menu

Text-to-Image

“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”

FLUX.1 [dev]
Grok Imagine Image

AI Judge Analysis

FLUX.1 [dev]

  • + Excellent layout with centered text and clean framing
  • + Accurate spelling for all requested menu items
  • + High clarity and contrast on the text surface
  • The text looks more like a digital marker or vector font than authentic dusty chalk
  • Contains some spelling errors in the footer text like 'drish' and 'aur'

Grok Imagine Image

  • + Highly realistic chalk texture with dusty smudges and varying pressure
  • + Perfectly captures the 'handwritten' request with natural slants and irregular letter heights
  • + The lighting and environment feel more integrated and cozy
  • The text is slightly less legible due to the authentic chalk grain
  • The third menu item overflows slightly toward the right edge

Verdict: While FLUX.1 [dev] produced a very clean and legible board, Grok Imagine captured the 'handwritten chalk' aesthetic much more convincingly, including realistic dust smudges and textured strokes. Grok Imagine followed the styling nuances of the prompt more effectively, whereas FLUX.1 [dev] output text that felt slightly too much like a clean digital font.

The Reversed Rodeo

Text-to-Image

“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”

FLUX.1 [dev]
Grok Imagine Image

AI Judge Analysis

FLUX.1 [dev]

  • + Excellent anatomical details on the horse's musculature
  • + Clean, cinematic lighting with a realistic planet curvature
  • + Highly consistent rendering of the astronaut suit and saddle
  • Failed the specific prompt instruction to have the 'horse on top'
  • Interpreted 'horse riding astronaut' as a standard configuration

Grok Imagine Image

  • + Successfully followed the difficult 'horse on top' instruction
  • + Vibrant, creative use of nebula colors and composition
  • + Maintains high detail in the astronaut suit and horse hair
  • The connection between the horse and astronaut is a bit ambiguous physically
  • Small artifact in the horse's front left hoof

Verdict: While FLUX.1 [dev] produced a more polished and photorealistic image, it completely ignored the specific negative constraint that the horse should be on top of the astronaut. Grok Imagine Image interpreted the surreal prompt correctly, showing a horse positioned above the astronaut in a more creative and literal way, making it the winner for prompt adherence.

The Capybara Taxi Driver

Text-to-Image

“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”

FLUX.1 [dev]
Grok Imagine Image

AI Judge Analysis

FLUX.1 [dev]

  • + Excellent fur texture rendering on the capybara
  • + High photographic realism in lighting and bokeh depth of field
  • + Strong cinematic color grading
  • The passenger is sitting in the front seat instead of the back seat as requested
  • The woman's hand holding the phone looks slightly distorted with unnatural red finger tips
  • The capybara's paws are not placed logically on the steering wheel

Grok Imagine Image

  • + Perfect adherence to the instruction for the passenger to be in the back seat
  • + Logical and realistic hand/paw placement on the steering wheel and phone
  • + Great background detail of a Manhattan street at night
  • The capybara's face is slightly less detailed compared to Model A
  • The capybara's hat is a bit small and sits awkwardly on its head

Verdict: While FLUX.1 [dev] produced a more cinematically beautiful image with superior textures, it failed a major part of the prompt by placing the passenger in the front seat. Grok Imagine followed all spatial instructions accurately, correctly placing the businesswoman in the back seat and providing a more realistic compositions of the hands and paws.

The Halloween Invitation

Text-to-Image

“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”

FLUX.1 [dev]
Grok Imagine Image

AI Judge Analysis

FLUX.1 [dev]

  • + Features a bold, vibrant jack-o-lantern illustration.
  • + Includes a thorny decorative border.
  • Failed significantly on text rendering with multiple typos like 'Lalloween Rantcy'.
  • Ignored several prompt elements including the 'parchment' texture and 'webs'.
  • The layout is more of a generic poster than an invitation card.

Grok Imagine Image

  • + Excellent text rendering with no spelling errors across all required sections.
  • + Successfully incorporated the vintage parchment texture and spider webs.
  • + The gothic title and scroll banner are perfectly executed according to the prompt.
  • The bats are slightly repetitive in shape.
  • The jack-o-lantern is a bit small relative to the text.

Verdict: Grok Imagine Image followed the prompt instructions near-perfectly, including the specific textures, border elements, and complex text requirements without any spelling errors. In contrast, FLUX.1 [dev] failed to include the requested parchment/webs and struggled significantly with the text, resulting in unintelligible words like 'Lalloween Rantcy'.

Isometric Miniature Diorama Scenes

Text-to-Image

“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”

FLUX.1 [dev]
Grok Imagine Image

AI Judge Analysis

FLUX.1 [dev]

  • + Excellent soft lighting and refined textures
  • + Accurate 45-degree isometric perspective
  • + Includes the requested diorama base
  • Text rendering is poor with incorrect characters and a typo
  • Sushi model is repetitive with identical toppings

Grok Imagine Image

  • + Perfect text rendering for 'JAPAN' and 'SUSHI'
  • + Clean 3D cartoon art style with vibrant colors
  • + Good variety of sushi types following the miniature theme
  • Flag icon is placed at the top instead of top-center according to typography hierarchy
  • Perspective feels slightly more front-on than a true top-down 45-degree isometric view

Verdict: Grok Imagine Image significantly outperformed FLUX.1 [dev] by accurately rendering the requested text without typos or gibberish. While FLUX.1 [dev] produced a more sophisticated miniature material look, Grok Imagine Image's adherence to the text prompt and overall clarity make it the superior choice for this specific graphic design challenge.

Adorable Baby Animals in Sunny Meadow

Text-to-Image

“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”

FLUX.1 [dev]
Grok Imagine Image

AI Judge Analysis

FLUX.1 [dev]

  • + Excellent soft lighting consistent with a golden sunrise.
  • + Clean, 3D animated-style aesthetic with expressive eyes.
  • + Cohesive arrangement of characters.
  • Failed to include the tabby kitten; it generated two puppies and two fox/rabbit hybrids instead.
  • Style is more '3D render' than 'hyper-photorealistic'.
  • The animals appear static and posed rather than playful.

Grok Imagine Image

  • + Successfully included all four requested animals (dog, cat, rabbit, fox).
  • + Dynamic posing captures the 'tumbling' and 'chasing' aspect of the prompt.
  • + Included noticeable dew sparkles in the grass.
  • Fur textures have a swirling 'AI-generated' artifact pattern rather than realistic hair.
  • Anatomical issues where the cat and fox bodies seem to merge or overlap confusingly.
  • Strong over-sharpening artifacts in the background.

Verdict: Both models struggled with the 'hyper-photorealistic' requirement, instead opting for a stylized, Pixar-like aesthetic. Grok Imagine succeeded in including all four specific animals, whereas FLUX.1 [dev] failed to generate the tabby kitten, though FLUX.1 [dev] produced a much cleaner image with fewer technical artifacts.

Vintage Cafe Logo

Text-to-Image

“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”

FLUX.1 [dev]
Grok Imagine Image

AI Judge Analysis

FLUX.1 [dev]

  • + Strong vintage aesthetic with woodcut-style shading on the banner
  • + Balanced composition for a classic emblem
  • + Applies the subtle texture requested in the prompt
  • Failed the primary text spelling with 'Café Flariláan'
  • Added hallucinated numbers '11011' and '1941'
  • The cloche is very small and lacks detail

Grok Imagine Image

  • + Perfect adherence to the primary text spelling 'Caffè Florian'
  • + Strong, clean vector emblem style with a prominent cloche graphic
  • + Clear and legible typography
  • Repeated the 'Est. 1720' text twice, which was not requested
  • The spoon/handle sticking out of the cloche is slightly awkward

Verdict: While FLUX.1 [dev] captures the vintage texture and 'vibe' better, it fails significantly on the actual text content and adds random numerical noise. Grok Imagine achieves perfect spelling for the brand name and provides a much cleaner, more professional vector graphic, making it much more useful for a logo design task despite repeating the date.

Apollo 11: Journey to Tranquility

Text-to-Image

“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”

FLUX.1 [dev]
Grok Imagine Image

AI Judge Analysis

FLUX.1 [dev]

  • + Excellent aesthetic composition with a central focus and radial flow.
  • + Smooth color integration using the requested palette.
  • + High-quality vector styling with clean line work.
  • Text is almost entirely illegible gibberish.
  • Inaccurate iconography, such as including Ringed planets (Saturn-like) for Earth and Moon orbits.

Grok Imagine Image

  • + Highly legible and mostly accurate text labels for all steps.
  • + Strong adherence to the requested NASA color palette (navy, red, white).
  • + Accurate iconography including the Earth, Moon, and Saturn V rocket.
  • The layout is a bit cluttered compared to a professional infographic.
  • The Moon icon erroneously has rings which doesn't match the prompt's request for accuracy.

Verdict: While FLUX.1 [dev] produces a more artistic and visually cohesive layout, it fails completely on text legibility and orbital accuracy. Grok Imagine Image successfully illustrates all six requested steps with readable text and a clear, albeit busy, educational layout, making it the better choice for an infographic task.

Next steps

Explore each model