Black Forest Labs' 12-billion parameter flow transformer for high-quality text-to-image generation, suitable for personal and commercial use with streaming support
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
FLUX.1 [dev]
#16 of 62 in Text-to-Image
Grok Imagine Image Pro
#17 of 62 in Text-to-Image
Where the votes landed
FLUX.1 [dev]
0%
win rate
Ties
0%
Grok Imagine Image Pro
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
FLUX.1 [dev]
- + Excellent photorealistic lighting and shallow depth of field.
- + Complex glass refractions and reflections look highly realistic.
- + Follows the relative positions of all objects accurately.
- − The sphere appears to be floating inside the cube rather than resting on the bottom.
- − The plant is very out of focus and less visible through the glass.
Grok Imagine Image Pro
- + Perfect adherence to the prompt regarding the plant being 'partially visible through the glass'.
- + Excellent text rendering on the book spine.
- + Realistic textures on the wooden table and clay pot.
- − The glass cube has some geometric inconsistencies, especially where the edges meet.
- − The reflection of the plant pot on the right is physically impossible/misplaced.
Verdict: Both models followed the prompt closely, but Grok Imagine Image Pro wins for its superior adherence to the 'plant behind the cube visible through the glass' instruction, which FLUX.1 [dev] obscured with heavy blur. While FLUX.1 [dev] has more realistic lighting and refractions, Grok Imagine Image Pro provides a more detailed scene with impressive text rendering on the book.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
FLUX.1 [dev]
- + Excellent shallow depth of field and bokeh
- + Subtle, realistic rain effects and reflections
- + High quality skin textures and clothing details
- − The man is just holding/looking at the bike rather than making a clear repair
- − Lacks the motion blur from cars requested in the prompt
Grok Imagine Image Pro
- + Better logic for the 'repairing' action with the crouched pose and wrench
- + Includes motion blur on the background vehicles as requested
- + Stronger 'candid' feel and imperfect framing as per the prompt
- − Distorted hand and wrench (clipping through the fingers and bike frame)
- − The 50mm shallow depth of field is less pronounced than in Image A
Verdict: Grok Imagine Image Pro followed the prompt instructions more closely by including motion blur and a clear repair action, though it suffered from anatomical issues in the hands. FLUX.1 [dev] produced a more aesthetically pleasing and high-fidelity image with superior skin textures, but it ignored the motion blur and the subject's interaction with the bicycle is less specific to 'repairing'.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
FLUX.1 [dev]
- + Excellent shallow depth of field and bokeh effect
- + High-quality skin texture and lifelike eyes
- − Failed to include 'small beads' in the hair braids
- − The plate armor lacks the 'ornate engraved' detail requested
- − Character looks too clean and youthful, missing the 'battle-worn' aesthetic
Grok Imagine Image Pro
- + Perfect adherence to specific details like hair beads, engravings, and scars
- + Superior detail on leather straps, armor texture, and cloth underlayers
- + Includes legible Latin-inspired text on the gorget which adds to the paladin theme
- − The bokeh sparks appear a bit more synthetic/repetitive than Model A
- − Facial structure and hair roots near the beads are slightly less naturalistic than Model A
Verdict: Grok Imagine Image Pro is the clear winner as it followed every specific instruction in the prompt, including the beads in the hair, the ornate engravings, and the battle-worn texture. FLUX.1 [dev] produced a high-quality portrait but missed several key descriptive elements, resulting in a character that looks like a clean model rather than a grizzled paladin.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
FLUX.1 [dev]
- + Excellent typography rendering with very clean, professional sans-serif fonts
- + Sophisticated layout that uses white space effectively in a minimalist style
- + Realistic food photography that blends seamlessly with the graphic elements
- − Does not feature the requested 'food photos in grid' layout
- − Missing the specific 'Pizza' section requested in the prompt
Grok Imagine Image Pro
- + Perfectly adheres to the 'grid' request for food photos
- + Includes all requested sections: Appetizers, Pizza, and Mains
- + Vibrant and appetizing food imagery
- − Font choices are slightly less professional and 'bubbly' compared to a modern minimalist aesthetic
- − Some text artifacts and spelling errors present in the descriptions
Verdict: Grok Imagine Image Pro followed the structural requirements of the prompt much more accurately by providing the grid layout and all three specific food sections. While FLUX.1 [dev] produced a more sophisticated and aesthetically pleasing design, it ignored the grid requirement and the 'pizza' section, making Grok the more successful model for this specific task.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
FLUX.1 [dev]
- + Excellent photorealistic texture on the burger patties and buns
- + Clean, minimalist composition for a high-end ad feel
- + Accurate rendering of the Euro symbol and requested price
- − Failed to include the main title 'MAGIC BURGER'
- − Missing the starburst element for the price
- − Lacks the sense of motion and 'exploded' energy requested in the prompt
Grok Imagine Image Pro
- + Strong adherence to all text requirements including 'MAGIC BURGER'
- + Dynamic, high-energy composition with great motion and 'exploded' effect
- + Effectively captures the fiery background and glowing starburst mentioned in the prompt
- − The cheese has somewhat unrealistic, stringy physics
- − The background is slightly more cluttered compared to the clean look of the competitor
Verdict: Grok Imagine Image Pro is the clear winner as it followed every instruction in the prompt, including complex text integration and specific design elements like the starburst. FLUX.1 [dev] produced a high-quality image but failed to include the primary brand name and lacked the dynamic energy requested.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
FLUX.1 [dev]
- + Excellent layout with clear spacing between menu items.
- + Consistent chalk stroke thickness throughout the board.
- + Successful rendering of the specific date and all requested prices.
- − The font looks a bit too clean and digital, lacking the requested messy chalk texture.
- − Lowercase 'u' and 'n' in the bottom text are poorly formed, making it look like gibberish in places.
Grok Imagine Image Pro
- + Superb authentic chalk texture with dusting and smudges on the board background.
- + Highly realistic handwriting that varies naturally in slant and size as requested.
- + The elegant cursive title perfectly matches the prompt's stylistic requirement.
- − The layout is slightly more cramped compared to Model A.
- − Includes a minor artifact at the very top of the frame where a light fixture hangs.
Verdict: Grok Imagine Image Pro significantly outperforms FLUX.1 [dev] by capturing the authentic feel of a chalk menu, including smudges, dust, and genuine handwriting texture. FLUX.1 [dev] produced a very clean image, but its text looks more like a digital font than actual chalk on a board.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
FLUX.1 [dev]
- + Excellent anatomical details on the astronaut suit.
- + Lighting and shadows are consistent with the celestial background.
- + Clean, high-resolution textures.
- − Completely failed the semantic prompt to have the horse on top of the astronaut.
Grok Imagine Image Pro
- + Successfully followed the difficult spatial prompt of placing the horse on top.
- + Vibrant, creative use of color and nebula effects.
- + Dynamic composition with detailed background elements like the ringed planet.
- − The horse's legs are slightly distorted where they meet the astronaut.
- − Lower realism in the suit textures compared to the competitor.
Verdict: While FLUX.1 [dev] produced a more polished and realistic single-subject render, it failed the logical constraint of the prompt. Grok Imagine Image Pro successfully interpreted the surreal 'horse on top' instruction, making it the winner for prompt adherence despite minor anatomical artifacts.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
FLUX.1 [dev]
- + Excellent photographic lighting and depth of field.
- + Accurately places the passenger in the back seat.
- + High-quality fur texture and realistic capybara facial anatomy.
- − The passenger is holding an incredibly large, warped phone.
- − The composition feels a bit tight, losing some of the Manhattan street context.
Grok Imagine Image Pro
- + Excellent text rendering on the hat with specific regional details.
- + Captures the bored, professional expression of the passenger perfectly.
- + Realistic taxi interior materials and textures.
- − Placed the passenger in the front seat instead of the back seat as requested.
- − The capybara paws look like strange claws grasping the wheel.
- − Minor anatomy issues on the passenger's fingers.
Verdict: FLUX.1 [dev] followed the spatial instructions much better by correctly placing the passenger in the back seat, whereas Grok Imagine moved her to the front. While Grok Imagine had superior text rendering on the hat and great internal details, FLUX.1 [dev] produced a more coherent scene that adhered to the complexity of the prompt's layout.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
FLUX.1 [dev]
- + Strong illustrative art style
- + Good contrast and cinematic lighting
- + Creative thorn-based frame
- − Significant text errors including 'Falloween Ranty' and 'You Tre'
- − Failed to include cobwebs in the border
- − Date formatting and repeated '7pm' are incorrect
Grok Imagine Image Pro
- + Near-perfect text accuracy for all requested fields
- + Closer adherence to the 'vintage dark parchment' aesthetic
- + Includes all requested elements like cobwebs, thorns, and scrolls
- − The 'Arches' text is slightly decorative but still legible
- − Illustration is slightly more generic than the artistic style of image A
Verdict: Grok Imagine Image Pro is the clear winner as it successfully rendered all the requested text with high accuracy, whereas FLUX.1 [dev] suffered from multiple spelling and formatting errors. Grok also followed the stylistic requirements more closely, incorporating both the parchment texture and the specific cobweb/thorn border elements requested in the prompt.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
FLUX.1 [dev]
- + Excellent soft lighting and sophisticated PBR textures.
- + Great miniature aesthetic with a clean professional diorama base.
- − Text rendering is hallucinated and misspelled ('SUSH CATON').
- − The sushi shapes are a bit repetitive and look slightly more like clay than distinct fish.
Grok Imagine Image Pro
- + Perfect text rendering of 'JAPAN' and 'SUSHI' with the correct flag icon.
- + Higher variety in sushi types (nigiri and maki) with very clean 'cartoon' 3D textures.
- + Precise adherence to the 45-degree isometric camera angle.
- − The diorama base is a simple wooden circle rather than a more architectural miniature base.
- − The lighting is a bit flatter compared to the soft shadows in Model A.
Verdict: Grok Imagine Image Pro is the clear winner as it perfectly followed the complex text instructions and accurately rendered the Japanese flag icon, whereas FLUX.1 [dev] failed significantly on text spelling. Grok also provided a better variety of sushi models that fit the '3D cartoon' prompt more effectively, even if FLUX had slightly more naturalistic lighting.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
FLUX.1 [dev]
- + Warm, pleasing color palette with soft golden light
- + Clear focus on subjects with a shallow depth of field
- − Failed the realism prompt, producing a stylized or Pixar-like 3D render
- − Missing several requested animals including the tabby kitten and distinct fox kit
- − Subjects look more like toys than living animals
Grok Imagine Image Pro
- + Successfully followed the hyper-photorealistic request
- + Included all requested animals: golden retriever, tabby kittens, bunny, and fox kit
- + Excellent rendering of light rays and dew sparkles on the wildflowers
- − The fox's anatomy is slightly awkward in its tumbling pose
- − Some minor blending issues where the animals meet the grass
Verdict: Grok Imagine Image Pro closely adhered to the prompt by providing a photorealistic scene with all requested animal species, whereas FLUX.1 [dev] produced a stylized, cartoon-like image that missed half of the specified animals. Grok's composition also captured the 'tumbling' and 'god rays' elements much more effectively than FLUX.1 [dev].
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
FLUX.1 [dev]
- + Strong vintage aesthetic with woodblock-style textures.
- + Includes all requested elements like the cloche and banner.
- − Major spelling errors including 'Flarilaan' and 'RESEAURANT'.
- − Includes extraneous numbers (1101, 1941) not requested in the prompt.
Grok Imagine Image Pro
- + Perfect text rendering with correct spelling and accent marks.
- + Clean vector emblem style that adheres well to the minimalist requirement.
- + Accurate interpretation of 'Est. 1720' on the banner.
- − The steam effect is a bit simple compared to the resto of the logo.
- − Composition is safe but slightly generic.
Verdict: While FLUX.1 [dev] captures a more authentic vintage texture, it fails significantly on text legibility and spelling, producing a 'hallucinated' name. Grok Imagine Image Pro followed the prompt instructions precisely, delivering a clean, professional logo with perfect typography and spelling.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
FLUX.1 [dev]
- + Excellent aesthetic adherence to the NASA-inspired soft palette.
- + Dynamic composition with a visually interesting central rocket and orbiting elements.
- + Crisp vector styling and sophisticated layout.
- − Text consists of nonsensical gibberish or hallucinated words.
- − Fails to follow the chronological sequence of the 6 steps requested.
- − Includes irrelevant planets with rings that don't belong in the Apollo mission story.
Grok Imagine Image Pro
- + Perfect adherence to the 6-step chronological sequence requested.
- + High quality text rendering with legible and accurate names for mission phases and crew.
- + Logical and clear iconography for each step of the mission.
- − The composition is a bit rigid and linear compared to Model A.
- − The 'Translunar' arc icon is slightly detached from the main vertical flow.
- − Slightly less creative use of the requested color palette across the background.
Verdict: Grok Imagine Image Pro is the clear winner because it correctly implemented the technical requirements of the infographic, including all six specific steps in order and rendering legible, accurate text. While FLUX.1 [dev] produced a more artistically pleasing 'NASA-style' poster, it failed entirely on the informational content, providing random orbiting planets and garbled text.
Explore each model
xAI's premium image generation model offering higher fidelity output and stronger performance on single-image editing benchmarks compared to the standard Grok Imagine model