Black Forest Labs' enhanced 12-billion parameter flow transformer with 6x faster generation than FLUX.1 [pro], delivering superior composition, detail, and artistic fidelity
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
FLUX1.1 [pro]
#50 of 62 in Text-to-Image
Grok Imagine Image
#26 of 62 in Text-to-Image
Where the votes landed
FLUX1.1 [pro]
0%
win rate
Ties
0%
Grok Imagine Image
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
FLUX1.1 [pro]
- + Excellent photorealistic texture on the book cover and glass edges.
- + Very clean reflections on the sphere showing window light.
- + Strong composition with vibrant, clear colors.
- − The 'cube' is significantly taller than it is wide, appearing more like a rectangular prism.
- − The sphere is quite large relative to the container.
Grok Imagine Image
- + Better adherence to the 'small blue sphere' description.
- + More realistic wooden table texture with visible grain and wear.
- + The container shape is closer to a cube than Model A.
- − The red book looks slightly flat and less detailed than in Model A.
- − Lighting is a bit more muted and less dynamic.
Verdict: Both models followed the prompt instructions accurately, but FLUX1.1 [pro] produced a more polished, high-contrast image with superior texture rendering on the glass and book. Grok Imagine Image captured the relative scale of the 'small' sphere and the 'cube' shape better, but lacked the crispness and vibrant lighting found in the FLUX1.1 output.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
FLUX1.1 [pro]
- + Excellent handling of wet pavement reflections and bokeh.
- + Strong atmospheric lighting and cinematic mood.
- + High level of detail in skin texture and clothing.
- − The bike anatomy is physically impossible with parts missing and floating.
- − The cars in the background lack the requested motion blur, appearing mostly static.
- − The subject is not actively 'repairing' the bike, just holding the handlebars.
Grok Imagine Image
- + Successfully captured the motion blur of passing cars as requested.
- + The 'imperfect framing' feels much more authentic to a candid street photograph.
- + The bike's geometry and the man's posture are more realistic for a repair task.
- − The subject's face is obscured, reducing the impact of the 'natural skin texture' prompt.
- − Overall resolution and sharpness are slightly lower than the competitor.
Verdict: FLUX1.1 [pro] creates a more visually stunning and high-resolution image, but fails significantly on the physical structure of the bicycle and the specific 'motion blur' request. Grok Imagine better captures the technical requirements of the prompt, including the motion blur and a more realistic repair scene, creating a more convincing 'candid' aesthetic despite lower fine detail.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
FLUX1.1 [pro]
- + Extremely lifelike eyes with realistic wetness and skin texture details
- + Strong cinematic lighting with convincing bokeh sparks
- + Excellent grit and dirt realism on the skin
- − The braiding with beads is very subtle and mostly obscured by messy hair
- − The engraving on the armor is less ornate than the other model
Grok Imagine Image
- + Incredibly detailed and beautiful engraving on the plate armor
- + Clear and explicit adherence to the braiding and beads requirement
- + Excellent texture on the cloth underlayer and leather straps
- − The facial features and hair look slightly more 'airbrushed' and less gritty than Model A
- − The torchlight is very bright and creates slightly distracting highlights on the face
Verdict: While FLUX1.1 [pro] excels at sheer hyper-realism and gritty character details, Grok Imagine provided a better adherence to the specific stylistic elements of the prompt, such as the ornate engravings and visible hair beads. Grok Imagine is the preferred choice here for its superior execution of the armor's complex textures and the overall composition.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
FLUX1.1 [pro]
- + Excellent presentation for a menu mockup with realistic 3D depth and silverware
- + High-quality, vibrant food photography within the grid sections
- + Clean and professional layout that feels like a real digital or printed graphic
- − Text consists largely of illegible gibberish characters
- − Design is more of a pamphlet/flyer than a standard single-page menu sheet
Grok Imagine Image
- + Strong typography with highly legible section headers and item names
- + Accurately includes all requested sections (Appetizers, Pizza, Mains)
- + Well-balanced grid layout with high-quality isolated food images
- − Repeated menu items (e.g., 'Grilled Salmon' and 'Steak Frites' appear multiple times)
- − Some anatomical errors in the food photography, specifically the fish and pasta details
Verdict: Grok Imagine Image provides a much more functional menu design with surprisingly legible bold sans-serif fonts and clear categorized sections despite repeating some menu items. FLUX1.1 [pro] creates a more visually appealing mock-up for a brand presentation with better lighting and photography, but the text is completely unreadable and fails to deliver the requested grid organization as effectively.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
FLUX1.1 [pro]
- + High level of photorealism in the food textures
- + Sophisticated depth of field and lighting integration
- + Clean text rendering for price and 'Limited Time Only'
- − Failed the 'exploded' instruction as the burger is largely intact
- − Missing the starburst element for the price
- − Text has repetitive titles and includes 'Magic Burger' twice in different styles
Grok Imagine Image
- + Followed all layout instructions including the exploded view and the starburst
- + Excellent adherence to the 'fiery, glowing effect' for all text elements
- + Dynamic composition with splashes of sauce and flying ingredients
- − Lighting on the burger feels slightly disconnected from the intense fire in the background
- − Some texture details on the meat patty are less realistic than the competitor
Verdict: Grok Imagine is the clear winner as it followed every specific detail of the prompt, including the exploded burger layout and the starburst price tag, which FLUX1.1 [pro] missed. While FLUX1.1 [pro] produced a more photorealistic burger, it failed to execute the core 'exploded' concept and struggled with the layout of the required text elements.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
FLUX1.1 [pro]
- + Excellent background blur and atmospheric lighting.
- + Clean, legible font style.
- − Significant text errors including 'Chipkies' and hallucinated prices like '$228'.
- − Failed to keep the text in the same style as requested, often looking more digital than chalk-drawn.
- − Split one menu item into two lines with repetitive text.
Grok Imagine Image
- + Perfect adherence to text content with zero spelling or pricing errors.
- + Masterful chalk texture with authentic smudges and varied letter weights.
- + Captured the specific 'elegant cursive' requirement for the title perfectly.
- − The frame of the chalkboard is slightly less sharp than the text itself.
- − The lighting is a bit dim overall, though fitting for a 'cozy cafe'.
Verdict: Grok Imagine is the clear winner as it followed every instruction, including several complex spelling and pricing requests, with absolute precision. FLUX1.1 [pro] struggled with the text generation, producing a hallucinated '$228' price tag and misspelling 'Cookies', whereas Grok Imagine produced a highly realistic chalk texture and perfectly accurate content.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
FLUX1.1 [pro]
- + Excellent cinematic lighting and texture on the horse's coat.
- + Strong composition with a sense of motion and scale.
- + Technically high-quality rendering of the spacesuit and landscape.
- − Completely failed the conceptual prompt requirement for the horse to be 'on top'.
- − Result is a cliché trope rather than the requested surreal subversion.
Grok Imagine Image
- + Successfully interpreted the surreal requirement of the horse being 'on top'.
- + Vibrant and colorful nebula background adds to the cosmic feel.
- + Good anatomical rendering of both the horse and the astronaut's gear.
- − The connection between the horse and astronaut is a bit physically awkward.
- − Slightly less 'cinematic' in lighting compared to the competitor.
Verdict: While FLUX1.1 [pro] produced a more visually polished and realistic image, it failed the specific instruction to have the 'horse on top'. Grok Imagine successfully followed the prompt's paradoxical logic to create a surreal scene where the horse is positioned above the astronaut, making it the better response to the specific request.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
FLUX1.1 [pro]
- + Excellent photographic lighting and skin texture
- + Stronger cinematic bokeh and atmosphere
- + Creative choice of profile-view composition
- − The capybara is not holding the steering wheel
- − The passenger appears to be in the front seat or right next to the driver instead of the back seat
- − Perspective is slightly confusing regarding the car's interior layout
Grok Imagine Image
- + Perfectly adheres to the 'back seat' positioning for the human
- + Very accurate depiction of hands on the steering wheel
- + Includes realistic NYC taxi details like the fare stickers and roof light on the HUD
- − Character expression is slightly more sad than bored
- − The steering wheel is positioned on the wrong side for a New York City vehicle
- − Texture on the capybara's claws looks a bit sharp/artificial
Verdict: Grok Imagine Image followed the technical instructions of the prompt much more accurately, correctly placing the passenger in the back seat and the capybara's paws on the wheel. While FLUX1.1 [pro] produced a more aesthetically pleasing, cinematic image with superior lighting, it failed to execute the spatial requirements of the scene. Grok Imagine Image is the winner for narrative accuracy despite the car being right-hand drive.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
FLUX1.1 [pro]
- + Excellent atmospheric lighting and moonlit backdrop
- + The frame design of thorns is intricate and elegant
- + High quality illustration of the central jack-o-lantern
- − Text rendering is poor with several typos and repeated information
- − Does not follow the specific 'Halloween Party Invitation' title request correctly in the banner
- − Background is more black than 'dark parchment' texture
Grok Imagine Image
- + Perfect text rendering for all requested fields with zero typos
- + Follows the layout instructions specifically, including the scroll banner location
- + Great interpretation of the dark parchment texture and spiderweb/thorn border
- − The bats have slightly more cartoonish anatomical features compared to Model A
- − The lighting on the pumpkin is a bit less cinematic than Model A
Verdict: Model B (Grok Imagine Image) is much better suited for an invitation task as it rendered all text perfectly without typos, whereas FLUX1.1 [pro] struggled with spelling and repeated lines. Model B also captured the vintage parchment aesthetic and the specific border elements more effectively while maintaining a clean, legible layout.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
FLUX1.1 [pro]
- + Excellent typography with a playful, rounded font that matches the cartoon style
- + Follows the request for a raised diorama base perfectly
- + Wide variety of sushi types that showcase the 'soft refined textures' prompt
- − Missed the small flag icon requirement
- − The garnish includes small tomatoes which is unconventional for traditional sushi
Grok Imagine Image
- + Includes all text and the flag icon as requested
- + Rendered more realistic 'PBR' materials, particularly the wood grain and lighting
- + Clearer rice grain detailing on the nigiri sushi
- − The 'JAPAN' text is simple white and lacks the 'large bold' presence of Model A
- − The composition of the sushi feels slightly crowded compared to Model A
Verdict: Both models followed the prompt well, but Grok Imagine (Model B) is the likely winner because it included all requested elements, specifically the flag icon which FLUX1.1 (Model A) missed. While Model A had more creative typography, Model B's adherence to the PBR material request and text layout was more complete.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
FLUX1.1 [pro]
- + Excellent handling of rim lighting and backlighting on the fur
- + High degree of realism in the fur textures and eye reflections
- + Beautifully soft, dreamy bokeh and color palette
- − Failed to include the red fox kit from the prompt
- − Animals are sitting still rather than 'playfully chasing' or 'tumbling'
Grok Imagine Image
- + Successfully included all four animals: golden retriever, kitten, bunny, and fox
- + Better captures the action of 'chasing and tumbling' described in the prompt
- + Captures the 'sunrise with god rays' and 'dew sparkles' more distinctly
- − Distinctly more 'AI-art' style rather than 'hyper-photorealistic'
- − Fur textures look somewhat stylized or painted rather than real fur fibers
- − Anatomy on the puppy's tail and fox's paws is slightly confused
Verdict: FLUX1.1 [pro] produced a much more realistic and aesthetically pleasing image with superior lighting, but failed the prompt's count requirement by omitting the fox and failed to depict the requested action. Grok Imagine followed the complex prompt more literally by including all four animals and the sense of movement, even though the final image has a more generic digital-art look. Grok Imagine is the winner for better prompt adherence regarding content and action.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
FLUX1.1 [pro]
- + Excellent vintage illustrative style with detailed engraving textures
- + High-quality decorative border adds to the restaurant's legacy feel
- + Perfectly captures the 'warm brown and cream' aesthetic requested
- − Serious spelling errors: 'Fratilian' instead of 'Florian'
- − Includes a redundant 'Cafe' text that wasn't explicitly requested
Grok Imagine Image
- + Perfect text rendering for 'Caffè Florian'
- + Clean, minimalist vector design that aligns with modern branding standards
- + Strong steam visualization above the cloche
- − Redundant 'Est. 1720' text appearing twice on the badge and below the name
- − The graphic elements like the handle behind the dome are slightly ambiguous
Verdict: Grok Imagine is the winner because it successfully spelled the requested name 'Caffè Florian', whereas FLUX1.1 [pro] significantly altered the name to 'Fratilian'. While FLUX1.1 [pro] captured the vintage texture and artistic complexity much better, the text accuracy makes Grok Imagine more useful as a logo mockup.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
FLUX1.1 [pro]
- + Excellent adherence to the 'subtle' and 'clean' aesthetic requested.
- + Strong color palette control matching the NASA-inspired theme.
- + Superior graphic design feel with high-quality vector-style art.
- − Text rendering is mostly gibberish or contains heavy spelling errors.
- − The numbered steps are disorganized and do not follow the sequence 1-6 correctly.
- − The infographic layout is confusing and lacks a clear logical flow.
Grok Imagine Image
- + Perfect adherence to the 6-step structure requested in the prompt.
- + Text is highly legible and remarkably accurate for AI, including names like Armstrong and Aldrin.
- + Excellent iconographic consistency and clear, professional layout.
- − The 'Translunar' icon includes redundant and misspelled text labels underneath.
- − The inclusion of text like 'NASA inspired' directly in the image is a bit literal.
- − A few small artifacts in the vector lines of the orbit rings.
Verdict: While FLUX1.1 [pro] captures a more sophisticated artistic 'modern vector' style, it fails significantly on the structural and textual requirements of the infographic. Grok Imagine followed the specific 6-step instructions perfectly, maintained a high level of text legibility, and produced a functional, well-organized educational poster.
Explore each model
An image generation model by xAI designed to generate highly aesthetic images from text descriptions.