Black Forest Labs' 12 billion parameter distilled image generation model optimized for speed, capable of generating high-quality images in just 4 inference steps
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
FLUX.1 [schnell]
#48 of 62 in Text-to-Image
Grok Imagine Image
#26 of 62 in Text-to-Image
Where the votes landed
FLUX.1 [schnell]
0%
win rate
Ties
0%
Grok Imagine Image
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
FLUX.1 [schnell]
- + Exceptional sharpness and clarity in textures.
- + Accurate and pleasing soft window light from the left.
- + Vibrant color palette.
- − Included an extra blue sphere on top of the book that was not requested.
- − The sphere inside the cube is resting on nothing, looking somewhat unnatural.
Grok Imagine Image
- + Followed the spatial instructions perfectly without adding extra objects.
- + Realistic wooden table texture and natural plant placement behind the glass.
- + Accurately represents the 'small' scale of the sphere relative to the cube.
- − The cube geometry is slightly warped/asymmetrical.
- − The blue sphere has a somewhat speckled texture that looks slightly off.
Verdict: While FLUX.1 [schnell] has higher raw image quality and better lighting, it failed the precise prompt instructions by adding a second blue sphere on top of the book. Grok Imagine followed all instructions accurately and handled the 'partially visible through glass' requirement for the plant more convincingly, making it the winner despite slightly less refined rendering.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
FLUX.1 [schnell]
- + Excellent detail on the man's face and natural skin texture
- + Strong color contrast with the vibrant red bicycle and warm light reflections
- + Cinematic composition with good depth of field feel
- − Lack of required motion blur on the passing cars
- − The hands and bicycle handlebars have slight structural inconsistencies (fusion)
- − The scene feels a bit too posed and static for a 'candid' prompt
Grok Imagine Image
- + Highly accurate adherence to the 'motion blur' and 'candid' prompt requirements
- + Authentic 'imperfect framing' that mimics real street photography
- + Very convincing wet pavement reflections and light rain atmosphere
- − The subject's face is obscured and less detailed than the other model
- − The red of the bicycle is slightly more muted
- − Low detail on the hands
Verdict: While FLUX.1 [schnell] captures a more aesthetically pleasing portrait with high detail, Grok Imagine Image followed the technical prompt much more accurately. Grok effectively included the requested motion blur and imperfect framing, which provided a more authentic 'candid street photo' feel compared to the more staged appearance of the FLUX.1 output.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
FLUX.1 [schnell]
- + Extremely high skin texture detail and realistic iris patterns
- + Intense, characterful expression that conveys 'battle-worn' well
- + Excellent hair and beard realism
- − The lighting in the background feels like a flat orange wash rather than a distinct torch
- − Armor engraving is less distinct and more generic than the competitor
- − Missed the specific 'beads' in the hair braids
Grok Imagine Image
- + Superior adherence to all prompt details including beads, bokeh sparks, and engraved plate
- + Beautifully intricate engraving on the armor which feels more 'ornate'
- + Better environmental storytelling with visible torch sources and shallow depth of field
- − Skin texture is slightly smoother and less realistic than the competitor
- − Slightly less 'battle-worn' in the facial expression despite the scars
Verdict: While FLUX.1 [schnell] produces a more convincing human face with superior skin texture, Grok Imagine is the clear winner for prompt adherence. Grok Imagine successfully included the specific hair beads and bokeh sparks requested, while also providing much more detailed and ornate engraving on the paladin's armor.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
FLUX.1 [schnell]
- + Strong minimalist aesthetic with plenty of white space
- + Clean professional layout that feels high-end
- + Excellent sans-serif font choices
- − Failed to include a 'Mains' section, using 'Orfefus' instead
- − Photos are somewhat repetitive with too much focus on pizza/orange-colored items
- − Large amount of illegible placeholder text
Grok Imagine Image
- + Successfully included all requested categories: Appetizers, Pizza, and Mains
- + Graphic design is vibrant and engaging with varied food photography
- + Fonts are bold and easy to distinguish across sections
- − Large number of spelling errors and repetitive menu items
- − Layout feels a bit cluttered compared to the minimalist request
- − Food photos are sporadically placed rather than in a strict grid
Verdict: Grok Imagine followed the prompt instructions more closely by including all requested sections (Appetizers, Pizza, Mains), whereas FLUX.1 [schnell] failed on the content categories. While FLUX.1 [schnell] captured a cleaner minimalist aesthetic, Grok Imagine provided a more complete design that better reflects a functional casual dining menu.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
FLUX.1 [schnell]
- + Excellent photorealistic texture on the meat and buns
- + Great lighting integration with the bottom fire
- + The burger itself looks highly appetizing
- − Failed the main text significantly with 'AGIC BURGER' and a redundant '€699' price
- − The burger is largely assembled rather than 'exploded' as requested
- − Floating crouton-like debris seems unrelated to a burger
Grok Imagine Image
- + Perfectly followed the 'exploded' instruction with all components separated in air
- + Accurate and high-quality rendering of all requested text and prices
- + Very effective fiery, glowing text effect that matches the theme
- − The starburst design looks a bit like a clipart asset
- − The lettuce and sauce splashes look slightly more CGI than photorealistic
Verdict: Grok Imagine Image is the clear winner as it followed every instruction in the prompt, including the complex layout of the 'exploded' burger and the specific text strings. FLUX.1 [schnell] failed to spell the product name correctly and provided a mostly intact burger instead of the requested motion-heavy exploded view.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
FLUX.1 [schnell]
- + Features a clean, legible layout.
- + Maintains a consistent handwriting style throughout.
- − Numerous spelling errors including 'Pril', 'Taffle', 'Mushmnctiomn', and 'Octtoopus'.
- − Text feels more like a digital pen than authentic chalk on a board.
- − Missed the word 'Today's' (using 'Today' instead).
Grok Imagine Image
- + Excellent prompt adherence with nearly perfect spelling of complex menu items.
- + Highly realistic chalk texture with smudges and dust that enhance the 'cozy café' atmosphere.
- + Perfectly captures the 'elegant cursive' requested for the title.
- − Very minor crowding towards the right edge of the board.
- − The price for the cookies is slightly detached from the main text line.
Verdict: Grok Imagine Image significantly outperforms FLUX.1 [schnell] by accurately rendering every specific menu item requested with near-perfect spelling and a much more authentic chalk texture. While FLUX.1 [schnell] struggles with nonsensical words and a sterile digital appearance, Grok Imagine Image successfully captures the requested 'handwritten style' and 'cozy' atmosphere.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
FLUX.1 [schnell]
- + Excellent adherence to the 'horse riding astronaut' concept with the horse physically seated on the gear.
- + Cinematic lighting with a clear focal point and realistic atmospheric depth.
- + High-quality texture on the horse's fur and the astronaut's suit.
- − Anatomical anomaly with a second horse head emerging from the first.
- − Composition is slightly crowded toward the center.
Grok Imagine Image
- + Beautiful, vibrant nebular background and surreal cosmic atmosphere.
- + Physically distinct separation between the astronaut and the horse while maintaining the requested spatial relationship.
- + Great sense of movement and dynamic posture for both subjects.
- − The horse appears to be drifting above the astronaut rather than 'riding' him.
- − Minor anatomical issues with the astronaut's glove connecting to the horse's hoof.
Verdict: Both models struggled with the physics-defying prompt, but FLUX.1 [schnell] followed the 'riding' instruction more literally by placing the horse directly on the astronaut's equipment. However, FLUX.1 [schnell] suffered a major artifact with a dual-headed horse, whereas Grok Imagine produced a much cleaner, more aesthetically pleasing surreal composition despite the horse being more of a companion than a rider.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
FLUX.1 [schnell]
- + Excellent fur texture and lighting on the capybara.
- + Captures a very cinematic, close-up perspective from the passenger side.
- + Accurate text on the taxi hat.
- − The passenger appears to be in the front seat or mid-cabin due to the perspective.
- − Only one paw is clearly on the steering wheel.
- − The passenger's scale feels a bit small compared to the capybara.
Grok Imagine Image
- + Perfectly adheres to the layout with the passenger clearly in the back seat.
- + Accurately depicts both paws on the steering wheel as requested.
- + The 'bored' expression on the businesswoman is very well executed.
- − The capybara's head shape is slightly elongated and less realistic than in Model A.
- − The perspective through the front windshield makes the car look like it's missing a hood or engine area.
- − The lighting is a bit flat compared to the cinematic feel of Model A.
Verdict: Grok Imagine Image followed the prompt's spatial instructions much better, correctly placing the businesswoman in the back seat and showing both paws on the wheel. While FLUX.1 [schnell] produced a more aesthetically pleasing image with superior textures, it failed the specific layout requirement by placing the passenger directly next to the driver.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
FLUX.1 [schnell]
- + Excellent color contrast and vibrant glowing effects
- + Good layout with a distinct border
- + Accurate date and location text
- − Significant spelling errors and repetitive text lines in the main body
- − The scroll banner is split and contains gibberish
- − The background is more digital-illustration style than vintage parchment
Grok Imagine Image
- + Perfect text rendering for all requested details and banners
- + Authentic vintage parchment texture and gothic aesthetic
- + Superior detailed border with thorns and cobwebs that feels more 'cinematic'
- − The composition is a bit crowded horizontally
- − The jack-o-lantern light has slightly less depth than Model A
Verdict: Grok Imagine Image is the clear winner as it followed all textual instructions perfectly, including the specific phrasing for the banners and event details. FLUX.1 [schnell] struggled significantly with text coherence, resulting in numerous typos and redundant lines, whereas Grok captured the vintage gothic parchment aesthetic with much higher fidelity.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
FLUX.1 [schnell]
- + Excellent clean 3D isometric diorama presentation
- + Very soft and refined textures on the base and food
- + Effective use of negative space for a minimalist aesthetic
- − Failed to include the word 'SUSHI' under 'JAPAN'
- − The sushi roll looks a bit flattened and less 'cartoonish' than the prompt suggested
Grok Imagine Image
- + Perfect adherence to text instructions including 'JAPAN' and 'SUSHI'
- + Includes a diverse and appealing variety of sushi pieces
- + Better interpretation of the 'cartoon scene' style with vibrant colors
- − The grain count on the rice appears slightly repetitive and artificial
- − Shadows on the base are a bit more harsh compared to the 'gentle lighting' request
Verdict: Grok Imagine Image followed the prompt much more accurately by including all the requested text and providing a more representative variety of sushi. While FLUX.1 [schnell] has a very sophisticated, clean 3D render feel, its failure to generate the second line of text makes it less successful for this specific prompt.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
FLUX.1 [schnell]
- + Excellent fur texture and realistic animal features
- + Beautiful soft lighting and color grading
- + Natural composition of animals nestled in the grass
- − Failed to include the baby bunny
- − Animals appear static rather than chasing butterflies
Grok Imagine Image
- + Strong prompt adherence including all four animals
- + Excellent sense of motion and 'tumbling' action
- + Clearly visible dew sparkles and god rays
- − Stylized 'AI' appearance with overly smooth fur rendering
- − Anatomical oddities in the animal legs and paws
- − Butterfly elements from the prompt are missing
Verdict: Grok Imagine Image followed the complex prompt more closely by including the bunny and capturing the sense of play, though it has a clearly artificial, stylized look. FLUX.1 [schnell] produced a more photorealistic image with superior fur textures and lighting, but it missed one of the four requested animals and lacked the dynamic motion described in the prompt.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
FLUX.1 [schnell]
- + Strong minimalist vector aesthetic
- + Excellent composition with balanced circular framing
- − Significant spelling error in the primary brand name
- − Incorrect year (7720 instead of 1720)
- − Missing the requested steam element
Grok Imagine Image
- + Perfect text accuracy for both name and date
- + Includes all requested elements including steam and cloche
- + Successfully captures the vintage texture and color palette
- − Redundant inclusion of 'Est. 1720' occurring twice
- − The spoon/handle element behind the cloche is slightly ambiguous
Verdict: Grok Imagine Image is the clear winner as it adhered perfectly to the text requirements, whereas FLUX.1 [schnell] failed significantly on spelling and date accuracy. While FLUX.1 [schnell] offered a cleaner vector layout, Grok Imagine Image correctly included the requested steam and provided a professional vintage aesthetic.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
FLUX.1 [schnell]
- + Features a clean, minimalist layout that feels like a modern conceptual poster.
- + Adheres strictly to the requested NASA-inspired color palette.
- + Includes a nicely stylized lunar surface at the bottom.
- − The text is largely illegible gibberish.
- − The diagram is confusing and does not clearly follow the 6 requested steps.
- − The icons are abstract and don't clearly represent the Saturn V or the Lunar Module.
Grok Imagine Image
- + Excellent adherence to the 6-step structure with clear, relevant icons for each phase.
- + Text rendering is highly legible and includes accurate names like Armstrong, Aldrin, and Collins.
- + Captures the flat-vector style perfectly with consistent iconography.
- − Includes the literal phrase 'NASA inspired' within the artwork.
- − Minor spelling errors in smaller text labels like '3rajoory' and 'Moom'.
Verdict: Grok Imagine Image is the clear winner as it successfully followed all instructional steps, providing a logical flow of information with specific, recognizable icons for each phase of the mission. While FLUX.1 [schnell] captures a nice aesthetic, it fails as an infographic because the diagram is nonsensical and the text is unreadable. Grok Imagine Image's ability to render legible and accurate names (Armstrong, Aldrin, Collins) adds significant value to the educational nature of the prompt.
Explore each model
An image generation model by xAI designed to generate highly aesthetic images from text descriptions.