FP8 quantized variant of Black Forest Labs' FLUX.1 [schnell] model, offering ~2x faster inference with reduced precision while maintaining high-quality image generation in 4 steps
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
FLUX.1 [schnell] FP8
#47 of 62 in Text-to-Image
Grok Imagine Image
#27 of 62 in Text-to-Image
Where the votes landed
FLUX.1 [schnell] FP8
0.0%
win rate
Ties
0.0%
Grok Imagine Image
100.0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
FLUX.1 [schnell] FP8
- + Excellent refraction and light play within the glass cube.
- + Modern, clean aesthetic with vibrant colors.
- + High visual clarity and smooth gradients.
- − The glass object is a rectangular prism rather than a cube.
- − The blue sphere is resting on a reflected plane that breaks the 'inside' logic slightly.
Grok Imagine Image
- + Perfect adherence to the object shapes, specifically the cube.
- + Renders the plant realistically visible through the glass as requested.
- + Natural wood texture and realistic soft lighting.
- − The blue sphere appears to be floating unnaturally without support or context.
- − The composition is a bit more muted and less 'premium' looking than the other model.
Verdict: Both models followed the complex spatial instructions well. FLUX.1 [schnell] FP8 produced a more visually striking and polished image with superior light handling, but it failed to create a true cube. Grok Imagine Image followed the 'cube' instruction perfectly and handled the 'plant behind the glass' prompt more effectively, despite the floating sphere looking somewhat magical.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
FLUX.1 [schnell] FP8
- + Excellent shallow depth of field and bokeh quality
- + High level of detail in skin texture and clothing
- + Strong cinematic lighting and reflections on the pavement
- − Static vehicles fail to show the requested motion blur
- − The bicycle geometry is slightly warped near the handlebars
- − Man appears somewhat younger than 'elderly'
Grok Imagine Image
- + Successfully captures motion blur of passing cars
- + More convincing 'imperfect' candid framing and pose
- + Authentic urban Japanese setting with realistic bicycle details
- − Face is obscured and less detailed than Image A
- − Lower resolution/clarity in fine textures
- − Foreground ground reflections are less pronounced than requested
Verdict: FLUX.1 [schnell] FP8 produces a more polished and high-detail cinematic portrait, but it fails to incorporate the specific 'motion blur' requested in the prompt. Grok Imagine Image much better captures the 'candid street photo' aesthetic with authentic motion blur and imperfect framing, making it feel more like a real photograph despite lower facial detail.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
FLUX.1 [schnell] FP8
- + Excellent intensity in the eyes and facial expression
- + Stronger, more dramatic torchlight highlights on the face
- + Higher local contrast and sharpness on skin textures
- − Missed the request for hair braids and beads
- − Armor is largely cut off at the bottom, losing the ornate details requested
- − Some strange artifacts/blurring on the lower right hair strand
Grok Imagine Image
- + Perfect adherence to specific details like braided hair with beads and ornate silver engraving
- + Excellent skin texture featuring both dirt and visible scars
- + Includes bokeh sparks and superior composition showing the leather straps and cloth underlayer
- − Lighting on the face is slightly flat compared to the backlighting
- − The hair on the left side of her face has some slightly messy blending with the armor
Verdict: Grok Imagine Image followed the prompt much more accurately, including specific details like the braided hair with beads, scars, and highly ornate silver armor that FLUX.1 [schnell] FP8 largely ignored or cut out of frame. While FLUX.1 [schnell] FP8 has a more intense and painterly look, Grok Imagine Image provided a more complete and technically accurate representation of the requested character elements.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
FLUX.1 [schnell] FP8
- + Clean, symmetrical layout that mimics a real bi-fold menu.
- + Consistent photographic style for food items.
- + Includes pricing alignment typical of professional menus.
- − Nonsensical text and spelling errors in section headers (e.g., 'ORCETERS', 'SECCER').
- − The food images are repetitive and lack variety in texture.
Grok Imagine Image
- + Legible and accurate English headers like 'APPETIZERS' and 'PIZZA'.
- + Highly vibrant and diverse food photography that fits the 'casual dining' prompt.
- + Creative use of whitespace and accents which makes the layout feel more dynamic.
- − Repetitive menu items (e.g., 'Grilled Salmon' and 'Steak Frites' listed multiple times).
- − Single-page flyer style rather than a traditional booklet layout shown in the other model.
Verdict: Grok Imagine Image is the superior choice because it generates legible English text for headers and menu items, which is critical for a design task. While FLUX.1 [schnell] FP8 has a clean bi-fold structure, its failure to spell basic categories like 'Appetizers' or 'Pizza' correctly makes it less functional as a design template.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
FLUX.1 [schnell] FP8
- + Clean layout for the central burger
- + High-quality photographic texture on the patty and bun
- − Significant spelling errors in the marketing text
- − Failed to apply the requested fiery, glowing effect to the typography
Grok Imagine Image
- + Perfectly executed 'exploded' view with dynamic ingredient suspension
- + Highly accurate text rendering with the requested fiery glowing effect
- − Burger ingredients look slightly more illustrative than photorealistic
- − A few minor digital artifacts in the sauce splashes
Verdict: Grok Imagine Image successfully follows all prompt instructions, including the complex text effects and the 'exploded' burger layout. FLUX.1 [schnell] FP8 fails on text accuracy and the specific 'exploded' requirement, keeping the burger largely assembled.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
FLUX.1 [schnell] FP8
- + Successfully renders the specific date requested
- + Captures a clean, well-lit café ambiance with a polished frame
- − Terrible text rendering with many spelling errors and redundant lines
- − Handwriting looks more like a digital marker font than authentic textured chalk
- − Fails to provide the elegant cursive style for the title
Grok Imagine Image
- + Excellent prompt adherence with near-perfect spelling of all menu items
- + Highly realistic chalk texture with authentic smudge marks and dust
- + Beautiful elegant cursive handwriting for the title as requested
- − Slightly cuts off the bottom of the sign in the composition
Verdict: Grok Imagine is the clear winner as it followed every instruction, including specific text strings, handwriting styles, and the requested chalk texture. FLUX.1 [schnell] FP8 struggled significantly with the text rendering, producing incoherent phrases and 'gibberish' while failing to capture the authentic look of chalk on a board.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
FLUX.1 [schnell] FP8
- + Successfully placed the horse physically on top of a technological object related to space travel.
- + Cinematic lighting and composition against the planet background.
- − Failed to render a recognizable human astronaut, substituting it with a generic machine/box.
- − Anomalous second horse head protruding from the back of the machine.
Grok Imagine Image
- + Clear representation of both a horse and an astronaut as requested.
- + Successfully adhered to the 'horse riding astronaut... horse on top' instruction by positioning the horse above the person.
- + Vibrant, surreal nebula aesthetic that fits the 'highly detailed' prompt requirement.
- − The astronaut is floating underneath rather than being physically 'ridden' in a traditional sense.
- − Slight anatomical inconsistencies in where the horse's legs meet the astronaut's hands.
Verdict: Grok Imagine Image followed the complex prompt much better by including both subjects and maintaining the requested 'horse on top' spatial relationship. FLUX.1 [schnell] FP8 failed to generate a human astronaut entirely, resulting in a confusing image of a horse riding a piece of equipment with a detached second horse head.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
FLUX.1 [schnell] FP8
- + Excellent fur texture and lighting on the capybara.
- + Clear, legible 'TAXI' text on the hat.
- + Vibrant color palette that captures a New York night vibe.
- − The passenger is holding two phones unnaturally.
- − The anatomy of the capybara's 'hands' on the wheel is slightly distorted.
- − The perspective makes the capybara look like it is sitting in the back or middle rather than the driver's seat.
Grok Imagine Image
- + Highly realistic 'bored' expression on the passenger.
- + Correct positioning of the driver and passenger in a standard vehicle cabin.
- + Superior photorealism in the car interior and background street details.
- − The capybara's claws are slightly too sharp/long for the species.
- − The hat is a plain yellow cap without the requested driver specific text.
Verdict: Model B (Grok Imagine Image) is the winner due to its superior composition and grounded realism. While Model A (FLUX.1 [schnell] FP8) captures the whimsical nature of the prompt well, it fails on basic logic by giving the passenger two phones and having an awkward interior perspective, whereas Grok perfectly captures the requested 'bored' atmosphere and realistic taxi environment.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
FLUX.1 [schnell] FP8
- + Strong cinematic lighting with a vibrant center glow
- + Good use of shadows to create a moody gothic atmosphere
- − Significant spelling errors in the scroll banner and event details
- − The border is simplified and lacks the requested thorns and webs
Grok Imagine Image
- + Excellent text rendering with near-perfect accuracy for the entire prompt
- + Includes all requested decorative elements like thorns, spiderwebs, and a moon
- + Authentic vintage parchment texture and layout
- − The pumpkin's lighting is a bit flat compared to the atmospheric background
Verdict: Grok Imagine Image significantly outperforms FLUX.1 [schnell] FP8 by adhering to every specific detail of the prompt, including the complex text sections and the decorative border elements. While FLUX.1 managed a moody atmosphere, its text is riddled with typos and it missed several visual requirements such as the thorns and webs.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
FLUX.1 [schnell] FP8
- + Excellent 3D miniature 'soft' aesthetic
- + Clean isometric perspective
- + High clarity and resolution
- − Failed text rendering, producing a garbled second line
- − Missed several prompt instructions including the word 'SUSHI'
Grok Imagine Image
- + Perfect adherence to text and icon requirements
- + Detailed textures on the rice and salmon
- + Substantially better lighting and shadows for 3D depth
- − The perspective is slightly lower than the requested 45-degree angle
- − A bowl of soy sauce was added which was not specifically requested
Verdict: While FLUX.1 [schnell] FP8 captures a very pleasing 'soft' cartoon style, it failed significantly on the text rendering. Grok Imagine Image followed every specific detail of the prompt, including the text hierarchy and the flag icon, while maintaining high-quality PBR-style textures and realistic lighting.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
FLUX.1 [schnell] FP8
- + Excellent depiction of backlighting and golden hour atmosphere
- + High detail in the flower meadow and butterfly design
- + Clean, sharp rendering of eyes and facial features
- − Failed to include a rabbit in the scene
- − Included extra cats/fox-like hybrids instead of the distinct specified animals
- − Static posing rather than the requested tumbling and chasing action
Grok Imagine Image
- + Correctly included all four requested animals: puppy, kitten, bunny, and fox
- + Successfully captured the dynamic movement and 'tumbling' requested in the prompt
- + Good implementation of dew sparkles and sunburst effects
- − Anatomy is a bit distorted, such as the fox having human-like hands/paws
- − The fur texture appears somewhat oily or over-sharpened compared to natural fur
- − The golden retriever puppy has slightly unusual facial proportions
Verdict: While FLUX.1 [schnell] FP8 offers a more polished and photorealistic aesthetic, it failed the prompt's count and species requirements by omitting the bunny and duplicating feline-like creatures. Grok Imagine Image successfully included the full variety of animals and captured the requested action, making it the better choice for prompt adherence despite some anatomical irregularities.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
FLUX.1 [schnell] FP8
- + Excellent interpretation of a vintage emblem
- + Beautiful warm tones and subtle texture
- + High-quality vector aesthetic
- − Major spelling errors in the brand name ('AFe FLAMILAN')
- − The central icon looks more like a dome or building than a food cloche
Grok Imagine Image
- + Perfect text rendering of 'Caffè Florian'
- + Clearer literal representation of a cloche dome with steam
- + Strong minimalist composition suitable for a logo
- − Redundant 'Est. 1720' text appearing twice
- − Spoon and handle elements are a bit thick compared to the rest of the stroke work
Verdict: While FLUX.1 [schnell] FP8 captured a more sophisticated vintage feel with its texture and emblem shape, it failed significantly on the text rendering. Grok Imagine produced a clean, professional logo with perfect spelling, which is essential for brand identity even though it included the date twice.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
FLUX.1 [schnell] FP8
- + Stronger adherence to the flat-vector, minimalist design aesthetic
- + Excellent use of the specified NASA-inspired color palette
- + Professional layout logic for an infographic
- − Nonsense filler text throughout the image
- − Iconography is overly abstract and doesn't clearly represent the Saturn V or specific steps
- − Failed to follow the requested 6-step sequence correctly
Grok Imagine Image
- + Successfully followed all 6-step instructions in order
- + Near-perfect rendering of mission names (Armstrong, Aldrin, Collins, Tranquility)
- + Clear, representational icons for every stage including the Saturn V and Lunar Module
- − Included 'NASA inspired' as literal text on the poster
- − Visual style is more illustrative and slightly cluttered compared to the requested minimalist vector look
- − Minor spelling errors on technical terms like '3rajoory'
Verdict: Grok Imagine Image is the clear winner for its superior prompt adherence, successfully mapping out all six requested stages of the mission with recognizable icons and accurate names for the crew and landing site. FLUX.1 [schnell] FP8 captured the requested aesthetic and color palette more accurately, but it failed to provide the specific content requested and filled the poster with illegible text.
Explore each model
An image generation model by xAI designed to generate highly aesthetic images from text descriptions.