FP8 quantized variant of Black Forest Labs' FLUX.1 [schnell] model, offering ~2x faster inference with reduced precision while maintaining high-quality image generation in 4 steps
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
FLUX.1 [schnell] FP8
#47 of 62 in Text-to-Image
GPT Image 1.5
#7 of 62 in Text-to-Image
Where the votes landed
FLUX.1 [schnell] FP8
0.0%
win rate
Ties
100.0%
GPT Image 1.5
0.0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
FLUX.1 [schnell] FP8
- + Excellent photorealistic lighting and window reflections
- + High clarity and sharp focus
- + Elegant, modern aesthetic
- − The glass container is a rectangular prism rather than a cube
- − The sphere appears to be floating on an internal shelf rather than sitting on the bottom
GPT Image 1.5
- + Accurately depicts a cubic shape for the glass
- + Correctly places the sphere on the bottom surface as expected
- + Good adherence to the requested spatial relationship with the plant directly behind the cube
- − Texture on the red book is slightly muted compared to the surroundings
- − Slightly less 'premium' photographic feel compared to the other model
Verdict: Both models followed the complex spatial instructions well. FLUX.1 [schnell] FP8 produced a more visually stunning, high-end photograph but failed to create a literal cube, opting for a tall prism instead. GPT Image 1.5 adhered better to the specific geometry of a cube and the placement of the internal object, making it the more accurate representation of the prompt.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
FLUX.1 [schnell] FP8
- + Excellent color vibrance and cinematic lighting.
- + Strong bokeh effect in the background.
- + Good composition using the frame to focus on the subject.
- − The man appears to be standing over the bike rather than actively repairing it.
- − The cars in the background are stationary and sharp, failing the 'motion blur' prompt.
- − Visible rain is almost non-existent compared to the wet floor.
GPT Image 1.5
- + Perfectly captures the action of 'repairing' with a tool box and crouching posture.
- + Excellent depiction of rain droplets on the jacket and bicycle frame.
- + Features subtle motion blur on the passing vehicle as requested.
- − The bicycle geometry is slightly warped near the rear gears.
- − Skin texture is a bit smooth despite the 'natural skin texture' prompt.
Verdict: GPT Image 1.5 is the clear winner as it adheres to nearly every technical aspect of the prompt, including the rain, motion blur, and the specific action of repairing. FLUX.1 [schnell] FP8 produces a high-quality portrait, but the subject is just holding a bike and the cars in the background are clearly standing still, missing the secondary prompt requirements.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
FLUX.1 [schnell] FP8
- + Excellent high-contrast dramatic lighting
- + Intense and lifelike eye detail
- + Very sharp texture on the skin and facial hair
- − Missed the request for scars and dirt on the skin
- − Hair is only minimally braided compared to the prompt
- − Leather and cloth textures are mostly obscured or out of frame
GPT Image 1.5
- + Perfect adherence to scars, dirt, and braided hair with beads
- + Beautifully rendered ornate plate armor with warm light reflections
- + Includes all requested texture elements like leather straps and cloth layers
- − The bokeh sparks appear slightly flat in some areas
- − The skin texture is slightly smoother/more processed than Model A
Verdict: GPT Image 1.5 is the clear winner as it followed every detail of the prompt, including the specific request for scars, dirt, and complex leather/cloth layering which FLUX.1 [schnell] FP8 largely ignored. While FLUX.1 produced a very striking facial portrait, GPT Image 1.5 captured the 'battle-worn' aesthetic and the armor details much more effectively.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
FLUX.1 [schnell] FP8
- + Provides a complete multi-page layout view
- + Includes many distinct food photos to fill the grid
- − Text is mostly gibberish with frequent typos
- − Significant spelling errors in headers (e.g., APPTIZERS, PIZZAL, SECCER)
- − Food images are repetitive and less realistic
GPT Image 1.5
- + Excellent text legibility and realistic copy
- + High-quality, appetizing food photography
- + Clean and professional modern minimalist UI
- − Crop is slightly tight on the left edge
- − Does not show the full page/booklet context like Model A
Verdict: GPT Image 1.5 is the clear winner as it produces a functional, professional menu with perfectly legible text and appetizing food photography. FLUX.1 [schnell] FP8 captures the layout of a menu booklet well, but fails significantly on text rendering and contains numerous spelling errors.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
FLUX.1 [schnell] FP8
- + Clean, simple layout with high contrast
- + Good lighting on the food ingredients
- − Serious spelling errors in the 'Limited Time' text
- − Incorrect price rendering and missing starburst effect on the main price
- − The burger is largely assembled rather than 'exploded' as requested
GPT Image 1.5
- + Perfect adherence to the 'exploded' burger concept with visible separation
- + Excellent adherence to stylistic font requests, including the fiery starburst
- + Exceptional detail on food textures like the seared patty and dripping sauce
- − The composition is very busy and crowded
- − Small sparks distract slightly from the main subject
Verdict: GPT Image 1.5 is the clear winner as it followed every part of the complex prompt, including the 'exploded' structure and the specific fiery effects for the text and starburst. FLUX.1 [schnell] FP8 failed significantly on the text rendering, with multiple spelling errors and the incorrect price.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
FLUX.1 [schnell] FP8
- + Crisp image quality and high resolution
- + Captures the aesthetic of a wooden-framed chalkboard effectively
- − Significant spelling errors and repeated words across the menu
- − Failure to use cursive style for the title as requested
- − Text rendering resembles a digital marker font rather than gritty chalk texture
GPT Image 1.5
- + Perfect adherence to spelling and specific menu items
- + Excellent chalk texture and realistic handwritten cursive style
- + Accurate layout following all prompt instructions including the date and prices
- − The frame of the chalkboard is slightly cropped at the edges
- − The lighting at the very top is a bit harsh
Verdict: GPT Image 1.5 is the clear winner as it followed all complex text instructions perfectly, including specific spelling, prices, and the requirement for cursive handwriting. FLUX.1 [schnell] FP8 struggled significantly with the text, producing numerous spelling errors and failing to render the title in cursive as requested.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
FLUX.1 [schnell] FP8
- + Successfully followed the difficult spatial instruction of having the horse on top of the astronaut/equipment.
- + High visual clarity and striking cinematic lighting.
- + Creative interpretation of a surreal concept including multiple horse features and space-tech.
- − Anatomy is a bit messy with multiple horse heads and necks emerging from the central figure.
- − The astronaut is represented more as a piece of equipment than a person.
GPT Image 1.5
- + Excellent anatomical detail on the horse and space suit.
- + Rich, busy background with impressive celestial objects and surface textures.
- + High level of technical detail in the harness and lunar lander.
- − Failed the primary spatial prompt instruction: 'horse on top, not vice versa'.
- − Followed the common cliche instead of the specific surreal request.
Verdict: FLUX.1 [schnell] FP8 is the clear winner because it actually attempted the 'horse on top' instruction, resulting in a surreal and unique composition. GPT Image 1.5 produced a much higher quality image in terms of detail and texture, but it completely ignored the specific spatial constraint to follow a standard 'astronaut riding horse' trope.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
FLUX.1 [schnell] FP8
- + Includes a clear 'TAXI' sign on another car to enhance the setting context.
- + High color saturation gives it a polished, vibrant look.
- − The passenger is holding two separate smartphones, which is redundant and looks like an error.
- − The capybara's hat looks more like a plastic toy than a professional taxi cap.
- − The capybara's paws are not correctly positioned on the steering wheel.
GPT Image 1.5
- + Excellent photorealistic texture on the capybara's fur and the leather jacket.
- + The taxi driver cap is much more realistic and detailed.
- + Properly depicts both paws on the steering wheel as requested.
- − The composition is a bit tight, making the passenger in the back slightly harder to see.
- − The lighting on the capybara's face is a bit flat compared to the backgrounds.
Verdict: GPT Image 1.5 is the clear winner for its superior realism and adherence to the specific details of the prompt, particularly the placement of the paws and the style of the driver's cap. FLUX.1 [schnell] FP8 suffers from significant logical errors, such as the woman holding two phones, and the overall image looks more like a 3D render than a photorealistic scene.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
FLUX.1 [schnell] FP8
- + Features a bold, central jack-o-lantern.
- + Captures the 'moody night sky' with a clean silhouette effect.
- − Several severe spelling errors in the banner and event details (e.g., 'A ai nigh tof friights', 'Theaches').
- − The border looks like a flat gradient rather than the requested 'webs and thorns'.
- − Text layout is cluttered and uses inconsistent fonts.
GPT Image 1.5
- + Excellent adherence to all text requirements with perfect spelling.
- + Highly detailed 'webs and thorns' border that matches the vintage gothic aesthetic.
- + Superior composition with a more natural integration of the pumpkin into the environment.
- − The layout is a bit dense with the background details.
- − The jack-o-lantern is slightly off-center.
Verdict: GPT Image 1.5 is the clear winner as it followed every complex text instruction perfectly while maintaining a high-quality vintage aesthetic. FLUX.1 [schnell] FP8 struggled significantly with the typography, resulting in numerous spelling errors and a much less cohesive gothic border.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
FLUX.1 [schnell] FP8
- + Excellent soft, refined 3D cartoon textures
- + Perfectly followed the isometric 45-degree angle
- + Clean, minimalist aesthetic that feels high-quality
- − Failed to render the word 'SUSHI'
- − Text is repetitive and contains a strange hybrid flag character
- − Sushi variety is limited compared to the competitor
GPT Image 1.5
- + Perfect text adherence with 'JAPAN', 'SUSHI', and flag icon correctly placed
- + Rich, realistic PBR materials on the wood, teapot, and food
- + Complex diorama base with moss and multiple accessories adds depth
- − Camera angle is a bit lower than the requested 45-degree top-down view
- − Slightly more cluttered than the 'minimal' request
- − Background blue is a bit darker than 'light blue'
Verdict: GPT Image 1.5 is the clear winner because it followed every text-based instruction perfectly, including the flag icon and 'SUSHI' text which FLUX.1 [schnell] FP8 failed to generate properly. While FLUX.1 [schnell] FP8 had a cleaner 'cartoon' style, it failed the core requirements for typography and icon placement.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
FLUX.1 [schnell] FP8
- + Excellent vibrant colors and high-contrast lighting.
- + Beautifully rendered butterflies with clear silhouettes.
- + Soft, appealing artistic style that fits the 'wholesome' vibe.
- − Failed to include the baby bunny entirely.
- − Includes two kittens and a strange hybrid animal instead of the requested variety.
- − Less realistic; has a distinct digital-art/AI-illustration look.
GPT Image 1.5
- + Followed the prompt accurately by including a puppy, kitten, bunny, and fox kit.
- + Superior 'hyper-photorealistic' textures and fur details.
- + Captured the specific 'tumbling' motion and 'dew sparkles' much more effectively.
- − The fox kit has three front paws visible on one side.
- − Lower contrast compared to model_a, with more hazy lighting.
Verdict: GPT Image 1.5 is the clear winner because it actually included all four requested animals, whereas FLUX.1 [schnell] FP8 failed to generate the bunny and instead duplicated other animals. GPT Image 1.5 also achieved a much higher level of realism and detail in the fur and environment, adhering better to the 'hyper-photorealistic' requirement.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
FLUX.1 [schnell] FP8
- + Strong minimalist vector aesthetic
- + Includes the specific 'Est. 1720' text
- − Significant spelling errors in the primary name ('Caffé AFe FLAMILAN')
- − Misinterpreted cloche dome as a building-like dome structure
GPT Image 1.5
- + Perfect text rendering for 'Caffè Florian' and 'Est. 1720'
- + Accurately depicts a food cloche dome with steam as requested
- + Excellent use of texture and shading to create a vintage feel
- − Ignored the 'light background' instruction, opting for a black background
Verdict: GPT Image 1.5 is the clear winner for its superior text accuracy and literal interpretation of the cloche dome. While FLUX.1 [schnell] FP8 followed the color scheme and background prompt better, its failure to spell the restaurant name correctly makes it unusable as a logo.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
FLUX.1 [schnell] FP8
- + Clean typography for the main title
- + Minimalist layout that feels more like a professional vector poster
- − Nonsense placeholder text for the body and labels
- − Icons do not accurately reflect the specific steps requested
- − Composition feels sparse and unfinished
GPT Image 1.5
- + Perfect text rendering of all labels and names
- + Accurate iconography for every mission step requested
- + Stronger narrative flow with a clear panels layout
- − Some minor color bleed on the edges of the vector boxes
- − The red is slightly more saturated than 'muted red'
Verdict: GPT Image 1.5 is the clear winner as it followed every instruction, including the specific iconography for each of the six mission steps and rendering the text accurately. FLUX.1 [schnell] FP8 failed to produce coherent text and the icons were generic and repetitive, failing to distinguish between technical stages like translunar injection and lunar orbit.
Explore each model
OpenAI's state-of-the-art image generation model with better instruction following and adherence to prompts