Black Forest Labs' 12 billion parameter distilled image generation model optimized for speed, capable of generating high-quality images in just 4 inference steps
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
FLUX.1 [schnell]
#48 of 62 in Text-to-Image
GPT Image 1
#32 of 62 in Text-to-Image
Where the votes landed
FLUX.1 [schnell]
100.0%
win rate
Ties
0.0%
GPT Image 1
0.0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
FLUX.1 [schnell]
- + Excellent rendering of light and reflections on the glass and wood
- + High visual quality with vibrant colors
- + Captures the 'plant partially visible through glass' instruction well
- − Added an extra blue glass sphere on top of the book that was not requested
- − The blue sphere appears to be floating inside the cube without gravity or support
GPT Image 1
- + Strict adherence to the prompt with no extra objects
- + More realistic glass texture with slight green tint on the edges
- + Clean composition and accurate lighting
- − The plant is mostly behind the cube but less visible 'through' the glass refraction compared to Image A
- − The sphere appears to be floating slightly above the bottom surface
Verdict: GPT Image 1 followed the prompt instructions more accurately than FLUX.1 [schnell], which added an unnecessary second blue sphere on top of the book. While FLUX.1 [schnell] produced more dramatic lighting and complex reflections, GPT Image 1 is preferred for its precise prompt adherence and realistic material properties.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
FLUX.1 [schnell]
- + Excellent depiction of wet pavement and light rain reflections
- + High resolution with sharp details on the bicycle and man's face
- + Good use of color with the red bicycle and vest pops against the grey street
- − Fails to include the requested motion blur from passing cars
- − The bicycle geometry is slightly warped, specifically the handlebars and pedals
- − The background cars look static and parked rather than passing by
GPT Image 1
- + Successfully captures the requested crouched 'repairing' pose
- + Includes subtle motion blur in the background traffic as requested
- + Achieves a very realistic, non-stylized skin texture and candid feel
- − The framing is very tight, cutting off much of the bicycle
- − Minor anatomical issues with the fingers merging near the gears
- − Slightly lower global contrast compared to Model A
Verdict: GPT Image 1 is the winner as it adheres much better to the specific technical requirements of the prompt, including the motion blur and the act of 'repairing' through a crouched posture. While FLUX.1 [schnell] produced a sharper image with better colors, it ignored the motion blur and the man appears to be simply holding the bike rather than fixing it.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
FLUX.1 [schnell]
- + extremely sharp detail on the iris and skin texture
- + very dramatic and intense facial expression
- + strong lighting contrast that highlights facial features
- − armor is cut off by the tight crop, losing the 'ornate' feel
- − braids appear somewhat thick and rope-like rather than fine
- − lacks the requested 'bokeh sparks' in the background
GPT Image 1
- + perfectly captures the ornate engraved plate armor as requested
- + excellent execution of bokeh sparks and warm torchlight atmosphere
- + most accurate depiction of battle-worn skin with realistic dirt and faint scars
- − the eyes are slightly less luminous than model a
- − overall image is a bit darker, which obscures some detail in the hair
Verdict: While FLUX.1 [schnell] provides a high-intensity facial portrait, GPT Image 1 follows the complex prompt much more faithfully. GPT Image 1 successfully incorporates the ornate armor, bokeh sparks, and battle-worn textures that were specifically requested, whereas FLUX.1 [schnell] cropped most of the armor out.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
FLUX.1 [schnell]
- + Strong minimalist layout that feels like a real physical menu.
- + Includes all three requested sections (Appetizers, Pizza, Mains).
- + Features a distinct grid arrangement for the food photos.
- − Text is largely gibberish or extremely distorted.
- − Some food photos look repetitive or low quality.
- − Includes a typo/strange word 'ORFEFUS' in a header.
GPT Image 1
- + Excellent food photography quality and realism.
- + Text is highly legible with clear prices and subheaders.
- + Successfully incorporates vibrant accents with the colored square icons.
- − Fails to include a specific section for 'Mains', instead mixing it into the 'Pizza' column.
- − Repeats the same placeholder description multiple times.
- − The grid layout is slightly uneven due to the text length.
Verdict: GPT Image 1 is the superior design due to its high-quality food photography and exceptionally legible bold sans-serif typography. While FLUX.1 [schnell] followed the structural layout of the prompt more closely by including all three requested sections, its rendering of text and food images is significantly less professional and contains more artifacts.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
FLUX.1 [schnell]
- + Excellent photorealistic lighting and texture on the burger bun and cheese.
- + Creative use of flying croutons and splatters to convey motion.
- + High-quality rendering of the fiery background embers.
- − Failed the text objective by clipping the 'M' in 'Magic' and duplicating the price incorrectly.
- − The burger is largely assembled rather than 'exploded' into mid-air components as requested.
GPT Image 1
- + Perfectly followed the 'exploded burger' instruction with clearly separated layers.
- + Exactly rendered the requested fiery glow effect on all text elements.
- + All prompt-required text is present, legible, and correctly placed.
- − The price in the starburst is missing the '6' before the decimal point (€.99).
- − The overall composition is somewhat static despite the exploded view.
Verdict: GPT Image 1 followed the complex prompt requirements much more accurately, successfully exploding the burger layers and applying the specific fiery text effects requested. While FLUX.1 [schnell] produced a more lifelike burger texture, it failed significantly on the text rendering ('AGIC BURGER') and did not truly explode the ingredients.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
FLUX.1 [schnell]
- + The text has a natural, imperfect handwritten feel.
- + Captures the chalkboard aesthetic well with high-contrast white on black.
- − Poor spelling throughout the menu (e.g., 'Taffle Mushmnctiomn', 'Octtoopus').
- − Fails to render the title in 'elegant cursive' as requested, using a basic print style instead.
- − The date text is truncated to 'Pril'.
GPT Image 1
- + Excellent text rendering with no spelling errors in the main items.
- + Realistic chalk texture with visible grain and slight pressure variations.
- + Follows the prompt details closely, including the specific date and menu items.
- − The title is in a stylized print rather than the requested 'elegant cursive'.
- − The prices for the cookies at the bottom are slightly cut off/incomplete.
Verdict: GPT Image 1 is the clear winner as it successfully rendered almost all the requested text without the severe spelling and gibberish errors present in FLUX.1 [schnell]. GPT Image 1 also provided a much more realistic chalk texture and followed the specific menu item instructions accurately.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
FLUX.1 [schnell]
- + Successfully followed the difficult spatial constraint of placing the horse on top of the astronaut.
- + High cinematic lighting with a pleasing warm glow from the planet.
- + Creative and surreal interpretation that matches the prompt's intent.
- − Anatomical glitch featuring a horse with two heads/necks.
- − The astronaut's pose is somewhat nonsensical and disconnected from the backpack.
GPT Image 1
- + Very high level of detail on the space suit and horse's mane.
- + Good composition with a clear sense of atmosphere and lighting.
- + Clean anatomical rendering of both the horse and the astronaut.
- − Completely failed the negative constraint/spatial instruction for the horse to be on top.
- − Produces a generic 'astronaut riding a horse' image despite the specific prompt instruction.
Verdict: While GPT Image 1 is higher quality in terms of anatomical correctness, it completely ignored the specific spatial instruction to place the horse on top of the astronaut. FLUX.1 [schnell] successfully followed the 'horse on top' constraint, making it a much better match for the surreal request despite the minor glitch of a double-headed horse.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
FLUX.1 [schnell]
- + Excellent fur texture rendering
- + Vibrant and cinematic lighting that fits the NYC night theme
- + Clear text on the cap and interior visor
- − The paws are not both on the steering wheel correctly
- − The capybara's head shape is slightly distorted to look more like a dog-capybara mix
GPT Image 1
- + Follows the instruction for both front paws on the steering wheel better
- + Perfectly captures the bored, indifferent expression of the human passenger
- + Highly accurate capybara facial anatomy
- − Lighting is a bit flat compared to the other model
- − The texture of the paws is a bit creepy and less realistic than the face
Verdict: GPT Image 1 followed the specific pose instructions more accurately, placing both paws on the wheel and capturing the 'bored' expression of the passenger perfectly. FLUX.1 [schnell] had superior lighting and fur detail, but the capybara's anatomy was slightly less convincing and it missed the specific paw placement.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
FLUX.1 [schnell]
- + Successfully captures a vibrant, high-contrast glow on the jack-o-lantern.
- + Includes clear, illustrative elements like the twisted trees and bats as requested.
- − Several major spelling and grammatical errors in the text (e.g., 'You invited to a a night', 'firiichts', 'butistigtion').
- − Poor layout of event details with redundant and incorrect time fields (Time: 17.2026).
GPT Image 1
- + Perfect text rendering for all requested strings, including the banner and event details.
- + Superior composition that feels like a cohesive, vintage gothic poster with a gritty parchment texture.
- + Elegant used of the thorny border and cobwebs that perfectly matches the prompt.
- − Lighting is slightly more muted compared to the requested 'cinematic' glow in Model A.
Verdict: GPT Image 1 is the clear winner as it perfectly follows all text instructions and maintains a sophisticated vintage gothic aesthetic. In contrast, FLUX.1 [schnell] struggles significantly with the requested text, producing numerous spelling errors and a cluttered bottom section.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
FLUX.1 [schnell]
- + Excellent high-clarity 3D rendering with realistic PBR properties.
- + Perfectly centered minimal diorama aesthetic.
- − Missed the 'SUSHI' text requested in the prompt.
- − Low contrast on top-center text makes it hard to read.
GPT Image 1
- + Followed all text instructions including 'JAPAN' and 'SUSHI'.
- + Clean isometric composition with attractive clay-like textures.
- − Added extra elements like chopsticks and ginger not specifically requested.
- − The lighting is slightly flat compared to the other model.
Verdict: FLUX.1 [schnell] produced a more sophisticated 3D render with superior material realism, but it failed to include the secondary text line correctly. GPT Image 1 followed the complex text instructions perfectly and captured the miniature diorama feel well, making it more accurate to the prompt despite slightly simpler textures.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
FLUX.1 [schnell]
- + Excellent fur texture and lighting details on the grass
- + Good color saturation and vibrance
- − Failed to generate a distinct bunny, creating a weird cat/bunny hybrid instead
- − The anatomy of the two center animals is merged and confusing
- − Static posing feels less like 'chasing' and more like a portrait
GPT Image 1
- + Successfully included all four distinct species listed in the prompt
- + Captures the action of 'playfully chasing' and 'tumbling' much better
- + Beautiful rendering of god rays and morning dew consistent with the prompt
- − The fox's front paws look a bit structurally simplified
- − The butterflies are less defined than in Model A
Verdict: GPT Image 1 followed the prompt significantly better by correctly identifying and rendering all four separate animals (dog, cat, bunny, and fox), whereas FLUX.1 [schnell] failed to create a distinct bunny and instead blended features. GPT Image 1 also better captured the dynamic energy of the animals playing and chasing, and accurately depicted the atmospheric 'god rays' requested.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
FLUX.1 [schnell]
- + Clean vector emblem style
- + Well-balanced circular composition
- − Serious spelling errors in the main name ('FRAMILAN')
- − Incorrect date ('7720' instead of '1720')
- − Missing the requested steam element
GPT Image 1
- + Perfect text rendering for name and date
- + Includes the requested steam and banner elements
- + Excellent texture application for a vintage look
- − Uses a black background instead of the requested light/cream background
- − Slightly less 'minimalist' than Model A's linework
Verdict: While GPT Image 1 missed the requested light background, it succeeded in every other aspect including correct spelling, accurate date, and the inclusion of steam. FLUX.1 [schnell] failed significantly on the text, changing the name to 'Framilan' and the date to '7720'.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
FLUX.1 [schnell]
- + Clean layout with a high-end vector feel
- + Good use of negative space and composition
- − Text is complete gibberish and ignores the specific step labels
- − The rocket design is generic and doesn't resemble a Saturn V
- − Fails to clearly represent the sequential 6-step infographic structure requested
GPT Image 1
- + Excellent adherence to the 6 steps with relevant icons for each
- + Text rendering is highly accurate and readable
- + Incorporated the specific mission crew and landing site names accurately
- − The layout is a bit cluttered with uneven spacing between labels
- − The Saturn V icon is slightly simplified compared to the rest of the professional vector art
Verdict: GPT Image 1 followed the complex multi-step prompt almost perfectly, including accurate text for the mission stages and crew. FLUX.1 [schnell] produced a visually striking image, but failed significantly on text legibility and ignored the specific sequence of events requested.
Explore each model
OpenAI's previous image generation model that accepts both text and image inputs and produces image outputs