Black Forest Labs' 12 billion parameter distilled image generation model optimized for speed, capable of generating high-quality images in just 4 inference steps
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
FLUX.1 [schnell]
#48 of 62 in Text-to-Image
GPT Image 2
#4 of 62 in Text-to-Image
Where the votes landed
FLUX.1 [schnell]
0%
win rate
Ties
0%
GPT Image 2
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
FLUX.1 [schnell]
- + Excellent photographic lighting and depth of field.
- + High-quality glass reflections and refraction effects.
- − Included a second blue sphere on top of the book which was not requested.
- − The sphere inside appears to be floating unnaturally.
GPT Image 2
- + Perfect adherence to all prompt elements with no extra objects.
- + Realistic physical interaction with the sphere resting on the cube floor.
- + Accurate representation of the green plant behind the glass.
- − The glass cube edges appear slightly thinner and less like solid glass than Model A.
Verdict: While FLUX.1 [schnell] produced a more visually striking image with superior glass physics, it failed the prompt by adding a second blue sphere on top of the red book. GPT Image 2 followed every instruction perfectly, including the specific placement of the plant and the single sphere inside the cube, making it the more accurate result.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
FLUX.1 [schnell]
- + Excellent handling of wet pavement reflections and rainy atmosphere.
- + Cinematic composition with beautiful bokeh and lighting.
- + The subject's face and hair texture are highly detailed.
- − Physical logic of the hands holding the handlebars is slightly messy.
- − Cars in the background lack the requested motion blur, appearing frozen instead.
GPT Image 2
- + Captures the motion blur of the passing car perfectly as requested.
- + Realistic crouched pose and inclusion of a tool kit adds to the storytelling.
- + Excellent source preservation of Japanese street aesthetic and signage.
- − The bike's frame and rear wheel assembly have significant structural glitches.
- − Skin texture on the hands is somewhat muddy compared to the face.
Verdict: Both models followed the prompt well, but GPT Image 2 better captured the technical request for motion blur and the 'candid' feel of a man actually performing a repair. While FLUX.1 [schnell] produced a more aesthetically pleasing and high-contrast 'cinematic' image, GPT Image 2 felt more like a genuine street photograph despite some structural issues with the bicycle's geometry.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
FLUX.1 [schnell]
- + Features very intense and sharp facial details
- + Includes a strong warm light source reflecting on surfaces
- − The textures look somewhat digital and overly sharpened
- − Braids lack the specific small beads mentioned in the prompt
GPT Image 2
- + Successfully captures both braids and beads with realistic hair texture
- + Excellent balance of ornate armor engraving and realistic skin dirt
- − The depth of field is slightly shallower than requested for a close portrait
- − Lighting is a bit more diffused than the requested torchlight
Verdict: While FLUX.1 [schnell] provides a very close and striking portrait, GPT Image 2 offers a more nuanced interpretation of the prompt requirements. GPT Image 2 better integrates specific elements like the hair beads and the interplay between leather and cloth materials.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
FLUX.1 [schnell]
- + Clean minimalist layout with ample white space
- + Good use of bold sans-serif headers
- − Text is mostly illegible gibberish
- − Food images are repetitive and low quality
- − Grid layout feels cramped and disjointed
GPT Image 2
- + Exceptional text rendering with coherent item names and descriptions
- + High-quality, distinct food photography for every item
- + Well-organized layout that follows all prompt instructions perfectly
- − Slightly more cluttered than a strict minimalist design
Verdict: GPT Image 2 is significantly better than FLUX.1 [schnell] in every functional category, providing fully legible text, high-quality food photography, and a professional layout. While FLUX.1 [schnell] attempted a clean aesthetic, it failed to produce usable text or clear food visuals, whereas GPT Image 2 created a production-ready menu design.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
FLUX.1 [schnell]
- + Great photographic texture on the burger bun and melting cheese.
- + Dynamic debris flying around the main subject adds a sense of explosion.
- + Realistic depth of field in the background.
- − Failed the primary text prompt, spelling it 'AGIC BURGER'.
- − The prices are inconsistent and cluttered at the bottom.
- − The burger is mostly assembled rather than 'exploded' into individual components.
GPT Image 2
- + Perfect adherence to text instructions including the fiery, glowing effect.
- + High-quality 'exploded' composition with all components clearly separated and suspended.
- + Professional ad layout with excellent balance between text and imagery.
- − The fiery effect on the text can make the smaller horizontal lines slightly harder to read.
Verdict: GPT Image 2 is the clear winner as it followed every instruction in the prompt, including complex text rendering and the specific 'exploded' layout. FLUX.1 [schnell] failed significantly on text accuracy and provided a largely assembled burger rather than the exploded view requested.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
FLUX.1 [schnell]
- + Legible handwriting style.
- + Correct date and pricing numbers provided in some lines.
- − Numerous spelling errors including 'Pril', 'Taffle Mushmnctiomn', and 'Octtoopus'.
- − Fails to render the title in the requested 'elegant cursive' style.
- − The chalk texture looks more like a digital marker than real chalk.
GPT Image 2
- + Excellent adherence to the 'elegant cursive' title and handwritten chalk aesthetic.
- + Flawless text rendering for all requested menu items with no spelling errors.
- + Realistic chalk texture with grains and smears that fit a cozy café atmosphere.
- − The date is slightly crowded against the right edge of the board.
- − The bottom text line introduces a slight slant variation compared to the main list.
Verdict: GPT Image 2 is significantly better, following every instruction including the specific cursive title and perfect spelling of complex menu items. FLUX.1 [schnell] struggles with spelling and fails to capture the requested chalk texture and cursive style, resulting in a more generic look.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
FLUX.1 [schnell]
- + Excellent cinematic lighting and color palette.
- + High visual quality with a clear focal point.
- + Successfully places the horse on top of the astronaut.
- − Anatomical failure resulting in a two-headed horse.
- − The astronaut's posture is confusing and appears detached from the horse.
- − The concept of 'riding' is less clear than Model B.
GPT Image 2
- + Perfect adherence to the prompt with a horse literally riding a saddled astronaut.
- + Highly detailed textures on the spacesuit and horse fur.
- + Clear and logical composition for the surreal request.
- − The astronaut's hands have too many fingers (6+).
- − The horse's front legs ending in stirrups is a bit of a strange anatomical glitch.
Verdict: Model B (GPT Image 2) is the clear winner for its superior prompt adherence, accurately depicting a horse saddled and riding an astronaut as requested. While Model A (FLUX.1 [schnell]) has much better lighting and a more cinematic feel, it fails physically by generating a two-headed horse and a disjointed astronaut. GPT Image 2's literal interpretation of the surreal request makes for a more successful image despite the finger artifacts.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
FLUX.1 [schnell]
- + Excellent fur texture and lighting on the capybara subject.
- + Captures the bored expression of the passenger perfectly.
- + High clarity and cinematic depth of field.
- − The capybara only has one paw clearly on the steering wheel, failing the prompt's request for both.
- − The cap is a simple beanie rather than a traditional driver cap.
GPT Image 2
- + Successfully placed both front paws on the steering wheel as requested.
- + The capybara's professional driver hat is more authentic to the 'taxi driver' trope.
- + Great background bokeh and rainy city atmosphere.
- − The passenger's face is slightly soft and lacks high detail.
- − The capybara's head looks a bit large and bulky compared to the body and seat.
Verdict: GPT Image 2 followed the specific instructions for the capybara's interaction with the steering wheel and hat style more accurately. While FLUX.1 [schnell] produced a cleaner, sharper image of the passenger and textures, it missed the key detail of having both paws on the wheel and used a generic cap instead of a driver's hat.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
FLUX.1 [schnell]
- + Includes all specific prompt elements like bats and twisted trees
- + The Jack-o'-lantern is bright and central as requested
- − Several typos in both the banner text and event details
- − The lighting and overall style feel consistent with a vector illustration rather than a vintage gothic poster
GPT Image 2
- + Excellent text rendering with no typos in the titles or main details
- + Highly detailed vintage gothic aesthetic with a complex border and atmospheric parchment texture
- + Cleverly incorporates 'The Arches' and the NYC skyline into the background
- − The border is a bit cluttered, potentially overshadowing the central subject
Verdict: GPT Image 2 is the clear winner as it perfectly captures the 'vintage gothic' aesthetic with professional-grade typography and zero spelling errors. While FLUX.1 [schnell] follows the layout instructions, it fails significantly on text legibility and generates several hallucinated text lines, making it unusable as an invitation.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
FLUX.1 [schnell]
- + Clean minimalist aesthetic
- + High-quality rendered textures on the salmon and base
- + Perfectly followed the 45-degree isometric angle request
- − Failed to include the word 'SUSHI' in the text
- − The flag icon is slightly distorted near the top
- − Very basic composition compared to the complexity of the prompt
GPT Image 2
- + Excellent adherence to full text requirements ('JAPAN', 'SUSHI', and flag)
- + Rich, detailed PBR materials and professional lighting
- + Strong interpretation of the 'miniature 3D cartoon scene' with the garden elements
- − Technically slightly busy despite the 'minimal garnish' request
- − The sushi is not as perfectly centered as Model A
Verdict: GPT Image 2 is the superior response as it followed all textual instructions, including the word 'SUSHI' which FLUX.1 [schnell] omitted. GPT Image 2 also provided a much more engaging 3D scene that captured the desired PBR materials and miniature aesthetic more effectively.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
FLUX.1 [schnell]
- + Excellent soft lighting and dreamy bokeh effect.
- + High level of fur detail and cleanliness in the character designs.
- − Failed to include a distinct bunny, instead creating a hybrid cat-like creature.
- − The composition is static and lacks the 'playfully chasing' action requested.
GPT Image 2
- + Successfully included all four distinct animal types requested.
- + Captures the activity of 'tumbling' and 'chasing' with dynamic poses.
- + Includes visible god rays and dew sparkles as specified in the prompt.
- − The fox's anatomy is slightly warped with a fifth leg or strange limb placement in the background.
- − Some of the distant butterflies are poorly defined.
Verdict: GPT Image 2 followed the prompt much more accurately by including all four specified animals and capturing a sense of movement, whereas FLUX.1 [schnell] failed to generate a bunny and produced a static, posed shot. While FLUX.1 [schnell] has a slightly cleaner artistic style, GPT Image 2 is the superior response for its adherence to the specific subjects and environmental details like god rays.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
FLUX.1 [schnell]
- + Strong minimalist vector aesthetic
- + Clean icon design for the cloche
- − Significant spelling errors in the brand name ('CAFEÉ FRAMILAN')
- − Incorrect year in the date ('7720' instead of '1720')
- − Missing the requested steam element
GPT Image 2
- + Perfect adherence to text prompt including name and date
- + Excellent vintage texture and detailed shading
- + Cohesive composition with all requested elements (cloche, steam, banner)
- − Slightly more complex than 'minimalist' might imply
- − The steam lines are somewhat stylized and whimsical
Verdict: GPT Image 2 followed the prompt perfectly, accurately rendering the difficult text 'Caffè Florian' and the correct historical date 'Est. 1720', while FLUX.1 [schnell] failed significantly on all text elements. GPT Image 2 also successfully included the steam and banner elements with a high-quality vintage texture that matches the requested aesthetic better than the overly simplified and misspelled version from FLUX.1 [schnell].
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
FLUX.1 [schnell]
- + Clean vector aesthetic that matches the requested flat-vector style.
- + Follows the color palette accurately.
- − Text is largely illegible gibberish.
- − Information layout is confusing and does not clearly follow the 6-step progression requested.
GPT Image 2
- + Excellent adherence to all 6 specific steps of the infographic.
- + High-quality text rendering and accurate data presentation.
- + Professional layout with high visual appeal and NASA-inspired theme.
- − Included images of the Moon and Earth use photographic textures rather than being purely flat-vector.
- − The scale of the Saturn V rocket is inconsistent with the icons next to it.
Verdict: While FLUX.1 [schnell] captures the 'flat-vector' aesthetic better, GPT Image 2 is the superior infographic because it actually follows the 6-step prompt requirements and features perfectly legible text. FLUX.1 [schnell] failed to create a logical flow or readable steps, making it ineffective as an information graphic.
Explore each model
OpenAI's state-of-the-art image generation model with arbitrary resolution up to 4K and strong instruction following