Head to head
Esc

Models · slot A

to navigate to pick

FLUX.1 [schnell] Black Forest Labs GPT Image 1.5 OpenAI

Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.

FLUX.1 [schnell]

18.7 arena score

#48 of 62 in Text-to-Image

Skill signature · Text-to-Image

GPT Image 1.5

27.1 arena score

#6 of 62 in Text-to-Image

Top 3 in Image Editing
Vote tally

Where the votes landed

FLUX.1 [schnell]

0.0%

win rate

Ties

0.0%

GPT Image 1.5

100.0%

win rate

0.0% 0.0% ties 100.0%
Shared challenges 13

Challenge by challenge

The strongest take from each model on every shared challenge, with the AI judge's read.

Geometric Composition

Text-to-Image

“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”

FLUX.1 [schnell]
GPT Image 1.5

AI Judge Analysis

FLUX.1 [schnell]

  • + Excellent vibrant colors and high-resolution textures.
  • + Creative interpretation with a floating sphere effect.
  • + Follows lighting directions well.
  • Includes an extra blue sphere on top of the book not requested in the prompt.
  • The glass cube looks more like a solid block of glass due to the thick edges.

GPT Image 1.5

  • + Perfect adherence to the prompt requirements without adding extra objects.
  • + Highly realistic glass transparency and reflections.
  • + Accurate physical placement of objects.
  • The plant in the background is slightly less detailed/distinct than in the other image.
  • Color palette is a bit more muted.

Verdict: GPT Image 1.5 is the clear winner as it followed every instruction in the prompt perfectly, whereas FLUX.1 [schnell] hallucinated an additional blue sphere on top of the book. While FLUX.1 [schnell] produced a more vibrant and artistic image, GPT Image 1.5 better managed the spatial relationships and transparency of the glass cube.

Candid Street Photography

Text-to-Image

“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”

FLUX.1 [schnell]
GPT Image 1.5
0% wins 0% ties 100% wins

AI Judge Analysis

FLUX.1 [schnell]

  • + Excellent reflection work on the wet pavement
  • + Strong color saturation on the red bicycle
  • + Good composition with the urban background bokeh
  • The man's interaction with the bicycle handlebars is physically illogical
  • Failed to include the requested motion blur for passing cars
  • Missing visible rain drops on the surfaces

GPT Image 1.5

  • + Captures visible rain drops and realistic wet textures on the clothing and bike
  • + Composition feels more 'candid' and 'imperfect' as requested
  • + More realistic tool usage and bodily posture for a repair
  • Lacks the requested motion blur for the car in the background
  • The red tones on the bike are a bit muted compared to Model A

Verdict: GPT Image 1.5 is the winner because it adheres much better to the specific technical atmospheric prompts like light rain and candid framing. While FLUX.1 [schnell] produced a clean image, the subject's interaction with the bike was nonsensical, and it lacked the rain elements requested.

Fantasy Warrior

Text-to-Image

“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”

FLUX.1 [schnell]
GPT Image 1.5

AI Judge Analysis

FLUX.1 [schnell]

  • + Extremely sharp facial details and skin texture
  • + Intense, lifelike eyes and gaze
  • Missed specific details like beads in the braids
  • Lighting feels a bit artificial/flat compared to a real torch
  • Minimal visible scars or dirt despite the prompt request

GPT Image 1.5

  • + Excellent adherence to specific details like hair beads and bokeh sparks
  • + Realistic application of scars, dirt, and battle-worn features
  • + Superior lighting and material rendering on the engraved armor
  • The cloth underlayer texture is slightly less crisp than the metal and skin

Verdict: GPT Image 1.5 is the clear winner as it captured every detail of the prompt, including the hair beads, bokeh sparks, and realistic battle damage, which FLUX.1 [schnell] largely ignored. While FLUX.1 [schnell] produced a high-quality portrait, GPT Image 1.5 offered a much more cinematic and accurate interpretation of a 'battle-worn' character.

Modern Clean Menu

Text-to-Image

“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”

FLUX.1 [schnell]
GPT Image 1.5

AI Judge Analysis

FLUX.1 [schnell]

  • + Clean minimalist layout with high-quality white space
  • + Professional typography for the main headers
  • Nonsense filler text for descriptions and price points
  • Photos are repetitive and do not accurately represent different food categories well
  • Hallucinated 'ORFEFUS' as a section header instead of 'MAINS'

GPT Image 1.5

  • + Perfect text rendering with legible, realistic menu items and descriptions
  • + Food photos are diverse and perfectly matched to the categories
  • + Strong use of vibrant accents and a clean grid layout that follows the prompt exactly
  • Slightly less 'minimalist' than Model A due to higher density of information

Verdict: GPT Image 1.5 is the clear winner as it produces a functional, realistic menu with perfectly rendered text and appropriate food imagery for each category. FLUX.1 [schnell] fails on the text details, providing garbled placeholder text and a hallucinated section header, making the design unusable.

Magic Burger Explosion: Fiery Photorealism Challenge

Text-to-Image

“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”

FLUX.1 [schnell]
GPT Image 1.5

AI Judge Analysis

FLUX.1 [schnell]

  • + Features a high-contrast, clean professional layout
  • + The central burger has excellent lighting and texture
  • + Incorporates the starburst element requested in the prompt
  • Significant text errors including 'AGIC' instead of 'MAGIC' and a cluttered price display
  • Failed to 'explode' the burger; it is mostly intact with random food chunks floating around it

GPT Image 1.5

  • + Perfectly captures the 'exploded' effect with all layers suspended in air
  • + All text is rendered accurately and with the requested fiery, glowing style
  • + Maintains a cohesive, high-energy fiery aesthetic throughout the entire image
  • The starburst shape is slightly less defined compared to more geometric starbursts
  • The overall composition is very busy with embers, which slightly obscures fine details

Verdict: GPT Image 1.5 is the clear winner as it followed every part of the prompt, specifically the 'exploded' burger requirement and accurate text rendering. FLUX.1 [schnell] failed to deconstruct the burger layers and suffered from multiple spelling and formatting errors in the ad copy.

Chalkboard Menu

Text-to-Image

“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”

FLUX.1 [schnell]
GPT Image 1.5

AI Judge Analysis

FLUX.1 [schnell]

  • + Good depth of field with cafe elements in background
  • + Clear, legible handwriting style
  • Significant spelling errors throughout including 'Pril', 'Taffle', and 'Mushmnctionm'
  • Failed to follow the cursive request for the title
  • The handwriting looks more like a digital marker than realistic chalk texture

GPT Image 1.5

  • + Excellent adherence to the requested text with zero spelling errors
  • + Beautiful chalk texture with natural smudging and varied pressure
  • + Perfectly followed the 'elegant cursive' instruction for the title
  • The composition is a very close crop, losing some of the 'cozy café' atmosphere requested
  • The bottom line of text is slightly smaller than naturally expected, though still handwritten

Verdict: GPT Image 1.5 is the clear winner as it followed every complex text instruction perfectly without a single typo, whereas FLUX.1 [schnell] struggled significantly with spelling and handwriting styles. GPT Image 1.5 also captured the requested chalk texture and cursive elements far more accurately than its competitor.

The Reversed Rodeo

Text-to-Image

“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”

FLUX.1 [schnell]
GPT Image 1.5

AI Judge Analysis

FLUX.1 [schnell]

  • + Successfully followed the difficult positional instruction 'horse on top'
  • + Clean, cinematic lighting with a surreal composition
  • + High structural integrity of the objects
  • The horse appears to have two heads or is two horses merged together
  • The astronaut's anatomy and helmet placement is confusing

GPT Image 1.5

  • + Excellent surface detail and textures on the moon and space suit
  • + Dynamic action with convincing dust and lighting effects
  • Failed the negative constraint entirely by placing the astronaut on top
  • Overly busy background with multiple planetary bodies and an asteroid

Verdict: FLUX.1 [schnell] is the clear winner for prompt adherence, as it successfully interpreted the 'horse on top' spatial instruction which is extremely difficult for most models. While GPT Image 1.5 produced a more detailed and aesthetically traditional space render, it failed to follow the core surreal premise of the prompt.

The Capybara Taxi Driver

Text-to-Image

“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”

FLUX.1 [schnell]
GPT Image 1.5

AI Judge Analysis

FLUX.1 [schnell]

  • + High resolution and clean visual style
  • + Accurate depiction of the passenger as requested
  • + Detailed capybara fur texture
  • Failed the specific instruction to have both paws on the steering wheel
  • Capybara is looking at the camera rather than the road
  • Camera angle makes the capybara look like a passenger rather than the driver (positioned in the middle/right)

GPT Image 1.5

  • + Strictly followed the instruction for both paws on the steering wheel
  • + Excellent composition that clearly shows the capybara in the driver's seat and the passenger in the back
  • + More authentic NYC taxi driver cap design and clothing detail
  • The passenger's face is slightly soft/out of focus
  • Slightly more grain/noise compared to model A

Verdict: GPT Image 1.5 is the clear winner because it accurately followed the complex spatial instructions of the prompt, specifically having both paws on the wheel and placing the businesswoman correctly in the back seat. FLUX.1 [schnell] failed to place the paws correctly and composed the shot in a way that makes the capybara look like it is sitting where the passenger should be.

The Halloween Invitation

Text-to-Image

“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”

FLUX.1 [schnell]
GPT Image 1.5

AI Judge Analysis

FLUX.1 [schnell]

  • + Clean layout with good visibility of the central elements
  • + Accurate representation of twisted trees and bats
  • Significant text hallucinations and repetition in the event details
  • The style feels more like a modern graphic design than a vintage gothic parchment

GPT Image 1.5

  • + Excellent adherence to the 'vintage gothic' aesthetic with highly detailed textures
  • + Perfect text rendering of all requested information including the date and location
  • + Superior border detail featuring thorns and spiderwebs as requested
  • The parchment texture is very busy, slightly reducing the legibility of the location text

Verdict: GPT Image 1.5 is the clear winner as it perfectly captures the desired vintage gothic aesthetic and accurately renders all text elements without the errors found in the competitor. FLUX.1 [schnell] failed significantly on text rendering, producing repetitive lines and nonsensical words in the event details section.

Isometric Miniature Diorama Scenes

Text-to-Image

“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”

FLUX.1 [schnell]
GPT Image 1.5

AI Judge Analysis

FLUX.1 [schnell]

  • + Excellent adherence to the 'minimal' requirement
  • + Very clean, high-clarity 3D aesthetic
  • + Perfectly centered and ultra-clean layout
  • Missed the word 'SUSHI' in the text overlay
  • The sushi roll looks a bit hybrid/confused rather than a standard piece

GPT Image 1.5

  • + Followed all text instructions including 'SUSHI' and the flag
  • + Rich, detailed PBR materials and textures
  • + Excellent composition for a diorama scene
  • Ignored the 'minimal' garnish and plate request by adding teapot and soy sauce
  • The scene feels a bit crowded compared to the 'ultra-clean' requirement

Verdict: GPT Image 1.5 adhered more closely to the specific text requirements, including both 'JAPAN' and 'SUSHI', and provided a much more detailed 3D miniature scene. FLUX.1 [schnell] produced a cleaner, more 'minimal' image as requested, but failed to include the word 'SUSHI'.

Adorable Baby Animals in Sunny Meadow

Text-to-Image

“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”

FLUX.1 [schnell]
GPT Image 1.5

AI Judge Analysis

FLUX.1 [schnell]

  • + Excellent fur texture and sharpness
  • + Vibrant colors and high-contrast lighting
  • + Clear, expressive eye details
  • Failed to include a bunny (substituted with a second kitten/fennec-like cat)
  • Anatomical merge where the two middle cats are touching
  • Fox kit looks slightly artificial/stylized

GPT Image 1.5

  • + Accurately included all four requested animals (dog, cat, bunny, fox)
  • + Captured the 'tumbling' and 'playful' motion better
  • + Beautiful god rays and realistic dew sparkles
  • Fox's front paws look slightly muddy or indistinct
  • Kitten's paw pads are a bit oversized and repetitive
  • Slightly more chaotic composition than Model A

Verdict: GPT Image 1.5 is the clear winner for prompt adherence as it successfully included the baby bunny which FLUX.1 [schnell] missed. While FLUX.1 [schnell] has slightly sharper fur textures, GPT Image 1.5 better captures the requested 'tumbling' action, 'god rays', and 'dew sparkles' to create a more complete realization of the prompt.

Vintage Cafe Logo

Text-to-Image

“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”

FLUX.1 [schnell]
GPT Image 1.5

AI Judge Analysis

FLUX.1 [schnell]

  • + Clean vector emblem style
  • + Well-balanced circular composition
  • Significant spelling errors in the brand name and year
  • Missing the steam element requested in the prompt

GPT Image 1.5

  • + Perfect text rendering for both brand name and date
  • + Includes the cloche dome with steam as requested
  • + Excellent use of texture and vintage shading
  • Ignored the 'light background' instruction, opting for black
  • Slightly less 'minimalist' than requested due to detailed shading

Verdict: GPT Image 1.5 is the clear winner because it correctly spells the brand name 'Caffè Florian' and the establishment year '1720', whereas FLUX.1 [schnell] produced 'Cafeé Framilan' and '7720'. While FLUX.1 adhered better to the background color request, GPT Image 1.5 captured the specific illustrative elements like the steam and the banner far more effectively.

Apollo 11: Journey to Tranquility

Text-to-Image

“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”

FLUX.1 [schnell]
GPT Image 1.5

AI Judge Analysis

FLUX.1 [schnell]

  • + Excellent adherence to the 'flat-vector' style and color palette.
  • + Clean and modern aesthetic with professional-looking iconography.
  • + Good use of negative space and balance in the layout.
  • Text is mostly gibberish and suffers from typical AI generation artifacts.
  • The step-by-step logic and specific icons requested are muddled into a single abstract graphic.
  • The rocket icon looks more like a generic stylized jet than a Saturn V.

GPT Image 1.5

  • + Perfectly follows all 6 specific steps requested in the prompt.
  • + Outstanding text legibility and accuracy, including names and mission stages.
  • + Captures the NASA-inspired aesthetic with much more identifiable iconography (Saturn V, Lunar Module).
  • The layout is a bit crowded compared to Model A's minimalist approach.
  • Small inconsistencies in lunar orbit/earth orbit graphic logic.

Verdict: GPT Image 1.5 is the clear winner as it precisely followed the complex logical structure of the prompt, including all six specific infographic steps with legible, accurate text. While FLUX.1 [schnell] captures a highly stylish aesthetic, its failure to render the requested steps and its reliance on gibberish text makes it less effective as an actual infographic.

Next steps

Explore each model