Head to head
Esc

Models · slot A

to navigate to pick

FLUX.1 [schnell] FP8 Black Forest Labs GPT Image 1 OpenAI

Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.

FLUX.1 [schnell] FP8

19.0 arena score

#47 of 62 in Text-to-Image

Skill signature · Text-to-Image

GPT Image 1

22.6 arena score

#32 of 62 in Text-to-Image

Vote tally

Where the votes landed

FLUX.1 [schnell] FP8

100.0%

win rate

Ties

0.0%

GPT Image 1

0.0%

win rate

100.0% 0.0% ties 0.0%
Shared challenges 13

Challenge by challenge

The strongest take from each model on every shared challenge, with the AI judge's read.

Geometric Composition

Text-to-Image

“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”

FLUX.1 [schnell] FP8
GPT Image 1

AI Judge Analysis

FLUX.1 [schnell] FP8

  • + Excellent handling of light refraction and reflections on the wood surface.
  • + Realistic blue globe with intricate glass textures.
  • + Follows the lighting direction prompt perfectly with bright window light.
  • The 'cube' is a tall rectangular prism rather than a cube.
  • The internal shelf is an hallucinated element not mentioned in the prompt.

GPT Image 1

  • + Accurately depicts a cubic shape for the glass container.
  • + Highly realistic texture for the matte blue sphere and the paper edges of the book.
  • + Shows the plant clearly behind the glass as requested.
  • The sphere appears to be floating inside the cube without touching the bottom surface.
  • The lighting is a bit more muted than the 'window light' description suggests compared to Image A.

Verdict: Both models followed the prompt instructions very well, including the specific spatial relations of the objects. GPT Image 1 is the winner because it actually rendered a cube and correctly placed the plant behind the glass, whereas FLUX.1 [schnell] FP8 rendered a tall rectangle and added an unnecessary internal shelf.

Candid Street Photography

Text-to-Image

“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”

FLUX.1 [schnell] FP8
GPT Image 1

AI Judge Analysis

FLUX.1 [schnell] FP8

  • + Excellent handling of wet pavement reflections and light artifacts
  • + Detailed mechanical rendering of the bicycle frame
  • + Good sense of environmental depth in an urban setting
  • The man appears to be posing with or leaning on the bike rather than 'repairing' it
  • Fails to significantly incorporate the requested motion blur for passing cars
  • The subject's ethnicity is somewhat ambiguous

GPT Image 1

  • + Strong adherence to the 'repairing' action with a realistic crouching pose
  • + Excellent skin texture and authentic elderly Japanese features
  • + Successfully captures the moody, desaturated color palette and cinematic atmosphere
  • The bicycle geometry becomes physically impossible/muddled around the rear wheel hub
  • Missing clear 'motion blur' on the cars in the background as requested

Verdict: GPT Image 1 captures the specific 'repairing' action much more effectively than FLUX.1 [schnell] FP8, which shows the man simply holding the handlebars. While FLUX.1 [schnell] FP8 has cleaner bicycle rendering and better reflections, GPT Image 1 feels more like a candid street photo with superior character work and skin texture, making it the better interpretation of the prompt's intent.

Fantasy Warrior

Text-to-Image

“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”

FLUX.1 [schnell] FP8
GPT Image 1

AI Judge Analysis

FLUX.1 [schnell] FP8

  • + Extremely sharp and detailed skin textures
  • + Intense and dramatic lighting
  • + Lifelike iris detail and coloring
  • Hair lacks the requested braided style with beads
  • The facial expression feels slightly over-dramatized/caricatured
  • Subtle text artifacts visible on the bottom right armor section

GPT Image 1

  • + Accurately depicts hair braided with small beads
  • + Ornate engraved plate armor is highly detailed and realistic
  • + Captures the battle-worn aesthetic with believable dirt and scars
  • Lighting is a bit darker, obscuring some details
  • The skin texture is slightly softer compared to the other image

Verdict: GPT Image 1 followed the specific prompt instructions much more closely, particularly regarding the hair braids and beads which FLUX.1 [schnell] FP8 missed entirely. While FLUX.1 [schnell] FP8 has higher contrast and skin clarity, GPT Image 1 creates a more authentic 'battle-worn paladin' look with better armor engraving and prompt adherence.

Modern Clean Menu

Text-to-Image

“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”

FLUX.1 [schnell] FP8
GPT Image 1

AI Judge Analysis

FLUX.1 [schnell] FP8

  • + Comprehensive multi-section layout with many items
  • + Great use of a grid for multiple food photos
  • + Captures the professional, spread-out feel of a real physical menu
  • Numerous spelling errors in headings like 'APPTIZERS' and 'ORCETERS'
  • Food photography is small and slightly repetitive in style

GPT Image 1

  • + High-quality food photography with vibrant colors
  • + Clean, bold sans-serif typography that is very legible
  • + Modern minimalist aesthetic with effective colored accents
  • Failed to include all requested sections (missing Mains)
  • Contains placeholder text like 'Apperoiation descrigion'

Verdict: Both models followed the prompt well, but they took different approaches. FLUX.1 [schnell] FP8 created a more expansive, realistic menu layout with multiple columns and categories, though its text rendering is poor. GPT Image 1 produced much higher quality food imagery and cleaner typography that feels more modern, making it more aesthetically pleasing overall despite having fewer menu items.

Magic Burger Explosion: Fiery Photorealism Challenge

Text-to-Image

“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”

FLUX.1 [schnell] FP8
GPT Image 1

AI Judge Analysis

FLUX.1 [schnell] FP8

  • + Excellent photorealistic rendering of the burger textures.
  • + Good use of actual fire in the background for atmospheric depth.
  • + Dynamic 'exploded' feeling with small debris and sauce droplets.
  • Significant text errors including 'LIIMITED' and 'NEEY'.
  • Incorrect price of €69 displayed prominently.
  • The primary burger is mostly assembled rather than fully exploded components.

GPT Image 1

  • + Perfect text rendering for all requested phrases including the fiery effect.
  • + Full adherence to the 'exploded' component layout showing individual layers.
  • + Highly cohesive color palette with consistent glowing embers.
  • The price is rendered as '€.99' instead of '€6.99'.
  • The bun texture is slightly less realistic than Model A.
  • Lighting on the burger is somewhat flat compared to the intensity of the background.

Verdict: GPT Image 1 is the superior choice because it accurately follows the structural prompt requirements of an exploded burger and renders the requested text effects with very few errors. While FLUX.1 [schnell] FP8 has higher surface realism on the ingredients, its repeated spelling failures and incorrect pricing make it unusable as an advertisement.

Chalkboard Menu

Text-to-Image

“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”

FLUX.1 [schnell] FP8
GPT Image 1
100% wins 0% ties 0% wins

AI Judge Analysis

FLUX.1 [schnell] FP8

  • + Excellent chalk-like texture with smudges
  • + Realistic wood frame and cafe environment context
  • + Captures the casual, messy nature of handwriting
  • Terrible text accuracy with many misspellings like 'Risortto' and 'Octemon'
  • Fails to render the elegant cursive requested for the title
  • Repeats menu items multiple times in a confusing layout

GPT Image 1

  • + Excellent text accuracy with no spelling errors
  • + Consistent and clear chalk texture across all letters
  • + Followed the specific menu item list and pricing accurately
  • Failed to provide 'elegant cursive' for the title
  • The handwriting looks a bit too uniform/digital rather than having 'natural variations'
  • The layout is very centered and static compared to a natural chalkboard

Verdict: GPT Image 1 is the clear winner due to its superior text rendering capabilities, accurately spelling every complex menu item requested. While FLUX.1 [schnell] FP8 captures a more authentic 'messy' chalkboard aesthetic, it fails significantly on spelling and composition, producing incoherent repetitions of the text.

The Reversed Rodeo

Text-to-Image

“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”

FLUX.1 [schnell] FP8
GPT Image 1

AI Judge Analysis

FLUX.1 [schnell] FP8

  • + Successfully interpreted the surreal instruction of placing the horse on top of the astronaut equipment.
  • + Cinematic lighting with a beautiful Earth backdrop.
  • + High resolution and clear textures on the horse's fur and the mechanical gear.
  • Anatomy is a bit confused with two horse heads/bodies merged into the scene.
  • The 'astronaut' is represented as a piece of equipment rather than a person in a suit.

GPT Image 1

  • + Excellent lighting and texture on the horse and space suit.
  • + Great composition that feels very cinematic and illustrative.
  • + Clear and coherent rendering of a horse and astronaut.
  • Failed the core prompt instruction to have the 'horse on top' of the astronaut.
  • Visual interpretation is literal and lacks the requested surrealism.

Verdict: While GPT Image 1 is a beautiful and well-composed image, it completely ignored the specific inversion instruction in the prompt. FLUX.1 [schnell] FP8 followed the challenging prompt logic by placing the horse on top of the astronaut's gear, resulting in a much more surreal and accurate response to the text instructions.

The Capybara Taxi Driver

Text-to-Image

“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”

FLUX.1 [schnell] FP8
GPT Image 1

AI Judge Analysis

FLUX.1 [schnell] FP8

  • + Excellent clarity and lighting contrast.
  • + Accurately depicts the taxi interior and the 'bored' expression of the passenger.
  • + High detail in the capybara's fur and clothing.
  • The passenger is holding two phones, which is a logic error.
  • The capybara's hat is a simple plastic-looking helmet rather than a traditional driver cap.

GPT Image 1

  • + Features a more authentic taxi driver cap with a brim.
  • + The passenger's bored expression perfectly matches the prompt's intent for normalcy.
  • + Better depth of field and more realistic night street bokeh through the rear window.
  • The passenger is slightly out of focus compared to Image A.
  • The steering wheel placement is a bit low and crowded near the capybara's chest.

Verdict: Both models followed the complex prompt very well, but GPT Image 1 feels more grounded and cinematic due to its superior cap design and natural lighting. While FLUX.1 [schnell] FP8 is sharper and has more vibrant colors, it suffered from a logical error where the passenger is holding two separate phones.

The Halloween Invitation

Text-to-Image

“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”

FLUX.1 [schnell] FP8
GPT Image 1

AI Judge Analysis

FLUX.1 [schnell] FP8

  • + Features a vibrant, glowing jack-o-lantern central to the composition.
  • + Includes stylized border flourishes that frame the image well.
  • Numerous spelling and grammar errors in the text like 'Tine', 'nigh tof friights', and 'Theaches'.
  • The text layout is messy with random colons and fragmented sentences.
  • Fails to correctly list the requested date and location clearly.

GPT Image 1

  • + Excellent typography with a perfect 'Halloween Party Invitation' title and banner text.
  • + Superior adherence to the vintage gothic aesthetic with a dark, moody parchment texture.
  • + High-quality details such as the spider webs in the corners and the subtle moon and bat silhouettes.
  • Combined the Time and Location fields into one line ('TIME: The Arches, NYC'), missing the '7pm' detail.

Verdict: GPT Image 1 is the clear winner as it successfully captures the 'vintage gothic' aesthetic with highly accurate and elegant typography, whereas FLUX.1 [schnell] FP8 struggles significantly with spelling ('A a nigh tof friights') and logical text placement. GPT Image 1 also follows the stylistic cues for the border, trees, and parchment texture much more effectively.

Isometric Miniature Diorama Scenes

Text-to-Image

“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”

FLUX.1 [schnell] FP8
GPT Image 1

AI Judge Analysis

FLUX.1 [schnell] FP8

  • + Excellent 3D rendering with soft shadows and realistic textures
  • + Accurate miniature diorama aesthetic
  • + High clarity and clean composition
  • Text is incorrect and fails to generate 'SUSHI'
  • Includes a weird second 'JAPAN' word with a dot instead of the requested layout

GPT Image 1

  • + Perfect text adherence with both 'JAPAN' and 'SUSHI' and a flag icon
  • + Accurate isometric diorama perspective and lighting
  • + Clean, high-quality cartoon style with pleasant textures
  • Background color is slightly darker than 'light blue' requested
  • The rice texture is a bit large for the scale

Verdict: While FLUX.1 [schnell] FP8 offers a more sophisticated 3D render with better subsurface scattering on the sushi, it fails the text prompt completely, rendering garbled words. GPT Image 1 (DALL-E 3) followed every instruction, including the specific text and layout requirements, resulting in a much more useful and accurate image.

Adorable Baby Animals in Sunny Meadow

Text-to-Image

“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”

FLUX.1 [schnell] FP8
GPT Image 1

AI Judge Analysis

FLUX.1 [schnell] FP8

  • + Includes a high number of animals and butterflies
  • + Vibrant golden lighting and colors
  • Failed to include a rabbit, replacing it with extra kittens/unknown kits
  • Anatomical glitches including a floating paw and a fifth animal head emerging from the fox
  • Looks more like a digital illustration than a photorealistic scene

GPT Image 1

  • + Perfect adherence to the prompt, including all four specific animals requested
  • + Excellent movement and dynamic composition that captures the 'chasing' aspect
  • + Realistic lighting, fur textures, and believable depth of field
  • One butterfly wing is slightly disconnected from its body
  • Fox paws are a bit dark and muddy in texture

Verdict: GPT Image 1 is the clear winner as it successfully included all four requested animals (Golden Retriever, kitten, bunny, and fox), whereas FLUX.1 [schnell] FP8 missed the bunny and produced several anatomical errors like a floating paw and a merged fifth head. GPT Image 1 also better captured the 'playfully chasing' action described in the prompt with a more realistic, photographic style.

Vintage Cafe Logo

Text-to-Image

“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”

FLUX.1 [schnell] FP8
GPT Image 1

AI Judge Analysis

FLUX.1 [schnell] FP8

  • + Successfully captures the requested warm brown and cream tones
  • + Includes a decorative banner and light background as requested
  • Significant spelling error in the primary brand name
  • The dome looks more like a building cupola than a food cloche
  • Composition is cluttered with overlapping text

GPT Image 1

  • + Perfect text rendering of the name and date
  • + Accurate depiction of a minimalist retro cloche dome with steam
  • + Excellent minimalist aesthetic and clean vector style
  • Ignored the 'light background' instruction, opting for black
  • Used a single brown tone instead of a brown and cream color scheme

Verdict: GPT Image 1 is the clear winner because it correctly follows the primary branding instructions and spells the name 'Caffè Florian' perfectly, whereas FLUX.1 [schnell] FP8 produced a severe misspelling ('AFE FLAMILAN'). While GPT Image 1 missed the light background requirement, its superior minimalist design and typography make it a more usable logo.

Apollo 11: Journey to Tranquility

Text-to-Image

“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”

FLUX.1 [schnell] FP8
GPT Image 1

AI Judge Analysis

FLUX.1 [schnell] FP8

  • + Strong composition that feels like a professional corporate infographic
  • + Excellent use of a dark navy palette with clean, sharp vector lines
  • + Attempts to include almost all requested complex text and numbered steps
  • Text is mostly gibberish with frequent spelling errors
  • Icons are abstract and don't clearly represent the Saturn V or specific mission stages

GPT Image 1

  • + Text is highly legible and correctly identifies the crew members
  • + Icons for Earth and the Lunar Module are very clear and recognizable
  • + Excellent flat vector illustration style with a cohesive color palette
  • Fails to follow the requested 6-step linear structure, combining/missing some phases
  • Layout feels a bit cluttered compared to a traditional infographic timeline
  • Contains some typos like 'EARLLUNAR'

Verdict: FLUX.1 [schnell] FP8 creates a more realistic infographic layout with a superior professional design, but its text is largely unintelligible. GPT Image 1 follows the stylistic instructions well and provides much better legibility and icon clarity, though it struggles with the specific 6-step sequential structure. GPT Image 1 is the likely winner for its ability to convey actual information and follow the NASA-themed iconography more accurately.

Next steps

Explore each model