Head to head
Esc

Models · slot A

to navigate to pick

FLUX.1 [schnell] FP8 Black Forest Labs GPT Image 2 OpenAI

Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.

FLUX.1 [schnell] FP8

19.0 arena score

#47 of 62 in Text-to-Image

Skill signature · Text-to-Image

GPT Image 2

27.7 arena score

#4 of 62 in Text-to-Image

Vote tally

Where the votes landed

FLUX.1 [schnell] FP8

0.0%

win rate

Ties

0.0%

GPT Image 2

100.0%

win rate

0.0% 0.0% ties 100.0%
Shared challenges 13

Challenge by challenge

The strongest take from each model on every shared challenge, with the AI judge's read.

Geometric Composition

Text-to-Image

“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”

FLUX.1 [schnell] FP8
GPT Image 2

AI Judge Analysis

FLUX.1 [schnell] FP8

  • + Excellent depiction of light refraction through the glass.
  • + Vibrant colors and high-quality rendering of the glass materials.
  • The cube shape is slightly elongated and contains an illogical internal shelf.
  • The sphere inside appears to be a globe or marbled rather than a simple blue sphere.

GPT Image 2

  • + Perfect adherence to the geometry of a cube and the placement of objects.
  • + Realistic texture on the red book and wooden table.
  • + Accurately depicts the plant visible through the transparent glass back.
  • The lighting is a bit flat compared to Model A.
  • Minor shadow inconsistency where the sphere touches the bottom glass.

Verdict: GPT Image 2 followed the prompt's spatial instructions much more accurately, creating a proper cube with the sphere resting on the bottom as expected. FLUX.1 [schnell] FP8 produced a more visually striking image with beautiful light effects, but it failed the geometric requirement by adding a shelf and creating a rectangular prism instead of a cube.

Candid Street Photography

Text-to-Image

“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”

FLUX.1 [schnell] FP8
GPT Image 2

AI Judge Analysis

FLUX.1 [schnell] FP8

  • + Excellent cinematic lighting and wet pavement reflections.
  • + Accurately depicts the requested red bicycle.
  • + Good bokeh effect and shallow depth of field.
  • The cars in the background are stationary despite the request for motion blur.
  • The man's ethnicity is somewhat ambiguous.
  • Lack of visible rain drops or rain protection on the subject.

GPT Image 2

  • + Successfully incorporates motion blur in the background traffic.
  • + The subject feels more authentic to a Japanese street setting with realistic skin texture.
  • + Strong adherence to the 'repairing' action with a toolbox and focused posture.
  • The bike frame construction is physically nonsensical near the rear wheel.
  • The shallow depth of field is less pronounced than requested.
  • Composition is a bit cluttered with the white post in the foreground.

Verdict: GPT Image 2 captured the 'candid' and 'repair' aspects of the prompt more effectively, successfully including the motion blur of passing cars which FLUX.1 [schnell] FP8 missed. While FLUX.1 [schnell] FP8 had superior lighting and reflections, GPT Image 2 felt more realistic and authentic to the location and requested actions.

Fantasy Warrior

Text-to-Image

“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”

FLUX.1 [schnell] FP8
GPT Image 2

AI Judge Analysis

FLUX.1 [schnell] FP8

  • + Strong, dramatic lighting with clear orange torchlight reflections
  • + Hyper-detailed iris and skin texture around the eyes
  • + Intense, battle-worn expression that fits the paladin archetype
  • The 'braided hair' request is poorly executed, appearing as thin loose strands rather than braids
  • Skin looks slightly processed or overly smoothed in some highlight areas

GPT Image 2

  • + Excellent adherence to hair braids with bead details
  • + Very realistic skin texture showing dirt and subtle scars as requested
  • + Detailed engraving on the plate armor and visible leather straps
  • Lighting is somewhat muted compared to the 'warm torchlight' request
  • Slightly less 'close' as a portrait than Model A

Verdict: GPT Image 2 followed the technical details of the prompt much more accurately, specifically regarding the braids, beads, and the 'battle-worn' texture of the skin. While FLUX.1 [schnell] FP8 captured a more dramatic lighting effect, its failure to render the braided hair and its slightly over-saturated look makes GPT Image 2 the more successful interpretation.

Modern Clean Menu

Text-to-Image

“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”

FLUX.1 [schnell] FP8
GPT Image 2

AI Judge Analysis

FLUX.1 [schnell] FP8

  • + Organized layout with clear columns
  • + Matches the requested sections for appetizers, pizza, and mains
  • Garbled text and nonsensical headings like 'ORCETERS' and 'SECCER'
  • Visual quality of food photos is low and repetitive
  • Layout feels cramped and dated rather than modern

GPT Image 2

  • + Exceptional text rendering with coherent menu items and descriptions
  • + High-quality, appetizing photography in a clean grid layout
  • + Follows all design constraints including bold sans-serif fonts and vibrant accents
  • None identified

Verdict: GPT Image 2 is the clear winner as it produces a professional-grade, functional menu design with perfectly legible text and high-quality food photography. FLUX.1 [schnell] FP8 fails on basic text generation and provides low-fidelity imagery that lacks the modern polish requested in the prompt.

Magic Burger Explosion: Fiery Photorealism Challenge

Text-to-Image

“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”

FLUX.1 [schnell] FP8
GPT Image 2

AI Judge Analysis

FLUX.1 [schnell] FP8

  • + Clean layout for the main title text
  • + Vibrant colors in the food elements
  • Significant text errors including 'LIIMITED' and 'NEEY'
  • Failed to apply the fiery glowing effect to the text as requested
  • Incorrect price rendering and redundant text placement

GPT Image 2

  • + Perfect text adherence with requested fiery, glowing effects
  • + Better 'exploded' burger composition showing all distinct layers
  • + Excellent integration of the starburst and pricing
  • The fiery background and embers are quite busy, reducing some contrast

Verdict: GPT Image 2 followed all prompt instructions perfectly, including the complex text rendering with fiery effects and the specific price format. In contrast, FLUX.1 [schnell] FP8 struggled with spelling, failed to apply the requested text styling, and missed the mark on the price and starburst integration.

Chalkboard Menu

Text-to-Image

“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”

FLUX.1 [schnell] FP8
GPT Image 2

AI Judge Analysis

FLUX.1 [schnell] FP8

  • + Very high image clarity and sharpness.
  • + Realistic lighting and wood texture on the frame.
  • Severely failed text rendering with repeated words and nonsensical phrases.
  • Incorrect prices and layout compared to the prompt.
  • Failed to render cursive for the title as requested.

GPT Image 2

  • + Perfect text rendering of all requested items and dates.
  • + Excellent adherence to the 'chalk texture' and 'cursive title' instructions.
  • + Consistent and realistic handwriting style throughout the board.
  • Slightly lower overall resolution/sharpness compared to the other model.
  • The lighting is a bit dim in the lower corners.

Verdict: GPT Image 2 is the clear winner as it followed every text and styling instruction perfectly, including complex menu items and specific dates. FLUX.1 [schnell] FP8 struggled significantly with the text, producing repetitive gibberish and failing to follow the price or cursive requirements.

The Reversed Rodeo

Text-to-Image

“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”

FLUX.1 [schnell] FP8
GPT Image 2
0% wins 0% ties 100% wins

AI Judge Analysis

FLUX.1 [schnell] FP8

  • + Excellent cinematic lighting and color palette.
  • + High visual clarity and artistic composition with the Earth background.
  • Fails the prompt requirement of a horse riding an astronaut.
  • Anatomical horror with a horse head emerging from another horse's back.

GPT Image 2

  • + Perfect adherence to the complex prompt instruction of the horse on top.
  • + Incredibly high detail on the space suit, NASA logo, and lunar surface.
  • + Accurate reflection in the astronaut's visor.
  • Minor leather strap clipping through the horse's front leg.

Verdict: GPT Image 2 followed the specific and difficult 'horse on top' instruction perfectly, showing a horse literally riding an astronaut on all fours. FLUX.1 [schnell] FP8 failed the core concept of the prompt, instead generating a surreal two-headed horse that ignores the 'riding astronaut' requirement entirely.

The Capybara Taxi Driver

Text-to-Image

“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”

FLUX.1 [schnell] FP8
GPT Image 2

AI Judge Analysis

FLUX.1 [schnell] FP8

  • + Excellent fur texture and lighting on the capybara.
  • + Clear, high-quality rendering of both the passenger and the background street lights.
  • The passenger is holding two phones simultaneously, which is a significant anatomical/logical error.
  • The 'TAXI' hat looks like a toy prop rather than a professional driver's cap.

GPT Image 2

  • + Perfect adherence to the 'bored expression' and 'businesswoman in a coat' prompt for the passenger.
  • + The taxi driver's cap is more realistic and detailed.
  • + Excellent photorealistic composition with natural-looking motion blur and bokeh.
  • The passenger's face is slightly out of focus compared to the foreground.
  • The capybara's paws on the wheel are slightly ambiguous in shape.

Verdict: GPT Image 2 is the superior image as it captures the specific tone of the prompt more effectively, particularly the bored expression of the businesswoman. While FLUX.1 [schnell] FP8 has higher sharpness, it suffers from a major logical error where the passenger is holding two different smartphones.

The Halloween Invitation

Text-to-Image

“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”

FLUX.1 [schnell] FP8
GPT Image 2

AI Judge Analysis

FLUX.1 [schnell] FP8

  • + Strong cinematic lighting with a vibrant center glow
  • + Clean layout for the main title text
  • Numerous spelling errors in the banner and specific event details
  • The border feels more abstract and lacks the requested web/thorn detail
  • Redundant and confusing time information provided

GPT Image 2

  • + Perfect text rendering for all requested titles and event details
  • + Intricate gothic border involving webs, thorns, and skulls
  • + Highly detailed background containing the bridges (arches) and NYC skyline
  • The parchment texture is very busy, slightly obscuring smaller details
  • The pumpkin light is a bit flatter compared to the glow effect in Model A

Verdict: GPT Image 2 is the clear winner for its superior text accuracy, following every specific instruction for dates and locations without spelling errors. It also captures the 'Vintage Gothic' aesthetic much better with its detailed thorns, webs, and thematic background, whereas FLUX.1 [schnell] FP8 struggled significantly with the spelling and the complexity of the border.

Isometric Miniature Diorama Scenes

Text-to-Image

“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”

FLUX.1 [schnell] FP8
GPT Image 2

AI Judge Analysis

FLUX.1 [schnell] FP8

  • + Excellent soft lighting and clean aesthetic
  • + Accurate 3D isometric diorama perspective
  • + Minimalist, professional composition
  • Failed text rendering completely, repeating 'JAPAN' and misspelling it as 'JA-AN'
  • The flag icon is integrated into the misspelled text rather than being separate
  • Lacks the word 'SUSHI' entirely

GPT Image 2

  • + Perfect text rendering for both 'JAPAN' and 'SUSHI'
  • + Includes a correct and separate flag icon as requested
  • + Detailed, high-quality PBR material textures for the sushi and environment
  • Composition is slightly more cluttered than 'minimal garnish' requested
  • The base includes extra environmental elements like a stone lantern not explicitly requested

Verdict: GPT Image 2 is the clear winner as it successfully followed the complex text instructions and icon placement. While FLUX.1 [schnell] FP8 captured the 'clean' and 'soft texture' aesthetic well, it failed significantly on the typography, which was a core part of the prompt.

Adorable Baby Animals in Sunny Meadow

Text-to-Image

“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”

FLUX.1 [schnell] FP8
GPT Image 2

AI Judge Analysis

FLUX.1 [schnell] FP8

  • + Excellent catchlights and expressive eyes.
  • + Rich, vibrant color palette with strong backlighting.
  • + High degree of 'fluffiness' in textures.
  • Failed to include a rabbit in the group.
  • Anatomical issues including a kitten with three ears/extra limbs and merged animal bodies.
  • Generated two kittens instead of following the variety list.

GPT Image 2

  • + Perfect adherence to the prompt, including all four specific animals.
  • + Dynamic 'chasing' and 'tumbling' poses that feel more active.
  • + Realistic lighting with distinct god rays and dew-like highlights.
  • The fox kit has slightly muddy textures on its face.
  • The kitten's tail is a bit thick, resembling a squirrel more than a cat.

Verdict: GPT Image 2 is the superior image because it successfully included all four requested animals (dog, cat, rabbit, and fox) with distinct, consistent anatomy. FLUX.1 [schnell] FP8 failed to include the rabbit and suffered from significant anatomical errors, including extra ears and fused limbs between the various animals.

Vintage Cafe Logo

Text-to-Image

“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”

FLUX.1 [schnell] FP8
GPT Image 2

AI Judge Analysis

FLUX.1 [schnell] FP8

  • + Successfully captured the requested warm brown and cream color scheme
  • + Included a vector-style emblem as requested
  • Major spelling errors in the brand name ('AFe FLAMILAN')
  • The central icon is a mosque-like dome rather than a food service cloche dome
  • Text layout is cluttered and poorly aligned

GPT Image 2

  • + Perfect text rendering for both 'Caffè Florian' and 'Est. 1720'
  • + Accurate representation of a cloche dome with stylized steam
  • + Excellent vintage texture and professional logo composition
  • Slightly more complex than 'minimalist' might suggest, though fits the 'vintage' style well

Verdict: GPT Image 2 followed the prompt with near-perfection, accurately depicting a restaurant cloche dome and rendering all text correctly. FLUX.1 [schnell] FP8 failed significantly on the typography and misinterpreted the 'cloche dome' as architectural, resulting in a confusing and misspelled logo.

Apollo 11: Journey to Tranquility

Text-to-Image

“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”

FLUX.1 [schnell] FP8
GPT Image 2

AI Judge Analysis

FLUX.1 [schnell] FP8

  • + Adheres moderately well to the flat-vector, minimalist aesthetic requested.
  • + Uses the specified navy, white, and red color palette effectively.
  • Text is largely illegible gibberish throughout.
  • The layout is disorganized with icons that don't clearly represent the Saturn V or the specific mission steps.
  • The horizontal timeline fails to complete the 6-step sequence meaningfully.

GPT Image 2

  • + Perfect adherence to all 6 specific mission steps with accurate illustrations for each.
  • + Excellent text rendering of titles and names like Armstrong, Aldrin, and Collins.
  • + Sophisticated composition that combines modern infographic design with the requested NASA-inspired palette.
  • Slightly more detailed than a 'flat-vector' style, leaning into 3D-shaded illustrations.
  • The landing site pin says 'Tranquility' but the text within the circle is a bit compressed.

Verdict: GPT Image 2 followed the complex instructions perfectly, including the specific 6-step timeline and providing highly legible and accurate text for the Apollo 11 mission. FLUX.1 [schnell] FP8 struggled with the sequence and text, producing a generic layout with nonsensical labels and less recognizable icons.

Next steps

Explore each model