Head to head
Esc

Models · slot A

to navigate to pick

GPT Image 1.5 OpenAI Stable Diffusion 3.5 Large Stability AI

Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.

GPT Image 1.5

27.1 arena score

#7 of 62 in Text-to-Image

Top 3 in Image Editing
Skill signature · Text-to-Image

Stable Diffusion 3.5 Large

22.8 arena score

#30 of 62 in Text-to-Image

Vote tally

Where the votes landed

GPT Image 1.5

69.2%

win rate

Ties

0.0%

Stable Diffusion 3.5 Large

30.8%

win rate

69.2% 0.0% ties 30.8%
Shared challenges 13

Challenge by challenge

The strongest take from each model on every shared challenge, with the AI judge's read.

Geometric Composition

Text-to-Image

“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”

GPT Image 1.5
Stable Diffusion 3.5 Large
29% wins 0% ties 71% wins

AI Judge Analysis

GPT Image 1.5

  • + Perfect adherence to the spatial arrangement requested.
  • + Highly realistic glass reflections and refractions of the background plant.
  • + Excellent material textures, especially on the book's canvas cover and the wooden table.
  • The blue sphere is relatively large compared to the prompt's 'small blue sphere'.

Stable Diffusion 3.5 Large

  • + Good lighting effects and sharp focus on the glass cube.
  • + Accurate 'small' scale for the blue sphere.
  • Incorrect object placement; the book is inside the cube rather than sitting on top of it.
  • Coherence issues where the glass cube seems to clip through the red book.
  • The plant is mostly to the side/front rather than behind the cube as requested.

Verdict: GPT Image 1.5 followed the complex spatial instructions perfectly, correctly placing the book on top of the cube and the plant behind it. Stable Diffusion 3.5 Large struggled with the spatial logic, placing the book inside the cube and failing to clearly place the plant behind the glass, which was a core element of the prompt.

Candid Street Photography

Text-to-Image

“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”

GPT Image 1.5
Stable Diffusion 3.5 Large
100% wins 0% ties 0% wins

AI Judge Analysis

GPT Image 1.5

  • + Excellent skin and fabric textures that look realistic.
  • + Strong adherence to 'imperfect framing' with a tight, candid composition.
  • + Detailed mechanical parts on the bicycle and repair tools.
  • The car in the background lacks the requested motion blur.
  • The raindrops appear as static white dots rather than falling streaks.

Stable Diffusion 3.5 Large

  • + Successfully captured motion blur on the background vehicles.
  • + Good representation of falling rain and wet pavement reflections.
  • + Accurate adherence to 'shallow depth of field' with a soft background.
  • Anatomical issues with the man's hands and arms which appear distorted.
  • The bicycle's structure is physically impossible with floating and merging parts.
  • The man appears slightly 'pasted' into the scene with mismatched lighting.

Verdict: GPT Image 1.5 is the superior image due to its high level of photorealism, particularly in the subject's skin texture and the mechanical detail of the bike. While Stable Diffusion 3.5 Large followed the 'motion blur' prompt better, it failed significantly on structural coherence, producing mangled hands and an impossible bicycle frame.

Fantasy Warrior

Text-to-Image

“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”

GPT Image 1.5
Stable Diffusion 3.5 Large
100% wins 0% ties 0% wins

AI Judge Analysis

GPT Image 1.5

  • + Excellent adherence to all prompt details including beads in hair and leather straps.
  • + Superior lighting effects with warm torchlight and realistic bokeh sparks.
  • + Highly detailed facial texture with convincing scars and dirt.
  • The composition is a bit tight on the forehead.

Stable Diffusion 3.5 Large

  • + Beautifully detailed ornate engraving on the plate armor.
  • + Strong character expression and clear facial features.
  • + Good interpretation of the braided hair requirement.
  • Missed the 'small beads' in the hair mentioned in the prompt.
  • The lighting feels more like daylight than the requested warm torchlight.
  • Lacks the specific bokeh sparks requested.

Verdict: GPT Image 1.5 is the clear winner as it followed every specific detail of the prompt, including the beads in the hair, the leather straps, and the specific warm torchlight atmosphere with bokeh sparks. Stable Diffusion 3.5 Large produced a high-quality image with impressive armor engraving, but it failed to include the beads and the lighting felt too cool and diffused compared to the torchlight requested.

Modern Clean Menu

Text-to-Image

“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”

GPT Image 1.5
Stable Diffusion 3.5 Large
100% wins 0% ties 0% wins

AI judge analyzing...

Magic Burger Explosion: Fiery Photorealism Challenge

Text-to-Image

“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”

GPT Image 1.5
Stable Diffusion 3.5 Large

AI Judge Analysis

GPT Image 1.5

  • + Excellent text rendering with all requested strings accurate.
  • + Perfect adherence to the 'exploded' layout requested.
  • + High photorealistic texture detail on the patty and bun.
  • The lighting is very intense, bordering on busy.

Stable Diffusion 3.5 Large

  • + Great visual quality on the food textures.
  • + Dynamic flame effects around the base of the burger.
  • Failed to include any of the requested text.
  • The burger is not 'exploded' as requested, but mostly assembled.
  • Missing the required starburst element.

Verdict: GPT Image 1.5 followed every instruction in the prompt, including the complex text integration and the 'exploded' view of the ingredients. Stable Diffusion 3.5 Large failed to include any text or the starburst, and only depicted a whole burger hovering over a fire rather than the requested exploded composition.

Chalkboard Menu

Text-to-Image

“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”

GPT Image 1.5
Stable Diffusion 3.5 Large

AI Judge Analysis

GPT Image 1.5

  • + Excellent text rendering with no spelling errors
  • + Highly realistic chalk texture and natural handwriting style
  • + Perfect adherence to the specific date and menu items provided

Stable Diffusion 3.5 Large

  • + Provides a wider architectural context of a cozy café environment
  • + Good composition with greenery and lighting
  • + Captures the general aesthetic requested
  • Numerous spelling errors including 'TODAAY' and '2024' instead of '2026'
  • Text is cluttered and doesn't follow the specific menu item list correctly
  • Uses a mix of styles that look more like digital fonts than natural handwriting

Verdict: GPT Image 1.5 followed the prompt instructions near-perfectly, delivering clean, legible, and correctly spelled text with an authentic chalk texture. In contrast, Stable Diffusion 3.5 Large struggled significantly with the text, introducing numerous typos and failing to render the requested date and menu names accurately.

The Reversed Rodeo

Text-to-Image

“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”

GPT Image 1.5
Stable Diffusion 3.5 Large

AI Judge Analysis

GPT Image 1.5

  • + Excellent high-contrast textures on the horse and suit
  • + Solid compositional layout with multiple space elements
  • + Sharp rendering of fine details like the lunar lander and asteroid dust
  • Failed the negative constraint to put the horse on top
  • The horse's front-right leg has an anatomical glitch at the joint

Stable Diffusion 3.5 Large

  • + Successfully followed the difficult spatial constraint of putting the horse on top of the astronaut
  • + Ethereal, dreamlike color palette and lighting
  • + Good sense of motion and scale against the planet's curve
  • The horse's head and neck anatomy are slightly distorted
  • Lower overall sharpness compared to the other model

Verdict: While GPT Image 1.5 produced a much higher quality image in terms of textures and sharpness, it completely ignored the specific spatial instruction for the horse to be on top. Stable Diffusion 3.5 Large correctly interpreted the 'horse on top' request, resulting in a much more surreal and accurate response to the prompt.

The Capybara Taxi Driver

Text-to-Image

“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”

GPT Image 1.5
Stable Diffusion 3.5 Large

AI Judge Analysis

GPT Image 1.5

  • + Excellent adherence to the full prompt including the passenger.
  • + Photorealistic lighting and textures that feel grounded in reality.
  • + Accurate capybara anatomy and driver accessories.
  • The passenger's face is slightly blurry.
  • The paws on the steering wheel look slightly human-like in their grip.

Stable Diffusion 3.5 Large

  • + High saturation and clean, modern colors.
  • + Detailed rendering of the capybara's fur and clothing.
  • Completely missed the secondary character in the back seat.
  • The capybara's head shape and ears look more like a rat or degu than a capybara.
  • Logical inconsistencies with the steering wheel placement and the animal's limbs.

Verdict: GPT Image 1.5 successfully followed all instructions, including the presence of a bored businesswoman in the background, creating a cohesive and cinematic scene. Stable Diffusion 3.5 Large failed to include the passenger and produced a subject that resembled a small rodent more than a capybara, while also struggling with the internal geometry of the vehicle.

The Halloween Invitation

Text-to-Image

“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”

GPT Image 1.5
Stable Diffusion 3.5 Large

AI Judge Analysis

GPT Image 1.5

  • + Excellent text rendering with no spelling errors across all required fields.
  • + Highly cohesive gothic aesthetic with a beautiful thorn and web border.
  • + Superior cinematic lighting and texture on the central jack-o-lantern.
  • The parchment texture is slightly less distinct than the torn paper edge look.

Stable Diffusion 3.5 Large

  • + Creative use of a torn parchment edge as a frame.
  • + Dynamic composition with a large, glowing moon.
  • Failed to include the specific event details (Date, Time, Location) at the bottom.
  • Text rendering on the scroll is messy with artifacts and incomplete words.
  • Composition is cluttered, making the 'Party' text feel squashed.

Verdict: GPT Image 1.5 is the clear winner as it followed every instruction, including the specific date and location text which Stable Diffusion 3.5 Large omitted entirely. GPT Image 1.5 also maintained a much more consistent vintage gothic atmosphere with superior typography and lighting.

Isometric Miniature Diorama Scenes

Text-to-Image

“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”

GPT Image 1.5
Stable Diffusion 3.5 Large

AI Judge Analysis

GPT Image 1.5

  • + Excellent typography rendering with clean, centered alignment.
  • + High-quality PBR textures for the wood, ceramic, and food items.
  • + Perfect adherence to the 45-degree isometric diorama perspective.
  • The lighting is a bit more contrasty than the 'gentle' request imply.

Stable Diffusion 3.5 Large

  • + Features a wider variety of sushi pieces.
  • + Good textures on the salmon and rice grains.
  • Failed to place the text at the top-center, putting it on a small sign instead.
  • The chopsticks are poorly rendered and overlapping incorrectly.
  • The perspective is not a true isometric 45-degree angle.

Verdict: GPT Image 1.5 significantly outperformed Stable Diffusion 3.5 Large by correctly following the layout instructions, specifically the placement of the text and the flag at the top-center. Stable Diffusion 3.5 Large struggled with the isometric perspective and had noticeable rendering artifacts in the chopsticks.

Adorable Baby Animals in Sunny Meadow

Text-to-Image

“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”

GPT Image 1.5
Stable Diffusion 3.5 Large
100% wins 0% ties 0% wins

AI judge analyzing...

Vintage Cafe Logo

Text-to-Image

“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”

GPT Image 1.5
Stable Diffusion 3.5 Large
100% wins 0% ties 0% wins

AI judge analyzing...

Apollo 11: Journey to Tranquility

Text-to-Image

“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”

GPT Image 1.5
Stable Diffusion 3.5 Large
25% wins 0% ties 75% wins

AI judge analyzing...

Next steps

Explore each model