Head to head
Esc

Models · slot A

to navigate to pick

GPT Image 1 OpenAI Stable Diffusion 3.5 Large Stability AI

Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.

GPT Image 1

23.2 arena score

#28 of 62 in Text-to-Image

Skill signature · Text-to-Image

Stable Diffusion 3.5 Large

22.8 arena score

#30 of 62 in Text-to-Image

Vote tally

Where the votes landed

GPT Image 1

0%

win rate

Ties

0%

Stable Diffusion 3.5 Large

0%

win rate

Shared challenges 13

Challenge by challenge

The strongest take from each model on every shared challenge, with the AI judge's read.

Geometric Composition

Text-to-Image

“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”

GPT Image 1
Stable Diffusion 3.5 Large

AI Judge Analysis

GPT Image 1

  • + Excellent adherence to the spatial requirements of the prompt.
  • + High visual quality with realistic textures on the wood and book.
  • + Clever use of depth of field and soft lighting.
  • The sphere appears to be floating mid-air without support or a ground plane.

Stable Diffusion 3.5 Large

  • + Realistic glass shader with subtle scratches and imperfections.
  • + Strong adherence to lighting direction (left side).
  • Failed the spatial arrangement; the book is inside the cube rather than on top of it.
  • The blue sphere is on top of the book rather than inside the cube in relation to the primary volume.
  • The scene is busier and more cluttered than requested.

Verdict: GPT Image 1 correctly followed all spatial instructions, placing the red book on top of the cube and the sphere inside. Stable Diffusion 3.5 Large failed the spatial logic by putting the book inside the cube and the sphere on the book, though its rendering of glass material was very convincing. Overall, GPT Image 1 is the superior choice for accurately following the descriptive prompt.

Candid Street Photography

Text-to-Image

“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”

GPT Image 1
Stable Diffusion 3.5 Large

AI Judge Analysis

GPT Image 1

  • + Excellent natural skin texture and facial realism
  • + Perfectly captures the lighting and moody atmosphere of a light rain
  • + Convincing cinematic depth of field
  • Lacks the requested motion blur on the background cars
  • Bicycle mechanics of the rear wheel area are slightly warped

Stable Diffusion 3.5 Large

  • + Better capture of the requested motion blur on passing vehicles
  • + Incorporates more visible raindrops as requested
  • + Detailed bike model with visible chain and basket
  • Anatomical issues with the man's arms, looking overly textured/distorted
  • Skin texture looks more artificial and stylized than requested
  • Composition feels slightly staged compared to Image A's 'candid' feel

Verdict: GPT Image 1 produces a far more believable and high-quality human portrait with natural skin and a genuine candid feel, though it ignores the request for motion blur. Stable Diffusion 3.5 Large adheres more closely to all technical prompt instructions like motion blur and rain visibility but at the cost of anatomical realism and a more 'processed' aesthetic. GPT Image 1 is the preferred winner due to its superior photographic realism and cohesive composition.

Fantasy Warrior

Text-to-Image

“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”

GPT Image 1
Stable Diffusion 3.5 Large

AI Judge Analysis

GPT Image 1

  • + Excellent character expression with highly realistic skin texture and visible scars.
  • + Superior lighting that captures the 'warm torchlight' atmosphere perfectly.
  • + Beautifully detailed engraving on the plate armor with believable metallic reflections.
  • The beads in the hair are present but visually subtle compared to the rest of the detail.
  • The depth of field is very shallow, almost blurring some of the armor details at the edges.

Stable Diffusion 3.5 Large

  • + Strong composition with a clear 'paladin' aesthetic and excellent braided hairstyle.
  • + Extremely intricate armor engravings that look very high-fantasy.
  • + Good adherence to the 'bokeh sparks' and 'beads' requirements.
  • The lighting feels more like daylight than the requested 'warm torchlight'.
  • The skin and eyes have a slightly digital/processed look compared to the cinematic quality of the first image.
  • The background contains blurred helmet shapes which were not requested and slightly clutter the composition.

Verdict: GPT Image 1 is the superior image due to its exceptional lighting and skin physics, which perfectly capture the 'battle-worn' and 'torchlight' elements of the prompt. While Stable Diffusion 3.5 Large provides more intricate braids and more visible beads, it fails to deliver the specific warm atmospheric lighting that gives GPT Image 1 its cinematic realism.

Modern Clean Menu

Text-to-Image

“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”

GPT Image 1
Stable Diffusion 3.5 Large

AI Judge Analysis

GPT Image 1

  • + Layout feels like a real, functional menu with clear sections and pricing.
  • + Text legibility is very high with clean sans-serif typography.
  • + Food photography is vibrant and high resolution with consistent lighting.
  • Nonsense 'Apperoiation descrigion' filler text is repetitive.
  • The 'grid' format is limited to just four large blocks rather than a gallery style.

Stable Diffusion 3.5 Large

  • + Excellent interpretation of a grid layout with a high volume of food photos.
  • + Design feels more like a sophisticated modern branding project.
  • + Strong visual hierarchy with the large 'Menu' header.
  • Text rendering is poor with many typographical errors like 'MAIMAES' and 'APPETIZRS'.
  • Layout is less functional as a menu due to the tiny, unreadable item details.
  • Some food photos in the grid appear slightly distorted or low-detail compared to Model A.

Verdict: GPT Image 1 is the superior choice because it functions as an actual menu with clear, legible sections and high-quality photography, despite the repetitive filler text. Stable Diffusion 3.5 Large offers a more creative, grid-based aesthetic, but fails on text accuracy and practical readability.

Magic Burger Explosion: Fiery Photorealism Challenge

Text-to-Image

“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”

GPT Image 1
Stable Diffusion 3.5 Large

AI Judge Analysis

GPT Image 1

  • + Excellent text rendering with the requested glowing/fiery effect.
  • + Perfectly follows the 'exploded burger' layout with clearly separated components.
  • + Includes the specific €6.99 price and starburst detail exactly as requested.
  • The price text is slightly truncated to '.99' instead of '€6.99' in the starburst.
  • The background is more static with embers rather than intense fire.

Stable Diffusion 3.5 Large

  • + Impressive photorealistic textures on the charred meat and melting cheese.
  • + Dynamic fire effects with good lighting interaction on the burger.
  • + High level of visual detail in the food components.
  • Failed to include any and all requested text ('MAGIC BURGER', price, etc.).
  • Ignored the 'exploded' instruction, providing a stacked burger instead.
  • Does not use a starburst or the requested secondary message.

Verdict: GPT Image 1 followed every instruction in the prompt, including complex text integration and layout requirements, though it slightly missed the full price digits. Stable Diffusion 3.5 Large produced a high-quality, appetite-appealing image but failed to follow any of the text or composition instructions, making it unusable as a 'Magic Burger' ad.

Chalkboard Menu

Text-to-Image

“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”

GPT Image 1
Stable Diffusion 3.5 Large

AI Judge Analysis

GPT Image 1

  • + Excellent text rendering with near-perfect spelling and legibility
  • + Convincing chalk texture with grainy edges consistent across all letters
  • + Followed the specific date and price instructions accurately
  • The 'elegant cursive' request for the title was not fully met, as it appears more like print
  • The chalkboard is a tight crop, missing the 'cozy café' environmental context

Stable Diffusion 3.5 Large

  • + Beautifully rendered environmental context showing the café interior
  • + Captures an artistic layout with decorative frames and plants
  • Significant spelling errors throughout the text (e.g., 'Todaay', 'Ottpups', 'Cholcalte')
  • Failed to follow the requested year (2024 instead of 2026)
  • The text is messy and illegible in several sections

Verdict: GPT Image 1 followed the complex text requirements with high precision, accurately spelling the menu items and adhering to the specific date and prices requested. While Stable Diffusion 3.5 Large provided a much better sense of atmosphere and environment, its failure to render coherent or correct text makes it less successful for this specific prompt.

The Reversed Rodeo

Text-to-Image

“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”

GPT Image 1
Stable Diffusion 3.5 Large

AI Judge Analysis

GPT Image 1

  • + Excellent anatomical detail on the horse and space suit
  • + Strong cinematic lighting and texture
  • + Clear focus on the subject with a balanced composition
  • Failed the negative constraint; the astronaut is riding the horse, not the other way around
  • Relatively standard interpretation despite the 'surreal' prompt keyword

Stable Diffusion 3.5 Large

  • + Dynamic sense of movement with cosmic dust and nebulous effects
  • + Vibrant colors and a sense of scale with the planetary curvature
  • + Highly detailed rendering of both characters
  • Failed the negative constraint; the astronaut is riding the horse
  • The harness/bridle lines are slightly messy and intersect unrealistically

Verdict: Both GPT Image 1 and Stable Diffusion 3.5 Large failed the specific prompt instruction to have the horse on top of the astronaut (a common logic-defying prompt test). While both models delivered high-quality cinematic visuals of an astronaut riding a horse, Stable Diffusion 3.5 Large is slightly more visually interesting due to the dynamic atmospheric effects in space.

The Capybara Taxi Driver

Text-to-Image

“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”

GPT Image 1
Stable Diffusion 3.5 Large

AI Judge Analysis

GPT Image 1

  • + Excellent adherence to all prompt elements including the passenger.
  • + Cinematic lighting and realistic taxi interior textures.
  • + Convincing 'bored' expression on the human passenger.
  • The capybara's paws look slightly human-like in their grip.

Stable Diffusion 3.5 Large

  • + High detail on the capybara's fur and whiskers.
  • + Vibrant colors and sharp focus on the central subject.
  • Completely missed the passenger in the back seat.
  • The capybara is not holding the steering wheel as requested.
  • The anatomy of the animal appears more like a rodent hybrid than a distinct capybara.

Verdict: GPT Image 1 followed the complex prompt instructions perfectly, capturing the specific humor of the bored passenger and the nighttime Manhattan atmosphere. Stable Diffusion 3.5 Large failed to include the passenger entirely and struggled with the specific positioning of the paws on the steering wheel, resulting in a less accurate and less creative composition.

The Halloween Invitation

Text-to-Image

“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”

GPT Image 1
Stable Diffusion 3.5 Large

AI Judge Analysis

GPT Image 1

  • + Excellent typography with perfect spelling and placement of the title, banner, and details.
  • + Moody and polished aesthetic that fits the 'vintage gothic' request perfectly.
  • + Very clean layout with clear cinematic lighting emanating from the jack-o-lantern.
  • Misidentified the event location as the 'Time' in the bottom text block.
  • The borders and thorns are quite dark and subtle, almost blending into the background.

Stable Diffusion 3.5 Large

  • + Captures the 'torn parchment' look very effectively with high-contrast borders.
  • + Dynamic composition with multiple glowing pumpkins and a large, detailed moon.
  • Failed to include the specific event details (Date, Time, Location) at the bottom.
  • The text rendering on the scroll banner is wobbly and contains artifacts/typos.
  • Includes multiple small pumpkins instead of the requested 'central' jack-o-lantern.

Verdict: GPT Image 1 is the clear winner as it successfully rendered almost all requested text elements with high legibility and a consistent gothic art style. While Stable Diffusion 3.5 Large has a nice parchment texture, it failed significantly on prompt adherence by omitting the specific event details and struggling with the banner text.

Isometric Miniature Diorama Scenes

Text-to-Image

“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”

GPT Image 1
Stable Diffusion 3.5 Large

AI Judge Analysis

GPT Image 1

  • + Excellent adherence to the 'cartoon' and 'soft refined texture' style requirements.
  • + Perfect text placement and rendering according to the prompt instructions.
  • + Clean, minimalist composition that fits the 'diorama' aesthetic.
  • The salmon texture is slightly over-simplified, looking more like plastic than a refined PBR material.

Stable Diffusion 3.5 Large

  • + High level of complex detail and variety in the sushi types.
  • + Good use of PBR materials with realistic textures on the rice and fish.
  • Failed the text placement instructions by putting text on a sign instead of 'top-center'.
  • The scene is cluttered with 'heavy' garnish, contradicting the 'minimal' requirement.
  • Includes a strange bowl of pink salt/grains that wasn't requested.

Verdict: GPT Image 1 followed every stylistic and layout instruction perfectly, including the specific text placement and the 'minimal' diorama aesthetic. Stable Diffusion 3.5 Large ignored the text placement instructions and the request for a minimal scene, resulting in a cluttered composition despite having high-quality individual textures.

Adorable Baby Animals in Sunny Meadow

Text-to-Image

“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”

GPT Image 1
Stable Diffusion 3.5 Large

AI Judge Analysis

GPT Image 1

  • + Excellent anatomical rendering with four distinct legs and paws for each animal
  • + Very strong adherence to the 'tumbling together' part of the prompt with dynamic posing
  • + Incredible detail in the fur texture and 'god rays' lighting
  • The fox kit lacks some of the characteristic facial markings usually seen in red foxes

Stable Diffusion 3.5 Large

  • + Beautiful bokeh and sparkling dew effects contribute to a magical atmosphere
  • + High-quality fur rendering and expressive facial expressions on the puppy
  • Anatomical errors such as the kitten missing its back body/legs and the puppy having a blurred limb
  • The kitten looks more like a caracal or generic feline than a domestic tabby
  • Composition is a bit flatter with animals lined up rather than tumbling together

Verdict: GPT Image 1 is the clear winner due to its superior handling of complex anatomy and interaction between the four animals. While Stable Diffusion 3.5 Large creates a beautiful atmosphere, it suffers from significant anatomical clipping (the kitten's body is missing) and less dynamic composition compared to the cohesive 'tumbling' scene in GPT Image 1.

Vintage Cafe Logo

Text-to-Image

“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”

GPT Image 1
Stable Diffusion 3.5 Large

AI Judge Analysis

GPT Image 1

  • + Excellent typography with correct spelling and accent mark.
  • + Clean minimalist aesthetic that resembles a real vector logo.
  • + Strong adherence to the warm brown color palette.
  • Failed the request for a light background, providing a black one instead.
  • The grainy texture is a bit uniform rather than subtle and organic.

Stable Diffusion 3.5 Large

  • + Successfully followed the instruction for a light, textured background.
  • + Includes a complex cloche illustration with internal steam details.
  • + Good use of the 'Est. 1720' request within a structured layout.
  • Spelling error in the main text ('Cafféé' instead of 'Caffè').
  • The composition feels a bit cluttered compared to a minimalist logo.
  • The cloche handle and steam shapes are slightly irregular.

Verdict: GPT Image 1 followed the stylistic and typographic requirements perfectly, producing a professional-looking logo, though it completely missed the light background instruction. Stable Diffusion 3.5 Large captured the background and texture well but failed on spelling ('Cafféé') and produced a more cluttered design. GPT Image 1 is the winner for its superior typography and clean vector-like execution.

Apollo 11: Journey to Tranquility

Text-to-Image

“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”

GPT Image 1
Stable Diffusion 3.5 Large

AI Judge Analysis

GPT Image 1

  • + Excellent adherence to the NASA-inspired color palette and flat-vector style.
  • + Clear and legible typography for most of the mission steps.
  • + Strictly followed the requested sequential steps and specific iconography.
  • Includes a few spelling errors like 'EARLLUNAR'.
  • The layout feels slightly disjointed with text separated far from its corresponding icon.

Stable Diffusion 3.5 Large

  • + Successfully captures the complex look of a technical space poster.
  • + Accurately represents the lunar surface with good detail.
  • Included a Space Shuttle instead of the Saturn V rocket requested.
  • Text is completely illegible gibberish and doesn't follow the specific step-by-step instructions.
  • Composition is cluttered and logical flow is non-existent.

Verdict: GPT Image 1 followed the instructions nearly perfectly, delivering a clean vector infographic with the correct NASA color scheme and mission steps. Stable Diffusion 3.5 Large failed on several fronts, most notably by generating a Space Shuttle instead of the Saturn V and failing to produce readable or instructional text.

Next steps

Explore each model