Head to head
Esc

Models · slot A

to navigate to pick

GPT Image 2 OpenAI Stable Diffusion 3.5 Large Stability AI

Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.

GPT Image 2

28.1 arena score

#3 of 62 in Text-to-Image

Top 3 in Text-to-Image
Skill signature · Text-to-Image

Stable Diffusion 3.5 Large

22.8 arena score

#30 of 62 in Text-to-Image

Vote tally

Where the votes landed

GPT Image 2

100.0%

win rate

Ties

0.0%

Stable Diffusion 3.5 Large

0.0%

win rate

100.0% 0.0% ties 0.0%
Shared challenges 13

Challenge by challenge

The strongest take from each model on every shared challenge, with the AI judge's read.

Geometric Composition

Text-to-Image

“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”

GPT Image 2
Stable Diffusion 3.5 Large

AI Judge Analysis

GPT Image 2

  • + Excellent adherence to the spatial requirements of the prompt.
  • + Highly realistic textures on the book cover and wooden table.
  • + Clean glass rendering with accurate reflections and refractions.
  • The plant is more 'behind' than 'partially visible through' the glass due to the angle.

Stable Diffusion 3.5 Large

  • + Strong lighting effects with high-contrast sunlight.
  • + Good representation of the glass cube and blue sphere.
  • Failed the spatial prompt: the red book is inside the cube rather than on top of it.
  • The glass edges show some digital artifacts and inconsistent alignment.
  • The sphere appears to be floating unnaturally above the book.

Verdict: GPT Image 2 is the clear winner as it followed every spatial instruction perfectly, placing the sphere inside the cube and the book on top. Stable Diffusion 3.5 Large failed the layout by placing the book inside the cube, and the overall composition feels less physically grounded.

Candid Street Photography

Text-to-Image

“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”

GPT Image 2
Stable Diffusion 3.5 Large

AI Judge Analysis

GPT Image 2

  • + Excellent shallow depth of field and motion blur in the background
  • + Very realistic skin and clothing textures
  • + Captures the 'imperfect framing' prompt well with foreground objects
  • The rain is virtually invisible despite the wet pavement
  • Anatomy issues with the man's hands blending into the bicycle chain

Stable Diffusion 3.5 Large

  • + Strong atmospheric effect with visible rain and pavement reflections
  • + Follows the red bicycle requirement accurately within a wide street shot
  • + Good cinematic lighting and color grading
  • The man's skin looks overly textured and slightly burnt/saturated
  • Lacks the requested 'motion blur' from passing cars
  • Anatomy of the feet and shoes is distorted

Verdict: GPT Image 2 captures the 'candid photography' feel much better with realistic textures, a shallow 50mm lens look, and effective motion blur. However, Stable Diffusion 3.5 Large follows the prompt's weather requirement more accurately by showing actual rain, whereas GPT Image 2 relies only on wet surfaces.

Fantasy Warrior

Text-to-Image

“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”

GPT Image 2
Stable Diffusion 3.5 Large

AI Judge Analysis

GPT Image 2

  • + Exceptional photographic realism in skin texture and eyes
  • + Complex and delicate braided hair with integrated beads
  • + Masterful use of warm torchlight and material physics on the armor
  • The 'battle-worn' effect on the armor is somewhat subtle compared to the face
  • Fewer visible beads in the hair braids

Stable Diffusion 3.5 Large

  • + Ornate engraving on the armor is very clear and high-contrast
  • + Good implementation of facial scars and grime
  • + Includes a clear bokeh army background that adds scale
  • Missed the request for beads in the hair entirely
  • The armor texture feels slightly more like a 3D render than a photograph
  • Skin texture is a bit oily/over-sharpened

Verdict: GPT Image 2 (Model A) provides a much more lifelike and cinematic interpretation of the prompt, with superior lighting and incredibly detailed skin and hair textures. While Stable Diffusion 3.5 Large captures the ornate engraving and the 'battle-worn' scars well, it fails to include the requested hair beads and has a slightly less realistic skin finish.

Modern Clean Menu

Text-to-Image

“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”

GPT Image 2
Stable Diffusion 3.5 Large
100% wins 0% ties 0% wins

AI Judge Analysis

GPT Image 2

  • + Excellent text rendering with no spelling errors.
  • + Perfect adherence to requested sections (Appetizers, Pizza, Mains).
  • + Professional, balanced layout suitable for a real-world use case.
  • The food photos, while clean, have a slightly generic 'stock photo' appearance.

Stable Diffusion 3.5 Large

  • + High-quality, artistic photography with good lighting.
  • + Creative interpretation of a grid layout using border columns.
  • Extensive text corruption and gibberish ('APPETIZRS FHOPEADRE', 'MAIMAES').
  • Poor adherence to specific prompt sections, omitting the 'Pizza' header entirely in favor of nonsensical text.
  • Layout feels more like a social media mood board than a functional restaurant menu.

Verdict: GPT Image 2 is the clear winner as it produces a fully functional, professional-grade menu with perfect typography and logical structure. In contrast, Stable Diffusion 3.5 Large fails significantly on text legibility and structural coherence, resulting in a layout that is visually interesting but practically unusable.

Magic Burger Explosion: Fiery Photorealism Challenge

Text-to-Image

“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”

GPT Image 2
Stable Diffusion 3.5 Large

AI Judge Analysis

GPT Image 2

  • + Perfect adherence to all text requirements with high legibility and creative fiery effects.
  • + Excellent 'exploded view' composition that creates a strong sense of motion and suspended components.
  • + Highly photorealistic textures on the beef patty, lettuce, and bun.
  • The glowing sparks around the text are slightly repetitive in pattern.

Stable Diffusion 3.5 Large

  • + Beautiful lighting effects with flames interacting with the cheese and bun.
  • + Strong atmospheric background with realistic hot coals and embers.
  • Completely failed to include any of the requested text (Magic Burger, Limited Time, Price).
  • Did not follow the 'exploded' prompt instruction, showing a mostly assembled burger instead.
  • The burger components are flattened and lack the dynamic motion requested.

Verdict: GPT Image 2 is the clear winner as it followed every instruction in the prompt, including complex text integration and a specific 'exploded' layout. Stable Diffusion 3.5 Large produced a visually appealing image of a burger over fire but failed to include any of the required marketing text or the specific suspended-component composition requested.

Chalkboard Menu

Text-to-Image

“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”

GPT Image 2
Stable Diffusion 3.5 Large

AI Judge Analysis

GPT Image 2

  • + Excellent prompt adherence with perfectly spelled text for all requested items.
  • + Authentic chalk texture and realistic handwriting variation.
  • + Natural, warm lighting creates a cozy cafe atmosphere.

Stable Diffusion 3.5 Large

  • + Aesthetically pleasing layout of the cafe environment.
  • + Good rendering of the wooden frame and interior textures.
  • + Correct date formatting despite spelling errors.
  • Numerous spelling errors including 'TODAAY', 'Cholcalte', and 'Cockies'.
  • Text style appears more digital/uniform rather than natural chalk handwriting.
  • The year is incorrect, showing 2024 instead of 2026.

Verdict: GPT Image 2 followed the prompt perfectly, rendering all the requested text accurately with a very convincing chalk texture. Stable Diffusion 3.5 Large struggled significantly with text accuracy, creating several typos and failing to complete the menu items as requested in the specific handwritten style.

The Reversed Rodeo

Text-to-Image

“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”

GPT Image 2
Stable Diffusion 3.5 Large

AI Judge Analysis

GPT Image 2

  • + Excellent adherence to the 'horse on top' spatial instruction
  • + High level of texture detail in the space suit and lunar surface
  • + Clever use of composition and lighting to create a surreal atmosphere
  • The horse's front legs/hooves look slightly malformed and short
  • Visible distortion in the NASA-style logo text

Stable Diffusion 3.5 Large

  • + Beautiful cinematic lighting and ethereal nebula effects
  • + Good sense of movement and dynamic action
  • + High visual quality and clear details on the astronaut
  • Complete failure to follow the specific 'horse on top' instruction
  • Clipped composition where the horse's tail meets the edge of the frame

Verdict: GPT Image 2 followed the complex spatial prompt perfectly, depicting a horse literally riding an astronaut on all fours. Stable Diffusion 3.5 Large produced a more aesthetically pleasing and cinematic image, but it failed the primary challenge of reversing the typical rider/steed dynamic. GPT Image 2 is the clear winner for its success in handling the surreal logic requested.

The Capybara Taxi Driver

Text-to-Image

“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”

GPT Image 2
Stable Diffusion 3.5 Large

AI Judge Analysis

GPT Image 2

  • + Successfully includes the businesswoman in the back seat as requested.
  • + The lighting and textures feel very cinematic and photorealistic.
  • + The capybara's expression and posture perfectly match the 'professional' prompt.
  • The capybara's left paw is blended awkwardly into the steering wheel.

Stable Diffusion 3.5 Large

  • + High resolution with vibrant colors.
  • + Captures the yellow jacket and cap specified in the prompt well.
  • Completely fails to include the businesswoman in the back seat.
  • The capybara's anatomy is slightly distorted, with human-like legs and jeans growing out of its torso.
  • Only one paw is on the steering wheel, despite the prompt asking for both.

Verdict: GPT Image 2 is the superior image as it adheres to all complex prompt requirements, including the interaction between the capybara driver and the passenger. Stable Diffusion 3.5 Large fails to render the passenger entirely and has significant anatomical errors in the lower half of the capybara.

The Halloween Invitation

Text-to-Image

“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”

GPT Image 2
Stable Diffusion 3.5 Large

AI Judge Analysis

GPT Image 2

  • + Excellent typography with perfect spelling and professional alignment.
  • + Superior artistic detail in the border, including thorns, webs, and a gothic frame.
  • + Includes the location 'The Arches' and 'NYC' which are visible in the background art.
  • The parchment look is applied to the image texture rather than the physical paper shape.

Stable Diffusion 3.5 Large

  • + Successfully captures the torn parchment paper aesthetic with ragged edges.
  • + High contrast lighting makes the moon and pumpkins pop.
  • Failed to include the specific event details like the date, time, and location.
  • The text within the scroll banner has garbled characters at the bottom.
  • The font choice for the main title is less 'gothic' and more modern/bold.

Verdict: GPT Image 2 is the clear winner as it adhered to all text requirements, including the specific date and location which Stable Diffusion 3.5 Large omitted entirely. GPT Image 2 also featured a more intricate and sophisticated gothic art style that better matched the 'vintage polish' requested in the prompt.

Isometric Miniature Diorama Scenes

Text-to-Image

“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”

GPT Image 2
Stable Diffusion 3.5 Large

AI Judge Analysis

GPT Image 2

  • + Excellent text rendering and placement as requested at top-center.
  • + Strong technical execution of the isometric 45-degree angle.
  • + High-quality textures on the fish and decorative diorama elements.
  • The diorama base is slightly more complex than the 'minimal garnish' request.

Stable Diffusion 3.5 Large

  • + Natural look to the rice and salmon textures.
  • + Includes requested flag icon and text elements.
  • Failed the text placement instruction, grouping it on a side card instead of the top-center.
  • Composition is slightly cluttered with elements like the bowl of salt/rice in the background.
  • Camera angle is not a strict 45-degree isometric view.

Verdict: GPT Image 2 followed the prompt instructions near-perfectly, specifically placing the text and flag in a clean, top-center graphic style on a solid background as requested. Stable Diffusion 3.5 Large struggled with the layout requirements, integrating the text into the scene rather than as a header, and failed to achieve the specified isometric perspective.

Adorable Baby Animals in Sunny Meadow

Text-to-Image

“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”

GPT Image 2
Stable Diffusion 3.5 Large

AI Judge Analysis

GPT Image 2

  • + Excellent adherence to the requested species with clearly identifiable features.
  • + Superior detail and clarity in fur texture and eyes.
  • + Better rendering of butterflies and meadow environment.
  • The fox looks a bit too similar to a dog in its facial expression.
  • Lighting is a bit intense, making some fur edges look overly sharpened.

Stable Diffusion 3.5 Large

  • + Captures a very charming, joyful mood with dynamic movement.
  • + Good use of bokeh and soft lighting to create a dreamlike atmosphere.
  • + Includes many sparkling dew points as requested.
  • The kitten looks like a miniature fox or squirrel rather than a distinct tabby cat.
  • Lower overall resolution and clarity compared to Model A.
  • Anatomical issues with the rabbit's ears and the puppy's paws.

Verdict: GPT Image 2 is the clear winner as it accurately depicts all four specific animals with high-fidelity detail, whereas Stable Diffusion 3.5 Large struggles with the 'tabby' requirement and has significant anatomical blurring. GPT Image 2 also demonstrates better technical quality in rendering textures like fur and the wings of the butterflies.

Vintage Cafe Logo

Text-to-Image

“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”

GPT Image 2
Stable Diffusion 3.5 Large

AI Judge Analysis

GPT Image 2

  • + Excellent typography with perfect spelling and professional spacing
  • + Sophisticated engraving-style illustation that fits the 'vintage aesthetic'
  • + Highly cohesive vector emblem composition
  • May be slightly more complex than 'minimalist' suggests, though it fits the 'vintage' theme well

Stable Diffusion 3.5 Large

  • + Successfully captures a more minimalist vector style
  • + Good use of negative space in the cloche
  • Spelling error in 'Cafféé'
  • The cloche dome is floating awkwardly and contains multiple clashing steam styles
  • Text layout is less balanced than Model A

Verdict: GPT Image 2 is significantly better, featuring perfect typography and a professional, cohesive vintage emblem design. Stable Diffusion 3.5 Large fails on basic spelling and has a disjointed central illustration that lacks the polish of a real logo.

Apollo 11: Journey to Tranquility

Text-to-Image

“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”

GPT Image 2
Stable Diffusion 3.5 Large

AI Judge Analysis

GPT Image 2

  • + Excellent typography with perfect spelling of all technical terms and names.
  • + Followed all six mission steps accurately with relevant iconography.
  • + Maintained a clean, modern aesthetic with a consistent professional layout.
  • The Saturn V rocket illustration is slightly stylized rather than strictly flat-vector.
  • The lunar module icons are quite detailed for a 'flat vector' request.

Stable Diffusion 3.5 Large

  • + Successfully captured the requested color palette and flat vector style.
  • + Includes a large, visually appealing moon graphic at the base.
  • Failed significantly on text rendering, resulting in illegible gibberish.
  • The rocket illustrated is a space shuttle/generic rocket rather than the requested Saturn V.
  • Completely missed the sequential six-step structure requested in the prompt.

Verdict: GPT Image 2 is the clear winner as it produced a fully functional infographic with perfect spelling, accurate mission steps, and high-quality iconography. Stable Diffusion 3.5 Large failed to follow the structural instructions and produced garbled text, making the infographic unusable.

Next steps

Explore each model