Head to head
Esc

Models · slot A

to navigate to pick

GPT Image 1.5 OpenAI Stable Diffusion 3.5 Large Turbo Stability AI

Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.

GPT Image 1.5

27.1 arena score

#7 of 62 in Text-to-Image

Top 3 in Image Editing
Skill signature · Text-to-Image

Stable Diffusion 3.5 Large Turbo

10.7 arena score

#61 of 62 in Text-to-Image

Vote tally

Where the votes landed

GPT Image 1.5

100.0%

win rate

Ties

0.0%

Stable Diffusion 3.5 Large Turbo

0.0%

win rate

100.0% 0.0% ties 0.0%
Shared challenges 13

Challenge by challenge

The strongest take from each model on every shared challenge, with the AI judge's read.

Geometric Composition

Text-to-Image

“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”

GPT Image 1.5
Stable Diffusion 3.5 Large Turbo

AI Judge Analysis

GPT Image 1.5

  • + Excellent adherence to the complex spatial prompt.
  • + High visual realism with convincing glass refractions and wood textures.
  • + Accurate lighting and shadow work consistent with a window source.
  • The sphere is slightly large for the cube's interior proportions.

Stable Diffusion 3.5 Large Turbo

  • + Clean, modern aesthetic with sharp lines.
  • + Good rendering of the glass transparency and lighting.
  • Failed the spatial part of the prompt by placing the book inside and the plant to the left instead of behind.
  • The black frame on the cube was not requested and alters the look.

Verdict: GPT Image 1.5 successfully followed all spatial instructions, correctly placing the book on top of the cube and the plant behind it. Stable Diffusion 3.5 Large Turbo struggled with the positioning of elements, placing the book inside the cube and the plant to the side. GPT Image 1.5 is the clear winner for its superior prompt adherence and realistic execution.

Candid Street Photography

Text-to-Image

“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”

GPT Image 1.5
Stable Diffusion 3.5 Large Turbo

AI Judge Analysis

GPT Image 1.5

  • + Excellent photorealism with natural skin textures and realistic lighting
  • + Strong adherence to prompt details like reflections on wet pavement and shallow depth of field
  • + Authentic storytelling through tools, bike grime, and weather effects
  • The bicycle frame geometry is slightly distorted where it meets the rear wheel
  • The motion blur on the car in the background is subtle rather than pronounced

Stable Diffusion 3.5 Large Turbo

  • + Successfully includes the red bicycle and the elderly man in light rain
  • + Maintains a clean subject matter with no major background distractions
  • Highly stylized and 'plastic' look that fails the 'no stylization' requirement
  • Significant anatomical and structural errors, particularly where the man's head and hat merge
  • Lacks the requested 50mm cinematic aesthetic and realistic car motion blur

Verdict: GPT Image 1.5 produces a high-quality, believable photograph that captures the specific mood and textures requested in the prompt. In contrast, Stable Diffusion 3.5 Large Turbo fails on nearly every technical instruction, delivering a stylized image with significant AI artifacts, such as the man's head merging into his clothing and a lack of realistic environment reflections.

Fantasy Warrior

Text-to-Image

“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”

GPT Image 1.5
Stable Diffusion 3.5 Large Turbo

AI Judge Analysis

GPT Image 1.5

  • + Exceptional photographic realism with highly detailed skin texture and lifelike eyes.
  • + Excellent adherence to all prompt elements, including small beads in the braids and battle-worn grit.
  • + Superior lighting effects with warm torchlight reflections and atmospheric sparks.
  • The bokeh sparks are a bit large and blurry in the immediate foreground.

Stable Diffusion 3.5 Large Turbo

  • + Ornate engraving on the armor is clear and distinct.
  • + Maintains the requested shallow depth of field and bokeh background.
  • The skin has a plastic, 'uncanny valley' smoothness that lacks the requested lifelike texture.
  • The 'scars and dirt' look more like face paint or blood smears rather than integrated facial details.
  • Missing the small beads in the hair braids mentioned in the prompt.

Verdict: GPT Image 1.5 significantly outperforms Stable Diffusion 3.5 Large Turbo by delivering a truly lifelike portrait with grit, realistic skin textures, and intricate details like the beads in the hair. While Stable Diffusion 3.5 follows the basic composition, its rendering is overly smooth and stylized, failing to capture the 'battle-worn' realism requested.

Modern Clean Menu

Text-to-Image

“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”

GPT Image 1.5
Stable Diffusion 3.5 Large Turbo

AI Judge Analysis

GPT Image 1.5

  • + Excellent text rendering with no spelling errors or gibberish.
  • + Highly logical layout that functions as a real-world menu.
  • + Food photography looks realistic and appetizing.
  • Layout is somewhat safe and traditional rather than highly 'modern minimalist'.
  • Top accent bar is slightly cut off at the edge.

Stable Diffusion 3.5 Large Turbo

  • + Creative top-down photography style that feels very modern.
  • + Strong grid arrangement for the food photos.
  • + Excellent use of white space to create a clean aesthetic.
  • Text is completely illegible and contains various nonsensical characters.
  • Visual artifacts present, such as floating green sprouts and distorted food shapes.
  • The pizza looks more like a cartoon or toy than actual food.

Verdict: GPT Image 1.5 produced a fully functional, professional-grade menu with perfect text and realistic food photography that strictly follows the prompt. In contrast, Stable Diffusion 3.5 Large Turbo failed to generate readable text and produced several uncanny artifacts in the food items, leading to a much lower quality result for a design-focused task.

Magic Burger Explosion: Fiery Photorealism Challenge

Text-to-Image

“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”

GPT Image 1.5
Stable Diffusion 3.5 Large Turbo
100% wins 0% ties 0% wins

AI Judge Analysis

GPT Image 1.5

  • + Excellent adherence to the 'exploded' layout request with clear separation of ingredients.
  • + All required text elements are rendered perfectly with the specified glowing effect.
  • + High level of photorealistic detail in the food textures and embers.
  • The composition feels a bit crowded with the large starburst overlapping the burger elements.

Stable Diffusion 3.5 Large Turbo

  • + Dynamic use of fire and smoke creates a strong atmospheric effect.
  • + Good vertical symmetry and lighting.
  • Failed to include any of the requested text elements.
  • The burger is not 'exploded' but mostly intact, missing the core layout instruction.
  • Visual style is more digital/illustrative than photorealistic.

Verdict: GPT Image 1.5 followed the prompt instructions near-perfectly, successfully including the specific text phrases, the exploded layout, and the price starburst. Stable Diffusion 3.5 Large Turbo failed to include any text or the exploded layout, resulting in a generic burger image that missed several key requirements.

Chalkboard Menu

Text-to-Image

“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”

GPT Image 1.5
Stable Diffusion 3.5 Large Turbo

AI Judge Analysis

GPT Image 1.5

  • + Excellent text accuracy with near-perfect rendering of the complex menu items.
  • + Authentic chalk texture and handwriting style that looks truly hand-drawn.
  • + Successfully followed the date and price requirements.
  • The composition is very tight, focusing only on the board rather than the 'cozy café' environment.

Stable Diffusion 3.5 Large Turbo

  • + Successfully rendered a 'cozy café' environment with plants and lighting.
  • + Clean composition with a realistic frame around the board.
  • Severe spelling errors and gibberish text in the menu items.
  • The text style looks like a digital font rather than natural chalk handwriting.
  • Failed to include the specific year and many requested menu details.

Verdict: GPT Image 1.5 is the clear winner as it perfectly captured the specific and complex text requested, maintaining a highly realistic chalk texture throughout. Stable Diffusion 3.5 Large Turbo struggled significantly with the text, producing illegible words and failing to adhere to the handwritten style requirement.

The Reversed Rodeo

Text-to-Image

“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”

GPT Image 1.5
Stable Diffusion 3.5 Large Turbo

AI Judge Analysis

GPT Image 1.5

  • + Excellent fine details on the space suit and horse textures
  • + Richly detailed cinematic background with nebula, asteroids, and lunar lander
  • + Dynamic composition with a sense of motion and dust effects
  • The astronaut is clearly on top of the horse, failing the specific surreal 'horse on top' spatial instruction

Stable Diffusion 3.5 Large Turbo

  • + Clean aesthetic with high contrast lighting
  • + Accurate depiction of an astronaut and horse in a space environment
  • Fails the prompt instruction for 'horse on top'
  • Anatomical issues with the horse's front legs and hoof structure
  • The background/Earth rendering is somewhat blurry and lacks the 'highly detailed' quality requested

Verdict: Both models failed the negative constraint/spatial instruction to place the 'horse on top' of the astronaut, both delivering the standard 'astronaut riding a horse' trope. GPT Image 1.5 is the clear winner as it is significantly more detailed, featuring professional-grade textures, lighting, and a complex cinematic background, whereas Stable Diffusion 3.5 Large Turbo contains anatomical distortions and a much simpler composition.

The Capybara Taxi Driver

Text-to-Image

“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”

GPT Image 1.5
Stable Diffusion 3.5 Large Turbo

AI Judge Analysis

GPT Image 1.5

  • + Excellent photorealism with realistic lighting and textures
  • + Perfect adherence to all prompt elements, including the passenger's expression and activity
  • + The capybara's anatomy and paws on the wheel are convincing
  • The passenger's hair has some slight blending issues with the seat background
  • The capybara's jacket is a bit generic

Stable Diffusion 3.5 Large Turbo

  • + Strong cinematic lighting with vibrant bokeh in the background
  • + Modern, clean aesthetic
  • Failed the prompt requirement for the passenger to be looking at a phone
  • The capybara's face/snout looks slightly mutated and less like a real capybara
  • The taxi cap text is distorted

Verdict: GPT Image 1.5 is the clear winner as it followed every instruction in the prompt, including the specific behavior of the passenger looking at her phone. Stable Diffusion 3.5 Large Turbo failed to include the phone and biological accuracy of the capybara's face was lower than GPT Image 1.5.

The Halloween Invitation

Text-to-Image

“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”

GPT Image 1.5
Stable Diffusion 3.5 Large Turbo

AI Judge Analysis

GPT Image 1.5

  • + Excellent typography with perfect adherence to the required text and date.
  • + Highly detailed atmospheric illustrations including the jack-o-lantern, bats, and cemetery background.
  • + Perfectly captures the 'vintage gothic' and 'dark parchment' aesthetic requested.
  • The text at the bottom is slightly crowded toward the edge of the frame.

Stable Diffusion 3.5 Large Turbo

  • + Clean graphic style with high contrast.
  • + Good use of the thorny border and webs around the edges.
  • Fails to include almost all requested text including the banner and event details.
  • Lacks the 'cinematic lighting' and 'vintage gothic' parchment texture requested.
  • The 'Party' text is misspelled as 'Partij'.

Verdict: GPT Image 1.5 completely outperformed Stable Diffusion 3.5 Large Turbo by successfully following every instruction, including complex text rendering and specific design elements like the parchment texture and scroll banner. Stable Diffusion 3.5 Large Turbo failed to include the event details and the banner, and the overall style felt more like a modern clip-art vector than a vintage gothic invitation.

Isometric Miniature Diorama Scenes

Text-to-Image

“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”

GPT Image 1.5
Stable Diffusion 3.5 Large Turbo

AI Judge Analysis

GPT Image 1.5

  • + Excellent text rendering with 'JAPAN' and 'SUSHI' spelled perfectly.
  • + Highly detailed PBR materials with realistic textures for the fish, wood, and ceramic.
  • + Correct inclusion of the Japanese flag icon as requested.
  • Includes more props (teapot, soy sauce bottle) than the 'minimal garnish' requested.
  • The background has a subtle gradient rather than being a strictly 'solid' light blue.

Stable Diffusion 3.5 Large Turbo

  • + Captures the 'minimal' aesthetic well with a simple plate and few items.
  • + Good isometric perspective and clean, rounded 3D cartoon style.
  • Text rendering is poor, misspelling 'SUSHI' as 'SIIHI'.
  • The flag icon does not accurately represent the Japanese flag.
  • Materials look more like plastic than realistic PBR textures.

Verdict: GPT Image 1.5 followed the complex prompt instructions much more accurately, particularly regarding text spelling and the inclusion of the correct flag. While Stable Diffusion 3.5 Large Turbo captured the 'minimal' request better, it failed significantly on text legibility and material realism.

Adorable Baby Animals in Sunny Meadow

Text-to-Image

“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”

GPT Image 1.5
Stable Diffusion 3.5 Large Turbo

AI Judge Analysis

GPT Image 1.5

  • + Successfully included all four requested animals (dog, cat, bunny, fox).
  • + Excellent hyper-photorealistic textures and realistic backlighting.
  • + Naturalistic interaction and posing that matches the 'joyful wholesome vibe' perfectly.
  • The fox kit has slightly unusual paw anatomy.
  • Some minor overlapping artifacts where the kitten and bunny meet.

Stable Diffusion 3.5 Large Turbo

  • + Vibrant colors and high contrast.
  • + Clean, stylized aesthetic.
  • Failed to include the bunny and the fox; it appears to show a dog and two kittens.
  • Lacks the 'hyper-photorealistic' quality requested, looking more like digital 3D art.
  • The anatomy of the dog is merged awkwardly with the cat in the foreground.

Verdict: GPT Image 1.5 is the clear winner as it followed all aspects of the prompt, including the specific list of animals and the requested lighting effects like 'god rays'. Stable Diffusion 3.5 Large Turbo failed to include half of the requested animals and produced a less realistic, more stylized image with significant anatomical merging errors.

Vintage Cafe Logo

Text-to-Image

“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”

GPT Image 1.5
Stable Diffusion 3.5 Large Turbo

AI Judge Analysis

GPT Image 1.5

  • + Excellent typography with correct spelling and accent placement
  • + Refined textures on the cloche and banner
  • + Strong vector emblem composition with a professional aesthetic
  • Failed to provide a 'light background', instead using a solid black background

Stable Diffusion 3.5 Large Turbo

  • + Correctly followed the 'light background' and 'warm tones' instruction
  • + Creative integration of a coffee mug shape with a cloche top
  • Spelling error in 'Caffeé' (double 'e')
  • Typography is slightly distorted on the right side of the banner
  • The cloche handle and steam are slightly off-center

Verdict: GPT Image 1.5 produced a much higher quality logo with professional typography and textures, though it failed the background color requirement. Stable Diffusion 3.5 Large Turbo followed the color scheme perfectly but suffered from spelling errors and less polished vector execution. GPT Image 1.5 is the winner for its superior graphic design quality and font rendering.

Apollo 11: Journey to Tranquility

Text-to-Image

“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”

GPT Image 1.5
Stable Diffusion 3.5 Large Turbo

AI Judge Analysis

GPT Image 1.5

  • + Excellent adherence to the 6-step chronological sequence requested.
  • + Perfect text rendering for all labels including names and mission phases.
  • + Clean, consistent flat-vector style that matches the NASA-inspired color palette perfectly.
  • The composition is a bit rigid and repetitive with rectangular boxes.
  • The orbital mechanics depicted in 'Earth Orbit' and 'Lunar Orbit' are stylized but basic.

Stable Diffusion 3.5 Large Turbo

  • + Visually striking artistic composition and layout.
  • + Good use of the requested color palette to create depth and mood.
  • + Effective use of white space and modern typography styles.
  • Failed to provide the specific 6-step sequence, merging and missing steps.
  • Poor text rendering with several typos and various instances of gibberish.
  • The iconography does not match the specific prompts (e.g., missing Saturn V launch icon).

Verdict: GPT Image 1.5 followed the prompt instructions with near-perfect accuracy, delivering all six requested steps with clear, legible text and appropriate iconography. While Stable Diffusion 3.5 Large Turbo created a more artistic and visually interesting poster, it failed on technical adherence, mispelled keywords, and ignored the specific sequential requirements. GPT Image 1.5 is the clear winner for its functionality and precise communication of the requested information.

Next steps

Explore each model