OpenAI's state-of-the-art image generation model with better instruction following and adherence to prompts
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
GPT Image 1.5
#7 of 62 in Text-to-Image
Stable Diffusion 3.5 Medium
#56 of 62 in Text-to-Image
Where the votes landed
GPT Image 1.5
100.0%
win rate
Ties
0.0%
Stable Diffusion 3.5 Medium
0.0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
GPT Image 1.5
- + Excellent adherence to all spatial instructions and prompt details.
- + Highly realistic lighting, reflections, and material textures.
- + Balanced composition with clear depth between the cube and the plant.
- − The sphere is slightly large compared to the description of a 'small' sphere.
Stable Diffusion 3.5 Medium
- + Successfully includes all requested elements in the scene.
- + Good color accuracy for the blue sphere and red book.
- − The blue sphere is levitating inside the cube, which defies physics without explanation.
- − The red book looks more like a felt block with significant blurring artifacts on its edges.
- − Lower overall photographic quality and clarity compared to the competitor.
Verdict: GPT Image 1.5 produced a much more realistic and professionally composed image that perfectly handled the glass refractions and spatial relationships of the objects. Stable Diffusion 3.5 Medium followed the prompt's content list, but the floating position of the sphere and the poor texture of the book made it look less coherent.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
GPT Image 1.5
- + Excellent adherence to the 'repairing' aspect of the prompt
- + Highly realistic skin textures and clothing details
- + Strong composition with realistic raindrops and wet pavement reflections
- − The motion blur of the passing car is somewhat static rather than streaky
Stable Diffusion 3.5 Medium
- + Good color contrast and mood
- + Reflections on the ground are vibrant
- − The man appears to be leaning on the bike rather than actively repairing it
- − Anatomical issues with the man's hands and face
- − The bike geometry is nonsensical with the pedal and chain missing or misplaced
Verdict: GPT Image 1.5 followed the prompt much more accurately, depicting a man crouched and actively repairing a mechanical component of a bicycle. Stable Diffusion 3.5 Medium struggled with the 'repairing' action and produced significant anatomical and mechanical distortions, particularly in the hands and bike frame.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
GPT Image 1.5
- + Exceptional photographic realism in skin texture and eyes
- + Highly detailed and accurate engraving on the plate armor
- + Successfully incorporates beads within the hair braids as requested
- − The lighting is a bit uniform across the face, lacking the high-contrast drama of a single torch light
Stable Diffusion 3.5 Medium
- + Excellent dramatic lighting with a strong warm glow contrasted against cooler tones
- + Captures a very intense and battle-worn emotional expression
- + Intricate hair braiding style
- − Missed the request for beads in the hair braids
- − Armor engraving and textures look slightly muddy compared to Model A
- − Artificial-looking skin texture on the forehead and nose
Verdict: GPT Image 1.5 offers superior detail in the armor engravings and achieves a much higher level of photorealism, especially in the eyes and skin. While Stable Diffusion 3.5 Medium has more dramatic cinematic lighting, it missed the specific detail of beads in the hair and lacks the crisp textural fidelity found in GPT Image 1.5.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
GPT Image 1.5
- + Excellent text rendering with no spelling errors or gibberish
- + Photorealistic food photography that aligns perfectly with the menu items
- + Clean, professional layout that matches the 'modern minimalist' prompt
- − Layout is a split column rather than a full-page grid of photos
Stable Diffusion 3.5 Medium
- + Good interpretation of a multi-photo grid layout
- + Accurate representation of a white background with vibrant food imagery
- − Text is completely illegible and contains nonsense characters
- − Food images contain visual artifacts and lack clarity compared to Model A
- − Menu layout is cluttered and non-functional for a real business
Verdict: GPT Image 1.5 wins this challenge decisively by producing a fully functional, professional-grade menu with perfect typography and high-quality food photography. Stable Diffusion 3.5 Medium struggles significantly with text rendering and overall image coherence, resulting in a layout that is visually messy and practically unusable.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
GPT Image 1.5
- + Excellent adherence to the 'exploded' and 'suspended' layout requested.
- + Superb integration of fiery, glowing text effects that match the background.
- + Highly detailed and realistic food textures with dynamic lighting.
- − The composition is very busy, bordering on cluttered.
Stable Diffusion 3.5 Medium
- + Clean text rendering with no spelling errors.
- + Realistic fire in the background and good depth of field.
- − Failed the 'exploded burger' requirement; the burger is mostly assembled.
- − The starburst and text are flat and do not match the requested 'fiery effect' style.
- − Layout feels static rather than dynamic.
Verdict: GPT Image 1.5 is the clear winner as it perfectly captured the 'exploded' burger concept and the specific fiery aesthetic for the typography. Stable Diffusion 3.5 Medium failed to separate the burger components in mid-air and provided a generic, flat graphic style for the pricing and text.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
GPT Image 1.5
- + Excellent text rendering with perfect spelling for the entire prompt description.
- + Highly realistic chalk texture with natural smudging and believable handwriting variations.
- + Strict adherence to the formatting and content requested in the prompt.
- − The composition is a bit plain, focusing only on the board face without showing the surrounding café context.
Stable Diffusion 3.5 Medium
- + Includes a wooden frame which adds to the 'cozy café' atmosphere.
- + Good implementation of chalky texture and varying stroke weights.
- − Extremely poor text rendering with numerous spelling errors and gibberish characters.
- − Failed to follow the specific title and menu item layout requested in the prompt.
- − The 'handwriting' looks more like digital bubble fonts than natural human script.
Verdict: GPT Image 1.5 followed the prompt with near-perfect accuracy, rendering every specific menu item and date exactly as requested with highly realistic handwriting. In contrast, Stable Diffusion 3.5 Medium failed significantly on text coherence, producing illegible words and incorrect formatting despite having a nice environmental frame.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
GPT Image 1.5
- + Excellent cinematic lighting and dynamic composition.
- + High level of texture detail on the astronaut suit and horse hair.
- + Coherent environmental details like the lunar lander and dust kicks.
- − Failed the negative constraint; the astronaut is riding the horse, not vice versa.
Stable Diffusion 3.5 Medium
- + Successfully placed the astronaut on a horse in a space environment.
- + Clean, minimalist composition.
- − Failed the negative constraint; the astronaut is riding the horse, not vice versa.
- − Anatomical issues with the horse's legs and the astronaut's lower body.
- − Lower overall resolution and detail compared to the competitor.
Verdict: Both models failed the specific spatial logic requested in the prompt ('horse on top, not vice versa'), defaulting to the standard image of a person riding a horse. GPT Image 1.5 is the superior image due to its significantly higher artistic quality, cinematic detail, and realistic textures, whereas Stable Diffusion 3.5 Medium produced anatomical distortions and a flatter aesthetic.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
GPT Image 1.5
- + Excellent photorealism with convincing textures on the capybara fur and the car interior.
- + Strong prompt adherence, showing the passenger looking at her phone as requested.
- + Realistic cinematic lighting that captures the New York night atmosphere well.
- − The capybara's front paws look slightly more like canine paws than true capybara feet.
Stable Diffusion 3.5 Medium
- + Features a more colorful and vibrant lighting style.
- + Accurately places the capybara in a driving position with a hat and jacket.
- − Failed the negative constraint/activity prompt for the passenger, who is looking at the camera instead of her phone.
- − The capybara's face looks slightly distorted and less photorealistic than Image A.
- − The perspective and depth of field make the passenger look unnaturally small or poorly integrated into the scene.
Verdict: GPT Image 1.5 is the clear winner as it followed all prompt instructions, including the specific behavior of the passenger in the back seat. Stable Diffusion 3.5 Medium failed to depict the woman looking at her phone and lacked the gritty, photorealistic texture found in the first image.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
GPT Image 1.5
- + Excellent typography with perfect spelling of all requested text.
- + High-quality cinematic lighting with a coherent vintage gothic aesthetic.
- + Superior composition that integrates the border, scroll, and central subject seamlessly.
- − The parchment texture is very dark, which may reduce readability for a physical card.
Stable Diffusion 3.5 Medium
- + Successfully includes the parchment, bats, and jack-o-lantern elements.
- + Uses a clear layout with distinct sections for text.
- − Serious spelling errors in almost every line of text ('Invillowen', 'Timme', 'Loccation').
- − Poor integration of elements, with pumpkins appearing to float or be stuck in trees awkwardly.
- − Failed to create a 'central' jack-o-lantern as requested, placing two on the sides instead.
Verdict: GPT Image 1.5 significantly outperforms Stable Diffusion 3.5 Medium by following every textual instruction and the stylistic intent perfectly. While Stable Diffusion 3.5 Medium struggled with severe spelling errors and disjointed composition, GPT Image 1.5 produced a polished, professional-looking invitation that is ready for use.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
GPT Image 1.5
- + Excellent adherence to the diorama base and miniature set piece request.
- + Accurate rendering of the flag icon and all requested text.
- + Highly detailed PBR materials on the wood, ceramic, and food textures.
- − Slightly more than 'minimal' garnish and plates by including more props like the teapot.
Stable Diffusion 3.5 Medium
- + Successfully rendered the 'SUSHI' text in a bold, stylized font.
- + Followed the soft blue background requirement well.
- + Clean, minimalist composition.
- − Failed to include the diorama base, flag icon, and correctly positioned header text.
- − Artifacts present on the 'S' of SUSHI and a random floating dot over the 'i'.
- − The 45-degree isometric angle is less defined compared to the other model.
Verdict: GPT Image 1.5 is the clear winner as it followed every instruction, including the complex request for a diorama base and specific text/icon arrangement. Stable Diffusion 3.5 Medium failed to generate the flag, the diorama, and the 'JAPAN' text correctly, and it showed noticeable artifacts in the typography.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
GPT Image 1.5
- + Successfully includes all four requested animals (dog, cat, bunny, fox).
- + Excellent rendering of textures, particularly the kitten's paw pads and individual fur follicles in the bokeh.
- + Masterful use of lighting with visible 'god rays' and dew sparkles that match the prompt perfectly.
- − The fox's paw in the bottom right corner looks slightly distorted/human-like.
Stable Diffusion 3.5 Medium
- + Vibrant colors and high contrast create a very cheerful atmosphere.
- + The butterfly wings have intricate, sharp patterns.
- − Failed to include the requested baby bunny animal.
- − The fox has an anatomical issue with its right leg appearing to merge into its chest.
- − The fur texture looks overly sharpened and digital rather than 'hyper-photorealistic'.
Verdict: GPT Image 1.5 is the clear winner as it adhered to all prompt instructions, including the specific list of four animals, whereas Stable Diffusion 3.5 Medium missed the bunny entirely. Furthermore, GPT Image 1.5 achieved a much more realistic lighting effect and superior fur textures, while Stable Diffusion 3.5 Medium suffered from anatomical merging and a more plastic, digital look.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
GPT Image 1.5
- + Excellent typography with perfect spelling and accent marks.
- + Clean vector emblem style with high-quality shading.
- + Adheres perfectly to all text elements in the prompt.
- − Ignored the 'light background' instruction, delivering a black background instead.
- − The cloche handle is slightly off-center.
Stable Diffusion 3.5 Medium
- + Captures the warm cream tones and light textured background beautifully.
- + Engraving-style illustration adds a sophisticated vintage feel.
- − Failed significantly on text accuracy, misspelling 'Florian' as 'Florrian' and '1720' as '170'.
- − The cloche dome is poorly defined and looks more like a decorative orb or tent.
Verdict: GPT Image 1.5 produced a professional, usable logo with perfect text rendering, though it failed the background color instruction. Stable Diffusion 3.5 Medium captured the requested background and texture better but failed basic spelling and the core structural elements of a cloche. GPT Image 1.5 is the preferred choice for a branding task where text accuracy is paramount.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
GPT Image 1.5
- + Excellent adherence to the sequential steps requested.
- + Perfect text rendering for all labels and names.
- + Clean, professional vector aesthetic that matches the 'infographic' requirement.
- − The rocket in the 'Launch' panel is missing its top half due to the crop.
- − Uses slightly more shading than a strict 'flat-vector' style would suggest.
Stable Diffusion 3.5 Medium
- + Bold color palette that fits the NASA-inspired theme.
- + Interesting abstract composition for a poster.
- − Text is largely nonsensical or misspelled.
- − Icons do not accurately correspond to the requested steps (e.g., 'Descant' for 'Descent').
- − The layout is cluttered and fails to function as a clear infographic.
Verdict: GPT Image 1.5 followed the prompt exactly, delivering a clear, legible, and logically ordered infographic with perfect text. In contrast, Stable Diffusion 3.5 Medium struggled with text legibility and failed to accurately represent the specific iconography and sequence requested.
Explore each model
Stability AI's 2.5-billion parameter Multimodal Diffusion Transformer with improvements (MMDiT-X) text-to-image model optimized for consumer hardware, featuring improved image quality, typography, and complex prompt understanding