OpenAI's previous image generation model that accepts both text and image inputs and produces image outputs
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
GPT Image 1
#29 of 62 in Text-to-Image
Stable Diffusion 3.5 Medium
#56 of 62 in Text-to-Image
Where the votes landed
GPT Image 1
0%
win rate
Ties
0%
Stable Diffusion 3.5 Medium
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
GPT Image 1
- + Excellent adherence to the complex spatial prompt.
- + High visual quality with realistic textures on the book, glass, and wood.
- + Natural and soft window lighting appearing from the requested direction.
- − The plant is more besides the cube than behind it, though it remains visible through the pane.
Stable Diffusion 3.5 Medium
- + Successfully places the plant behind the object as requested.
- + Clean rendering of the glass material.
- − Completely failed the spatial instruction for the red book, placing it inside the cube instead of on top.
- − The blue sphere is sitting on the book rather than just being inside the cube.
- − The composition is cropped poorly at the top.
Verdict: GPT Image 1 followed all spatial instructions accurately, placing the red book on top of the cube and the sphere inside. Stable Diffusion 3.5 Medium struggled with the logic of the prompt, incorrectly placing the red book inside the cube and resting the sphere upon it. GPT Image 1 also features superior texture realism and lighting.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
GPT Image 1
- + Excellent skin texture and hyper-realistic facial details
- + Accurate shallow depth of field with 50mm look
- + Consistent lighting and rain effect on surfaces
- − Anatomical issue with the hand merging into the bicycle chain and spokes
- − Missing motion blur from passing cars requested in prompt
- − Car headlights in the background are out of focus but static rather than blurred
Stable Diffusion 3.5 Medium
- + Stronger composition for a 'candid' street photo
- + Better implementation of the 'motion blur from passing cars' request
- + Captures the atmosphere of a rainy Japanese street very well
- − Poor facial detail and skin texture compared to Model A
- − Artificial-looking artifacts around the man's hands and the bicycle basket
- − Low clarity and a somewhat 'smeared' look rather than intentional shallow focus
Verdict: GPT Image 1 produces a much more intimate and technically detailed portrait, though it fails to include the requested motion blur and has some structural errors where the hand meets the bike. Stable Diffusion 3.5 Medium captures the 'candid street' energy and car motion much better, but falls behind significantly in terms of facial realism and overall image clarity. GPT Image 1 is the winner for its superior visual quality and adherence to the 'natural skin texture' and 'shallow depth of field' requirements.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
GPT Image 1
- + Excellent depiction of battle-worn textures including skin grime and aged metal.
- + Atmospheric lighting with consistent warm torchlight reflections.
- + Superior focus and shallow depth of field effects.
- − Lacks visible leather straps mentioned in the prompt.
- − The 'beads' in the hair are somewhat subtle and blend into the braid.
Stable Diffusion 3.5 Medium
- + Clear representation of leather straps and cloth underlayers.
- + Included distinct hair beads and more complex braiding patterns.
- + Sharp facial features and lifelike eye color.
- − The lighting feels more artificial and less like ambient torchlight.
- − The armor engraving and skin texture look too clean for a 'battle-worn' character.
- − Bokeh sparks appear as floating circles that don't integrate naturally with the scene.
Verdict: GPT Image 1 captures the 'battle-worn' atmosphere much better through grittier skin textures and realistic metallic aging. Stable Diffusion 3.5 Medium follows the specific structural elements of the prompt (like leather straps and beads) more literalally, but the overall image feels cleaner and more synthetic compared to the cinematic quality of GPT Image 1.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
GPT Image 1
- + Excellent typography rendering with clean, readable sans-serif fonts
- + High-quality, appetizing food photography with vibrant colors
- + Logical, minimalist layout that adheres perfectly to the grid request
- − Minor spelling errors in descriptive text like 'Apperoiation descrigion'
- − Limited number of menu items displayed compared to a full page
Stable Diffusion 3.5 Medium
- + Successfully creates a full-page menu concept
- + Good use of space and multiple food photos
- + Captures the professional casual dining aesthetic
- − Text is completely illegible and visually garbled
- − Food images lack fine detail and clarity
- − Poor source preservation on fonts, looking more like symbols than actual characters
Verdict: GPT Image 1 produces a high-fidelity, professional-looking design with clear text and beautiful photography that matches the modern minimalist prompt. In contrast, Stable Diffusion 3.5 Medium fails to generate readable text and the image quality of the food items is significantly lower, making it unusable as a menu design.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
GPT Image 1
- + Excellent adherence to the 'exploded' concept with all components clearly separated and suspended.
- + High-quality text rendering with the requested glowing, fiery effect.
- + Photorealistic textures on the food items, particularly the bun and the lettuce.
- − The price text is incorrect, displaying '.99' instead of '€6.99'.
- − The starburst around the price is a bit simplified compared to the rest of the professional-grade typography.
Stable Diffusion 3.5 Medium
- + Accurate text spelling including the price and the currency symbol.
- + Strong background dynamics with literal flames and flying debris.
- + Clean, modern graphic design for the price starburst.
- − Failed the core 'exploded burger' requirement; the burger is mostly assembled rather than suspended in components.
- − The text does not have the requested 'fiery, glowing effect' and looks like flat overlays.
- − Visible artifacts and blurring where the meat meets the cheese and lettuce.
Verdict: GPT Image 1 followed the creative direction much more effectively by accurately depicting the 'exploded' burger with high-detail textures and fiery glowing text. While Stable Diffusion 3.5 Medium got the price text correct, it failed to deliver on the suspended component layout and the specific requested text effects, resulting in a more generic-looking advertisement.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
GPT Image 1
- + Excellent text legibility and accuracy
- + Convincing chalk texture on the letters
- + Perfect adherence to the specific menu items and prices requested
- − Font looks slightly too uniform for 'natural variations'
- − Title is not in cursive as requested
Stable Diffusion 3.5 Medium
- + Beautiful artistic chalk style and layout
- + Included a wooden frame that enhances the cozy café aesthetic
- + Better variation in letter size and artistic flourishes
- − Severe spelling errors throughout the board
- − Failed to follow the requested menu text and prices
- − Layout is disorganized and cluttered with gibberish
Verdict: GPT Image 1 followed the prompt's text requirements with near-perfect accuracy, resulting in a functional and realistic menu, though it missed the request for a cursive title. Stable Diffusion 3.5 Medium produced a more visually artistic chalk effect but failed significantly on legibility and prompt adherence, rendering the text as nonsensical gibberish.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
GPT Image 1
- + Excellent cinematic lighting and texture on the space suit and horse hide.
- + Higher artistic coherence with a consistent 'surreal' aesthetic.
- + Better anatomical rendering of the horse compared to the competitor.
- − Prompt specified 'horse on top' (implying the horse riding the astronaut), but the astronaut is riding the horse.
Stable Diffusion 3.5 Medium
- + Clear, high-contrast composition against the planet's atmospheric glow.
- + Accurate space suit detailing and vibrant color palette.
- − Failed the negative constraint; the astronaut is riding the horse instead of the inverse.
- − Significant anatomical errors with the horse's legs, including extra segments and warped hooves.
- − The horse's neck and head join the body at an awkward, unrealistic angle.
Verdict: Both models failed the negative constraint to have the 'horse on top' of the astronaut, both defaulting to the standard interpretation of an astronaut riding a horse. GPT Image 1 is the superior image as it features high-quality cinematic lighting and anatomical accuracy, whereas Stable Diffusion 3.5 Medium has severe anatomical glitche with the horse's legs and neck.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
GPT Image 1
- + Perfect adherence to the phone-using businesswoman instruction
- + High photorealistic quality with naturalistic lighting and textures
- + Accurate representation of capybara anatomy holding a steering wheel
- − The perspective makes the capybara look slightly large for the cab interior
Stable Diffusion 3.5 Medium
- + Good centered composition that highlights the main subject
- + Vibrant colors and clear details on the capybara's face
- − Failed to include the businesswoman looking at her phone
- − Capybara's paws look more like bird claws or generic paws rather than capybara feet
- − The steering wheel is missing or poorly defined
Verdict: GPT Image 1 followed every detail of the prompt, including the specific behavior of the passenger and the placement of the capybara's hands on the wheel. Stable Diffusion 3.5 Medium failed to include the passenger's phone interaction and struggled with the anatomy and steering wheel placement, making GPT Image 1 the clear winner for prompt adherence and realism.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
GPT Image 1
- + Flawless text rendering with correct spelling and typography.
- + Cohesive and elegant vintage gothic aesthetic.
- + All requested elements including bats, trees, and scroll are integrated naturally into the composition.
- − Lighting is a bit dark, making some of the background details subtle.
Stable Diffusion 3.5 Medium
- + Bright and vibrant colors that pop against the sky.
- + Good use of the parchment paper texture and torn edges.
- − Numerous spelling errors in every line of text.
- − The composition feels cluttered with two jack-o-lanterns instead of one central one as requested.
- − The font style is more playful than the requested elegant gothic style.
Verdict: GPT Image 1 perfectly followed all instructions, including difficult text rendering and specific layout requirements, creating a professional-looking invitation. Stable Diffusion 3.5 Medium failed significantly on the text/spelling and didn't follow the 'central' jack-o-lantern instruction, resulting in a much lower-quality output for this specific task.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
GPT Image 1
- + Excellent adherence to the 'small raised diorama base' instruction.
- + Perfect text rendering and inclusion of the flag icon as requested.
- + Very clean 3D cartoon aesthetic with soft, refined textures.
- − The salmon texture on the nigiri is slightly repetitive.
Stable Diffusion 3.5 Medium
- + Successfully renders the sushi on a plate with requested text.
- + Good color vibrancy in the food items.
- − Missed the 'small raised diorama base' and 'flag icon' instructions.
- − Text layout is poor, with 'JAPAN' being much smaller than 'SUSHI' contrary to the prompt.
- − Low visual clarity with noticeable artifacts and blurry text elements.
Verdict: GPT Image 1 followed every instruction in the prompt, including the specific text hierarchy, the flag icon, and the diorama base. Stable Diffusion 3.5 Medium failed to include the base and flag, and the overall image quality was significantly lower with less clarity in the textures and text.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
GPT Image 1
- + Excellent adherence to the prompt by including all four distinct animals.
- + Dynamic composition with a strong sense of movement and 'tumbling' interaction.
- + Beautiful lighting with visible god rays and realistic textures on the fur.
- − The fox's front paw has a slightly awkward anatomical structure.
- − Minor clipping where the kitten's paw overlaps the rabbit.
Stable Diffusion 3.5 Medium
- + Vibrant color palette with high-contrast floral elements.
- + Clean, sharp rendering of the animals' faces.
- − Failed to include the baby bunny, only showing three animals.
- − The kitten's ears are unusually long and pointed, looking more like a caracal or a hybrid.
- − Static composition that lacks the 'tumbling' and 'chasing' action requested in the prompt.
Verdict: GPT Image 1 successfully followed all aspects of the prompt, including the specific list of four animals and the requested action of tumbling together. In contrast, Stable Diffusion 3.5 Medium missed the rabbit entirely and the animals appear to be posing rather than playing. GPT Image 1 also achieved a more naturalistic 'wholesome' lighting effect compared to the saturated, slightly artificial look of Stable Diffusion 3.5 Medium.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
GPT Image 1
- + Excellent text rendering with no spelling errors
- + Follows the minimalist vector emblem style perfectly
- + High contrast and clean lines suitable for a logo
- − Ignored the 'light background' instruction using black instead
- − Minimal visual texture compared to the request
Stable Diffusion 3.5 Medium
- + Successfully used a light background with subtle texture
- + Sophisticated vintage illustration style
- + Includes the 'banner' element and 'cream' tones
- − Multiple spelling errors in the main name and date
- − The cloche is poorly defined and resembles a helmet or building
- − Overly complex for a 'minimalist' logo request
Verdict: GPT Image 1 followed the core design principles of a logo and perfectly rendered all requested text, though it failed to use a light background. Stable Diffusion 3.5 Medium captured the requested aesthetic and color palette more accurately, but suffered from significant spelling failures and a cluttered composition that ignores the 'minimalist' requirement. GPT Image 1 is the winner for its functional utility as a logo and perfect typography.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
GPT Image 1
- + Excellent typography with legible names and headings
- + Clean flat-vector aesthetic that matches the requested NASA palette
- + Includes accurate supporting details like the astronaut names
- − Confusing layout order where labels don't always align with the icons below them
- − Spelling error in 'EARLLUNAR'
Stable Diffusion 3.5 Medium
- + Dynamic composition with a strong sense of trajectory and scale
- + Follows the sequential list format relatively well on the vertical axis
- − Poor typography with significant illegible 'loreum ipsum' style text
- − Failed to follow the specific iconography requests (e.g., Saturn V icon, Lunar Module on surface)
- − Messy visual elements that deviate from the requested 'crisp lines' flat-vector style
Verdict: GPT Image 1 followed the aesthetic and content instructions much more effectively, producing legible names and high-quality flat-vector icons, whereas Stable Diffusion 3.5 Medium struggled with text rendering and icon accuracy. While GPT Image 1 has a somewhat disjointed layout and a spelling error, it is far more usable as an infographic than the illegible output of Stable Diffusion 3.5 Medium.
Explore each model
Stability AI's 2.5-billion parameter Multimodal Diffusion Transformer with improvements (MMDiT-X) text-to-image model optimized for consumer hardware, featuring improved image quality, typography, and complex prompt understanding