OpenAI's previous image generation model that accepts both text and image inputs and produces image outputs
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
GPT Image 1
#32 of 62 in Text-to-Image
Stable Diffusion 3.5 Large Turbo
#61 of 62 in Text-to-Image
Where the votes landed
GPT Image 1
0%
win rate
Ties
0%
Stable Diffusion 3.5 Large Turbo
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
GPT Image 1
- + Perfect adherence to spatial relationships defined in the prompt
- + Highly realistic glass refraction and textures
- + Natural lighting and shadows that enhance the depth of the scene
- − The glass cube has thick, somewhat irregular seams
Stable Diffusion 3.5 Large Turbo
- + Clean, sharp edges for the cube's frame
- + Good lighting contrast with sharp shadows
- − Fails to place the book on top of the cube, putting it inside instead
- − The plant is to the left rather than behind the cube
- − The cube frame and glass panels appear physically disconnected or illogical in construction
Verdict: GPT Image 1 followed every instruction perfectly, accurately placing the red book on top and the sphere inside while maintaining realistic glass transparency to show the plant behind. Stable Diffusion 3.5 Large Turbo failed the spatial logic of the prompt, placing the book inside the cube and the plant to the side instead of behind.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
GPT Image 1
- + Excellent realism with natural skin textures and realistic rain droplets on the character and bike.
- + Strong adherence to the cinematic 50mm lens look with effective shallow depth of field and bokeh.
- + The candid, 'imperfect' framing and lighting create a convincing street photography aesthetic.
- − The red bicycle's rear wheel and chain guard geometry are slightly warped.
- − The background car lights are slightly more static than 'motion blurred' but still effective for the scene.
Stable Diffusion 3.5 Large Turbo
- + The bicycle model is technically clean with realistic proportions.
- + Fulfills the basic prompt elements of a red bicycle and an elderly man in the rain.
- − Lacks the requested 'natural skin texture,' appearing smooth and plastic-like.
- − The rain is represented by simple vertical white lines that do not interact realistically with the subjects.
- − The composition feels like a cutout against a background rather than a cinematic street photo.
Verdict: GPT Image 1 significantly outperforms Stable Diffusion 3.5 Large Turbo in terms of photographic realism and prompt adherence. While the red bicycle in GPT Image 1 has minor structural flaws, its superior handling of skin texture, light reflections, and atmospheric depth creates a much more convincing 'candid' street photo compared to the flat, illustrative look of Stable Diffusion 3.5.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
GPT Image 1
- + Exceptional level of cinematic skin texture and realistic grit
- + Incredible detail in the engraved metal armor and leather neck guard
- + Atmospheric lighting that perfectly captures the orange torchlight and bokeh sparks
- − The beads in the hair are small and somewhat blended into the braid
Stable Diffusion 3.5 Large Turbo
- + Strong adherence to the braiding request
- + Clearer visibility of the green cloth underlayer
- + Distinct shallow depth of field in the background
- − The skin has a plastic, over-smoothed digital appearance
- − Blood/dirt marks look like flat digital brushes rather than being integrated into the skin
- − Light reflections on the nose and face look artificial
Verdict: GPT Image 1 is superior in terms of photorealism and artistic mood, successfully capturing the 'battle-worn' aesthetic with complex skin textures and sophisticated lighting. Stable Diffusion 3.5 Large Turbo follows the prompt instructions well but produces an output that looks like a high-end video game render rather than a lifelike photograph.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
GPT Image 1
- + Excellent typographic legibility and layout
- + High-quality realistic food photography
- − Nonsense filler text for descriptions
- − Redundant pricing and item names
Stable Diffusion 3.5 Large Turbo
- + Strong creative interpretation of grid-based food photography
- + Includes defined sections like 'Mians'
- − Food graphics appear overly stylized and plastic
- − Garbled text in the menu lists
Verdict: GPT Image 1 produces a more professional and usable layout with a clean aesthetic and high-quality photography, whereas Stable Diffusion 3.5 Large Turbo feels more like a conceptual collage with less realistic food. GPT Image 1 is preferred for its superior balance of design and imagery.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
GPT Image 1
- + Excellent text rendering with all requested phrases included correctly.
- + Highly photorealistic textures on the patty, lettuce, and bun.
- + Perfectly captures the 'exploded' view with clear separation of ingredients.
- − The price text is missing the '6' resulting in '.99' instead of '6.99'.
Stable Diffusion 3.5 Large Turbo
- + Strong sense of heat and fire with vibrant embers.
- + Good dynamic lighting on the cheese and sauce.
- + Captures the mid-air suspension concept effectively.
- − Fails to include any of the requested text elements.
- − The burger is not truly 'exploded' but largely stacked, missing the clear ingredient separation.
- − Less photorealistic, leaning towards a slightly digital or artificial look.
Verdict: GPT Image 1 followed the prompt requirements much more closely, successfully rendering the fiery text, starburst, and the specific 'exploded' layout. In contrast, Stable Diffusion 3.5 Large Turbo completely ignored the text prompts and failed to separate the ingredients into a dynamic exploded view, though it provided a more intense fiery atmosphere. GPT Image 1 is the clear winner for its superior prompt adherence and realistic food photography style.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
GPT Image 1
- + Excellent text rendering with perfect spelling and high legibility
- + Authentic chalkboard texture with realistic chalk dust and stroke variations
- + Consistent handwritten style across the entire board
- − The 'elegant cursive' request for the title was not fully met, as it appears more like print-script
- − Limited environmental context shown compared to the other model
Stable Diffusion 3.5 Large Turbo
- + Includes the cafe environment and lighting as requested
- + Used a cursive style for the main title heading
- − Numerous spelling errors including 'trulale', 'ocotpgs', and 'apri 31'
- − The handwriting looks like a digital font or a clean vector brush rather than realistic chalk
- − Layout is cluttered and fails to follow the specific text content requested
Verdict: GPT Image 1 followed the prompt perfectly, rendering all text with zero spelling errors and a highly convincing chalk texture. Stable Diffusion 3.5 Large Turbo struggled significantly with text legibility and accuracy, producing multiple gibberish words and failing to represent the specific date and prices correctly.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
GPT Image 1
- + Excellent anatomical depth and texture on the horse.
- + Cinematic lighting with a grit that matches the surreal theme.
- − Failed the specific positional prompt; the astronaut is on top.
Stable Diffusion 3.5 Large Turbo
- + Clean, high-contrast visual style.
- + Dynamic composition with the curve of the Earth's atmosphere.
- − Failed the specific positional prompt; the astronaut is on top.
- − Noticeable anatomy errors on the horse's front legs and hooves.
Verdict: Both models failed the negative constraint to put the horse on top of the astronaut, defaulting to a standard rider configuration. GPT Image 1 is superior due to its significantly higher level of textural detail and cinematic realism compared to the flatter, anatomically flawed Stable Diffusion 3.5 Large Turbo.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
GPT Image 1
- + Excellent adherence to the passenger description with a bored expression and phone.
- + High level of photorealism in fur texture and lighting.
- + Superior composition that clearly shows both the driver and the passenger as requested.
- − The capybara's paws lack anatomical accuracy and appear slightly mutated.
- − The steering wheel placement is slightly low relative to the animal's torso.
Stable Diffusion 3.5 Large Turbo
- + Clean, vibrant colors with a modern digital look.
- + Good placement of the capybara's paws on the steering wheel.
- − The passenger is heavily blurred and does not show the requested bored expression or phone clearly.
- − The capybara's head shape is slightly distorted and less characteristic of the species.
- − Missing the realistic taxi interior feel compared to the competitor.
Verdict: GPT Image 1 followed the prompt much more accurately, particularly regarding the passenger's expression and actions in the back seat. While Stable Diffusion 3.5 Large Turbo produced a clean image, it failed to clearly render the businesswoman as specified, making GPT Image 1 the more successful interpretation of the scene.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
GPT Image 1
- + Perfect text rendering for all requested details including date and location
- + Atmospheric and cinematic lighting that matches the 'vintage gothic' mood
- + Excellent composition with a professional scroll banner and border integration
- − The dark parchment makes the background details like the trees very subtle
Stable Diffusion 3.5 Large Turbo
- + Dynamic border featuring intricate webs and thorns
- + Strong contrast and clear jack-o-lantern rendering
- − Failed to include most of the required text like the location and date
- − Missed the specific scroll banner and the full 'Halloween Party Invitation' title
- − The bright background reduces the gothic atmospheric 'moody night sky' feel requested
Verdict: GPT Image 1 followed the prompt requirements perfectly, including all requested text, the scroll banner, and the specific event details accurately. Stable Diffusion 3.5 Large Turbo failed to include the event details and the scroll, and opted for a much cleaner, less 'vintage' aesthetic that deviated from the Gothic theme.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
GPT Image 1
- + Excellent text rendering with 'JAPAN' and 'SUSHI' spelled perfectly.
- + Highly polished 3D aesthetic with soft, professional lighting.
- + Accurate representation of the Japanese flag icon.
- + Strict adherence to the 45-degree isometric composition.
- − The scale of the chopsticks is slightly large compared to the plate.
Stable Diffusion 3.5 Large Turbo
- + Good use of the isometric diorama base.
- + Detailed rice texture on the maki rolls.
- − Misspelled text as 'SIIHI'.
- − The flag icon is incorrect, resembling a vertical Poland or Indonesia flag rather than Japan.
- − The text placement is on a billboard-style sign rather than being integrated at the top-center of the frame as requested.
Verdict: GPT Image 1 followed the instructions almost perfectly, delivering clean typography and a professional 3D render. In contrast, Stable Diffusion 3.5 Large Turbo failed to spell 'SUSHI' correctly and generated an inaccurate flag, resulting in a much less cohesive final image.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
GPT Image 1
- + Expertly captures all four requested animals (dog, cat, bunny, fox) with distinct anatomical features.
- + Excellent adherence to the 'golden sunrise' and 'god rays' lighting effects.
- + Superior fur texture and hyper-photorealistic rendering of the meadow and butterflies.
- − The kitten has an extra-wide mouth position that looks slightly unnatural.
Stable Diffusion 3.5 Large Turbo
- + Good use of backlighting and rim light on the animal silhouettes.
- + Vibrant colors and a clean, illustrative style.
- − Failed to include all four animals, missing the bunny and fox entirely.
- − The image has a CGI/cartoonish look rather than the requested hyper-photorealistic style.
- − Anatomy issues, such as the kitten's paws merging strangely with the dog.
Verdict: GPT Image 1 followed the complex prompt requirements perfectly, successfully including all four specific animals in a cohesive and photorealistic scene. In contrast, Stable Diffusion 3.5 Large Turbo failed to include the bunny and fox, and the overall image quality was more stylized and artificial with several anatomical merging errors.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
GPT Image 1
- + Excellent typography with perfect spelling and correct accent marks.
- + Clean vector emblem style that adheres to the minimalist request.
- + Accurate representation of a cloche dome and a classic banner.
- − Ignored the 'light background' request, providing a black background instead.
- − Minimalist approach might appear a bit too simple for some 'vintage' interpretations.
Stable Diffusion 3.5 Large Turbo
- + Successfully used a light background with subtle vintage textures.
- + High-quality vector shading and classic illustrative style.
- + Creative combination of a cloche and a coffee cup shape.
- − Spelling error in the name ('Caffeé Florin' instead of 'Caffè Florian').
- − Typography is slightly cluttered and less legible than Model A.
- − The steam lines are somewhat thin and fragile compared to the bold emblem.
Verdict: GPT Image 1 followed the core technical details and text requirements perfectly, though it failed to use a light background as requested. Stable Diffusion 3.5 Large Turbo captured the 'vintage' aesthetic and background light/texture much better but suffered from significant spelling errors and missed the 'minimalist' instruction by creating a more complex illustration. GPT Image 1 is the winner for its professional clarity and accurate text rendering despite the background color error.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
GPT Image 1
- + Excellent typography and readability for the key mission phases.
- + Clean iconographic style that perfectly matches the 'flat-vector' request.
- + Accurate NASA-inspired color palette and balanced layout.
- − Typos present in labels such as 'EARLLUNAR'.
- − The layout of labels and icons is slightly disorganized, with some labels far from their icons.
Stable Diffusion 3.5 Large Turbo
- + Strong professional composition that feels like a real infographic poster.
- + Excellent use of the retro NASA color scheme and stylistic texture.
- + Good balance between the main illustration and the sidebar information.
- − Significant text issues including many spelling errors and illegible filler text.
- − Failed to include all 6 specifically requested steps in the infographic sequence.
- − The central illustration looks more like a comet or meteor than a translunar trajectory.
Verdict: GPT Image 1 followed the specific structural instructions much better, providing distinct icons for most of the requested mission steps and including legible names. While Stable Diffusion 3.5 Large Turbo created a more visually striking 'poster' composition, it failed significantly on technical details, text accuracy, and adhering to the specific 6-step list.
Explore each model
Distilled version of SD 3.5 Large that generates high-quality images in just 4 steps, offering faster inference and reduced costs