OpenAI's state-of-the-art image generation model with arbitrary resolution up to 4K and strong instruction following
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
GPT Image 2
#4 of 62 in Text-to-Image
Stable Diffusion 3.5 Large Turbo
#61 of 62 in Text-to-Image
Where the votes landed
GPT Image 2
100.0%
win rate
Ties
0.0%
Stable Diffusion 3.5 Large Turbo
0.0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
GPT Image 2
- + Excellent adherence to all spatial instructions in the prompt.
- + Highly realistic photographic quality with naturalistic lighting and textures.
- + Clear and coherent refraction through the glass cube.
- − The sphere appears slightly matte, though it matches the scene well.
Stable Diffusion 3.5 Large Turbo
- + Clean, sharp aesthetic with high-contrast lighting.
- + The blue sphere has a pleasing glossy finish.
- − Failed the spatial prompt by putting the book inside the cube instead of on top.
- − The plant is not visible behind/through the glass, only to the side.
- − Unnatural shadows and perspective on the table surface.
Verdict: GPT Image 2 followed the prompt perfectly, placing all objects in their correct relative positions with a high degree of realism. In contrast, Stable Diffusion 3.5 Large Turbo failed several spatial requirements, notably placing the book inside the cube and the plant to the side rather than behind the glass.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
GPT Image 2
- + Excellent adherence to the 'candid street photo' aesthetic with realistic lighting and environment.
- + High level of detail in skin texture, tools, and background elements like Japanese text.
- + Successfully captures motion blur from a passing car while maintaining clarity on the subject.
- − The framing feels slightly crowded with the foreground pillar and sign.
- − Minor structural issues with the bike's rear wheel spokes upon close inspection.
Stable Diffusion 3.5 Large Turbo
- + The red of the bicycle is vibrant and captures the requested color well.
- + Clean subject-background separation.
- − Faces and hands are poorly formed, with significant anatomical distortions.
- − Lacks the requested 'natural skin texture' and 'no stylization', looking more like a computer render than a photo.
- − The rain effect is composed of simple vertical white lines that look superimposed.
Verdict: GPT Image 2 is the clear winner as it authentically captures a cinematic, real-world atmosphere with natural textures and complex background details. In contrast, Stable Diffusion 3.5 Large Turbo produces a highly stylized, plastic-looking image with significant anatomical errors in the man's hands and face, failing the 'no stylization' and 'natural' requirements.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
GPT Image 2
- + Excellent photorealistic texture and skin detail.
- + Masterful lighting that accurately conveys warm torchlight on metal.
- + Ornate engraving on the armor is intricate and realistic.
- − The beads in the hair are somewhat subtle compared to the prompt's request.
Stable Diffusion 3.5 Large Turbo
- + Follows structural prompts like braids and scars clearly.
- + Good contrast and color saturation.
- − The facial features and skin have an uncanny, plastic-like 'airbrushed' quality.
- − The scars look like flat, painted-on textures rather than physical wounds.
- − Lack of depth in the engraving and material textures.
Verdict: GPT Image 2 is significantly superior, producing a cinematic and realistic portrait with nuanced lighting and complex material textures. In contrast, Stable Diffusion 3.5 Large Turbo produces a result that looks like a digital illustration or a video game character, failing to capture the lifelike quality and realistic 'battle-worn' grit requested.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
GPT Image 2
- + Excellent text rendering with clear, readable fonts for names, descriptions, and prices.
- + Comprehensive adherence to the prompt, including all three requested sections (Appetizers, Pizza, Mains).
- + Professional commercial-grade layout that looks like a real functional menu.
- − The 'Pizza' section images are slightly repetitive in their circular framing compared to the variety in other sections.
Stable Diffusion 3.5 Large Turbo
- + Good use of vibrant colors in the food photography elements.
- + Interesting overhead composition that feels artistic and modern.
- − Unreadable gibberish text throughout the menu sections.
- − Failed to include an 'Appetizers' section as requested in the text.
- − Visual artifacts and illogical food representations, such as the square pizza with floating leaves.
Verdict: GPT Image 2 is the clear winner as it produced a fully functional, professional menu with perfectly rendered text and logical categorizations. Stable Diffusion 3.5 Large Turbo failed on the fundamental task of menu design, producing illegible text and a confusing layout that resembles a mood board rather than a restaurant document.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
GPT Image 2
- + Perfect adherence to complex text requirements including the starburst and specific currency symbol.
- + Highly realistic textures on the beef patty, fresh vegetables, and melting cheese.
- + Excellent sense of motion and 'exploded' composition as requested in the prompt.
- − The composition is very crowded with little room for the background to breathe.
Stable Diffusion 3.5 Large Turbo
- + Bold and vibrant color palette with high-contrast lighting.
- + Creative use of smoke and fire to ground the burger in the environment.
- − Completely failed to include any of the requested text elements.
- − Failed to create an 'exploded' view; the burger components are mostly stacked.
- − The burger has an artificial, plastic-like texture compared to Model A.
Verdict: GPT Image 2 is the clear winner as it followed every instruction in the prompt, including the complex text rendering and the specific exploded layout. Stable Diffusion 3.5 Large Turbo failed to generate any text at all and missed the primary 'exploded' theme of the burger components.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
GPT Image 2
- + Excellent text rendering with perfect spelling and realistic chalk texture.
- + Authentic hand-drawn aesthetic that follows the prompt's request for a cohesive style.
- + Great composition that accurately captures a cozy café vibe.
- − None notable; it followed every instruction including specific dates and items.
Stable Diffusion 3.5 Large Turbo
- + Bright and clean visual composition.
- + Good contrast between the board and the background décor.
- − Significant spelling errors like 'specils', 'ocopus', and 'luea'.
- − Text looks like a digital font rather than natural chalk handwriting.
- − Failed to include the price for the first item and the full title date format.
Verdict: GPT Image 2 is the clear winner as it flawlessly rendered all the requested text with a realistic chalk texture and perfect spelling. Stable Diffusion 3.5 Large Turbo failed across multiple dimensions, including legibility, spelling accuracy, and the specific 'chalk handwriting' stylistic requirement, resulting in a digital-looking output with gibberish words.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
GPT Image 2
- + Excellent adherence to the complex spatial instruction of horse on top.
- + High level of photographic realism in the textures of the spacesuit and horse hair.
- + Coherent interaction between the horse, reins, and saddle on the astronaut.
- − The astronaut's hands/gloves have an incorrect number of fingers and look slightly mangled.
Stable Diffusion 3.5 Large Turbo
- + Clean, cinematic lighting and a dynamic sense of motion.
- + Good interpretation of 'in space' with a curved horizon of a planet.
- − Completely failed the negative constraint/specific instruction for the horse to be on top.
- − Severe anatomical issues with the horse's legs and hooves.
Verdict: GPT Image 2 followed the specific, surreal instruction to have the horse on top of the astronaut, whereas Stable Diffusion 3.5 Large Turbo ignored it and generated a standard astronaut on a horse. Additionally, GPT Image 2 achieved a much higher level of textural detail and realism, making it the clear winner.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
GPT Image 2
- + Excellent photorealism with convincing textures on the capybara's fur and jacket.
- + Perfectly captures the 'bored businesswoman' expression in the background.
- + Stronger composition with the capybara facing forward-right, as if actually steering.
- − The passenger's phone is slightly warped.
- − The steering wheel placement is a bit physically awkward relative to the seat.
Stable Diffusion 3.5 Large Turbo
- + Successfully includes the capybara, driver cap, and dark jacket.
- + Clean lighting and vibrant bokeh in the background.
- − The passenger is out of focus and does not appear to be looking at a phone as requested.
- − The capybara's fur texture and features look slightly more artificial/digital than Model A.
- − The woman is sitting in the front passenger seat instead of the back seat.
Verdict: GPT Image 2 is much more successful, as it correctly places the passenger in the back seat with the specific bored expression and phone requested. Stable Diffusion 3.5 Large Turbo fails on several prompt details, such as the passenger's location and action, and lacks the convincing photorealistic grit of the New York night scene found in GPT Image 2.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
GPT Image 2
- + Excellent typography rendering all requested text accurately
- + High level of detail in the gothic border, webs, and thorns
- + Rich atmospheric lighting and professional layout
- − The sky is slightly cluttered with many elements overlapping
Stable Diffusion 3.5 Large Turbo
- + Clean vector-like aesthetic
- + Includes therequested thorns and spiderwebs in the border
- − Failed to include most of the required text, including the scroll and date/location
- − Typography is plain and lacks the requested 'elegant gothic' style
- − The composition feels empty and lacks the 'vintage parchment' texture.
Verdict: GPT Image 2 is the clear winner as it followed every instruction, including the specific date, location, and multiple text banners with perfect spelling. Stable Diffusion 3.5 Large Turbo failed to include the event details and the scroll banner, providing a much simpler image that lacked the requested cinematic and vintage quality.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
GPT Image 2
- + Excellent text rendering with correct spelling of 'JAPAN' and 'SUSHI'.
- + High-quality PBR textures on the fish and wood materials.
- + Strong adherence to the requested diorama aesthetic with complex 3D details.
- − The garnish and elements are slightly more crowded than the 'minimal' request suggested.
Stable Diffusion 3.5 Large Turbo
- + Successfully captures a clean isometric diorama style.
- + Satisfies the solid background and soft lighting requirements.
- + Good 3D cartoon stylization.
- − Misspelled text with 'SIIHI' instead of 'SUSHI'.
- − The flag icon is generic/incorrect and does not look like the Japanese flag.
- − Layout is less sophisticated with awkward text placement on a sign.
Verdict: GPT Image 2 followed all prompt instructions perfectly, including complex text rendering, a correct Japanese flag icon, and realistic yet stylized PBR materials. Stable Diffusion 3.5 Large Turbo failed on text accuracy and the specific flag icon, resulting in a less polished final image.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
GPT Image 2
- + Successfully included all four requested animals (dog, cat, rabbit, fox).
- + Excellent realistic lighting with clear god rays and rim lighting on fur.
- + Dynamic composition with animals in motion, capturing the 'chasing' and 'tumbling' aspect.
- − The fox's front right paw has anatomical issues/extra joints.
- − Some butterflies are slightly blurry or missing wing detail.
Stable Diffusion 3.5 Large Turbo
- + Clean, sharp focus on the animals' faces.
- + Vibrant colors and high contrast.
- − Failed the prompt by only including three animals, missing both the fox and the bunny.
- − The art style is overly 'AI-plastic' and lacks the requested hyper-photorealistic texture.
- − Static pose does not capture the 'playfully chasing' or 'tumbling' action.
Verdict: GPT Image 2 is the clear winner as it followed all instructions, including the specific list of four different animals and the complex lighting effects. Stable Diffusion 3.5 Large Turbo failed to include the bunny or fox and produced a more generic, illustrated look rather than the requested hyper-photorealistic masterpiece.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
GPT Image 2
- + Excellent typography with perfect spelling and accent marks.
- + Sophisticated cross-hatch shading and fine line-work for a truly vintage feel.
- + Well-balanced composition with an elegant border.
- − The scale of the cloche relative to the text is slightly small.
Stable Diffusion 3.5 Large Turbo
- + Strong vector emblem style with high-contrast shading.
- + Creative integration of a coffee mug shape into the cloche design.
- − Spelling error in 'Caffe' (added an extra 'e') and 'Florin' (incorrect suffix).
- − Typography is warped and poorly rendered on the banner.
- − The background texture looks messy rather than subtle.
Verdict: GPT Image 2 is the clear winner as it perfectly follows all prompt instructions, including the correct spelling of 'Caffè Florian' and the inclusion of the establishment date. Stable Diffusion 3.5 Large Turbo fails on the text rendering, introducing spelling errors and messy typographical artifacts that make the logo unusable.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
GPT Image 2
- + Excellent typography and nearly perfect text rendering across all labels
- + Strictly followed all six sequential steps from the prompt with matching iconography
- + Professional, balanced infographic composition that feels like a finished educational product
- − Styling leans slightly toward 3D rendering rather than the requested flat-vector style
- − The Saturn V rocket in step 1 is missing its first stage fins
Stable Diffusion 3.5 Large Turbo
- + Captures the 'muted' retro-modern vector aesthetic and color palette color very well
- + Dynamic layout with interesting diagonal divisions
- − Failed to include the specific six-step sequence requested
- − Contains significant text gibberish ('Apoll.o', 'Laurch - Orbiit', 'Transoluna')
- − Included an astronaut character which was not part of the step-by-step landing sequence prompt
Verdict: GPT Image 2 (Model A) is the clear winner as it directly followed the complex multi-step instructions and produced a highly legible, accurate infographic. Stable Diffusion 3.5 Large Turbo (Model B) failed to render the requested steps and suffered from significant spelling errors and incoherent text sections.
Explore each model
Distilled version of SD 3.5 Large that generates high-quality images in just 4 steps, offering faster inference and reduced costs