OpenAI's state-of-the-art image generation model with arbitrary resolution up to 4K and strong instruction following
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
GPT Image 2
#3 of 62 in Text-to-Image
Stable Diffusion 3.5 Medium
#56 of 62 in Text-to-Image
Where the votes landed
GPT Image 2
100.0%
win rate
Ties
0.0%
Stable Diffusion 3.5 Medium
0.0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
GPT Image 2
- + Perfectly captures all spatial relationships including the plant behind the glass.
- + Superior photographic quality with realistic textures on the book and wooden table.
- + Accurate lighting and reflections that respect the physics of the glass cube.
- − The glass cube has double-walled edges that look slightly more like a display case than a solid glass object.
Stable Diffusion 3.5 Medium
- + Achieves a bright, airy aesthetic with vibrant colors.
- + Shows the plant behind the glass as requested.
- − The blue sphere is floating unnaturally in the center of the cube without support.
- − The book appears flattened and lacks realistic perspective/dimensions.
- − The glass cube geometry is distorted, with the bottom face appearing curved or misaligned.
Verdict: GPT Image 2 followed the prompt perfectly, placing the sphere on the bottom of the cube and rendering a highly realistic red book. In contrast, Stable Diffusion 3.5 Medium struggled with physics and geometry, resulting in a floating sphere and a distorted cube. GPT Image 2 is the clear winner for its photographic realism and coherent composition.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
GPT Image 2
- + Excellent adherence to technical prompts like 50mm lens feel and shallow depth of field
- + Natural skin textures and highly detailed mechanical components on the bicycle
- + Effective use of motion blur in the background to suggest a candid street environment
- − The white sign in the immediate foreground is a bit distracting
- − Hand interaction with the wheel spokes is slightly messy upon close inspection
Stable Diffusion 3.5 Medium
- + Atmospheric lighting and very strong reflections on the wet pavement
- + The color palette captures a moody, rainy urban aesthetic
- − Significant anatomy/geometry issues with the bicycle frame and handlebars
- − Lack of 'repairing' action; the subject appears to just be holding the bike
- − Image quality is grainier and lacks the requested 'natural skin texture' detail
Verdict: GPT Image 2 is the clear winner as it successfully follows the complex technical requirements of the prompt, including the specific focal length feel and the action of repairing. Stable Diffusion 3.5 Medium captures a nice rainy atmosphere, but fails on the physical logic of the bicycle and the clarity of the subject's features.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
GPT Image 2
- + Excellent photorealism with very natural skin texture and eyes
- + Highly detailed engraving on the plate armor that looks physically plausible
- + Subtle and effective hair braiding with beads that feels integrated into the style
- − Lighting is a bit soft and lacks the intense 'warm torchlight' contrast described
Stable Diffusion 3.5 Medium
- + Strong prompt adherence regarding the warm torchlight and bokeh sparks
- + Good representation of the braids and battle-worn skin texture
- + Vibrant colors and high contrast create a dramatic mood
- − The eyes and hair appear slightly over-sharpened and digital
- − The armor engraving is a bit flat and less intricate compared to Model A
- − Occasional artifacting in the hair and fire effects
Verdict: GPT Image 2 provides a significantly more lifelike and cohesive portrait with superior material rendering, particularly on the armor and skin. While Stable Diffusion 3.5 Medium captures the lighting and 'sparks' aspect of the prompt more aggressively, it lacks the anatomical and textural realism found in the first image.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
GPT Image 2
- + Exceptional text rendering with perfect spelling and realistic descriptions.
- + Highly professional and clean layout that matches the 'modern minimalist' prompt exactly.
- + Consistent high-quality food photography that fits a grid pattern perfectly.
- − None identified, it functions as a ready-to-use professional graphic design.
Stable Diffusion 3.5 Medium
- + Good adherence to the requested grid-based layout.
- + Includes a white background and colorful food photography as requested.
- + Attempts to follow the sectioning requested in the prompt.
- − Text is nonsensical and garbled throughout the image.
- − Visual quality of the food photos is lower with some strange artifacts.
- − The layout is cluttered and contains excessive pricing columns that don't make sense.
Verdict: GPT Image 2 is significantly superior to Stable Diffusion 3.5 Medium, producing a professional-grade restaurant menu with flawless text, coherent typography, and high-quality photography. Stable Diffusion 3.5 Medium struggles with the 'Mains' and 'Appetizers' sections and fails completely on text legibility and graphic clarity.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
GPT Image 2
- + Excellent adherence to the 'exploded' view requirement with clearly separated, dynamic components.
- + Integrated text perfectly matches the requested fiery, glowing effect and starburst design.
- + High level of photorealistic detail in food textures, such as the sear on the patty and moisture on the vegetables.
- − The composition is very crowded, leaving little room for the background to breathe.
Stable Diffusion 3.5 Medium
- + Good background lighting and fire effects that create a strong sense of heat.
- + Clean, legible text for the price starburst and title.
- − Failed to produce an 'exploded' burger, showing an almost fully assembled burger instead.
- − The 'fiery, glowing effect' on the text was ignored in favor of flat white/black graphics.
- − Text rendering for 'LIMITED TIME ONLY' is slightly garbled.
Verdict: GPT Image 2 followed the prompt instructions much more closely, successfully delivering the requested 'exploded' burger layout and the specific fiery text effects. Stable Diffusion 3.5 Medium failed the core structural requirement of the prompt by presenting a static, assembled burger and utilized simple graphic overlays instead of the requested integrated glowing text.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
GPT Image 2
- + Perfect spelling and adherence to the specific text requested in the prompt.
- + Highly realistic chalk texture with natural variations in handwriting and dust details.
- + Exquisite composition with a warm, cozy café atmosphere reinforced by the lighting and wood framing.
- − The handwriting, while realistic, is quite consistent which might be seen as less 'messy' than real chalkboards.
Stable Diffusion 3.5 Medium
- + Captures a stylized, artistic chalk aesthetic with decorative flourishes.
- + Good contrast between the board and the wooden frame.
- − Severe spelling errors and nonsensical text throughout the entire board.
- − Failed to follow the Date and Price instructions accurately, including jumbled numbers.
- − The layout is cluttered and incoherent compared to the prompt's request for a clear menu structure.
Verdict: GPT Image 2 is the clear winner as it followed every textual requirement of the prompt with perfect spelling and a highly convincing chalk-on-blackboard aesthetic. In contrast, Stable Diffusion 3.5 Medium failed significantly on text rendering, producing garbled words and incorrect dates that diverged heavily from the prompt instructions.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
GPT Image 2
- + Excellent adherence to the counter-intuitive instruction of the horse riding the human
- + High-quality textures on the space suit and lunar surface
- + Clever details like the saddle designed for a human back and the horse holding reins
- − The astronaut's hands/gloves have an incorrect number of fingers
- − The earth in the background is slightly blurry compared to the foreground
Stable Diffusion 3.5 Medium
- + Beautiful cinematic lighting and composition
- + Dynamic pose with a sense of motion in the horse's mane
- − Completely failed the negative constraint to put the horse on top
- − Anatomical issues with the horse's legs and the astronaut's leg placement
- − Low-resolution blurring on the earth's surface
Verdict: GPT Image 2 followed the complex prompt instruction perfectly, depicting the surreal sight of a horse riding an astronaut. Stable Diffusion 3.5 Medium fell into a common bias and placed the astronaut on top of the horse, failing the primary challenge of the prompt despite having a pleasant color palette.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
GPT Image 2
- + Excellent photorealism with convincing textures on the fur and leather jacket.
- + Natural composition that clearly shows both the capybara driving and the passenger using her phone.
- + Cinematic lighting and realistic background bokeh that captures the New York night atmosphere.
- − The capybara's paws look slightly human-like in their grip on the wheel.
Stable Diffusion 3.5 Medium
- + Strong adherence to the capybara's facial features and whiskers.
- + Good placement of the yellow taxi driver cap.
- − Missed the prompt instruction for the passenger to be looking at a phone.
- − The composition is a bit flat and the car interior looks less realistic than the competitor.
- − The paws are positioned awkwardly on top of the wheel rather than gripping it.
Verdict: GPT Image 2 is the clear winner as it followed every detail of the prompt, including the passenger's activity and the specific 'bored' expression. Stable Diffusion 3.5 Medium failed to include the passenger's phone and resulted in a less realistic, more staged-looking composition.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
GPT Image 2
- + Excellent typography with perfect spelling in all requested text fields.
- + Superior atmospheric lighting and intricate gothic detailing in the border and background.
- + Highly cohesive composition that feels like a professional invitation.
- − The jack-o-lantern is central but slightly less 'glowy' than the lanterns nearby.
Stable Diffusion 3.5 Medium
- + Successfully includes the twisted trees and parchment aesthetic.
- + Distinctive jack-o-lantern expressions with high contrast.
- − Multiple spelling errors in almost all text fields including 'Halloweeen' and 'Inviloween'.
- − The layout is less polished and lacks the cinematic lighting requested.
- − Text is poorly centered and various elements feel disconnected.
Verdict: GPT Image 2 provides a masterful execution of the prompt, delivering perfect text rendering and a rich, atmospheric gothic aesthetic. In contrast, Stable Diffusion 3.5 Medium struggles significantly with the text requirements and offers a much flatter, less professional composition. GPT Image 2's attention to detail in the border and the clever integration of NYC-themed architecture (the bridge/arches) makes it the clear winner.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
GPT Image 2
- + Perfect adherence to text requirements, including 'JAPAN', 'SUSHI', and the flag icon.
- + Highly detailed isometric diorama with realistic 3D textures and materials.
- + Excellent miniature aesthetic with a clean, professional finish.
- − The diorama base is more complex than the 'small' and 'minimal' request, though it fits the theme well.
Stable Diffusion 3.5 Medium
- + Successfully placed the requested text and used a solid blue background.
- + Clean, simple plate composition that adheres to the 'minimal' request.
- − Failed to include the requested flag icon.
- − The text rendering is messy, with characters overlapping and a strange 'i' on 'SUSHI'.
- − Lacks the 'isometric miniature diorama' feel, appearing more like a standard photo on a flat surface.
Verdict: GPT Image 2 perfectly executed the prompt, providing a high-quality isometric miniature with flawless text and the requested flag icon. Stable Diffusion 3.5 Medium struggled with the text rendering, missed the flag icon entirely, and failed to capture the diorama aesthetic requested in the prompt.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
GPT Image 2
- + Successfully included all four requested animals (dog, cat, bunny, fox).
- + Excellent lighting with clear god rays and rim lighting on the fur.
- + Dynamic composition with a sense of movement and 'tumbling' as requested.
- − The fox's front right paw is anatomically slightly off/distorted.
- − The butterflies appear somewhat flat against the background light.
Stable Diffusion 3.5 Medium
- + Vibrant colors and high-contrast lighting.
- + Very cute, expressive eyes on the animals.
- − Failed to include the requested baby bunny animal.
- − The fox has an extra leg/paw visible near its chest.
- − The kitten's ears are unusually large and pointed, resembling a different species.
Verdict: GPT Image 2 followed the prompt much more accurately by including all four specified animals, whereas Stable Diffusion 3.5 Medium missed the bunny entirely. Furthermore, GPT Image 2 achieved a much better sense of depth and realistic fur texture under the golden sunrise lighting, while the Stable Diffusion output had notable anatomical errors like extra limbs.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
GPT Image 2
- + Excellent typography with perfect spelling of the brand name and date.
- + Highly sophisticated vector emblem style with clean lines and balanced composition.
- + Effective use of shading and subtle paper texture to match the vintage prompt.
- − The steam is a bit more illustrative and less minimalist than it could be.
Stable Diffusion 3.5 Medium
- + Bold, high-contrast illustration style that stands out.
- + Interesting integration of a coffee cup at the bottom.
- − Failed to spell the name correctly, adding an extra 'r' in 'Florrian'.
- − Incorrect date 'Est 170' instead of '1720'.
- − The cloche dome is awkwardly shaped and overlaps poorly with the text box.
Verdict: GPT Image 2 followed every aspect of the prompt with high precision, delivering professional-grade typography and a cohesive layout. Stable Diffusion 3.5 Medium struggled significantly with the text rendering, misspelling the name and omitting a digit from the date, while also failing to produce a clean vector aesthetic.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
GPT Image 2
- + Excellent typography with perfect spelling and clear hierarchy.
- + Strict adherence to the 6-step timeline with accurate iconography for each phase.
- + Highly professional layout that mimics a real NASA technical infographic.
- − The style leans more toward detailed illustration than pure flat-vector.
- − Included additional sections (Crew, Landing Site) not explicitly requested in the prompt.
Stable Diffusion 3.5 Medium
- + Captures a more minimalist flat-vector aesthetic.
- + Uses the requested color palette effectively.
- − Text is largely gibberish and suffers from significant spelling errors.
- − Visual layout is cluttered and fails to communicate the 6-step sequence clearly.
- − Icons are abstract and do not represent the specific mission phases requested.
Verdict: GPT Image 2 (Model A) is the clear winner as it successfully creates a high-quality, legible, and accurate infographic documenting the Apollo 11 mission steps. Stable Diffusion 3.5 Medium (Model B) failed to follow the sequential instructions and produced unintelligible text and confusing visuals.
Explore each model
Stability AI's 2.5-billion parameter Multimodal Diffusion Transformer with improvements (MMDiT-X) text-to-image model optimized for consumer hardware, featuring improved image quality, typography, and complex prompt understanding