Head to head
Esc

Models · slot A

to navigate to pick

GPT Image 1 Mini OpenAI Stable Diffusion 3.5 Large Stability AI

Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.

GPT Image 1 Mini

24.9 arena score

#13 of 62 in Text-to-Image

Skill signature · Text-to-Image

Stable Diffusion 3.5 Large

22.8 arena score

#30 of 62 in Text-to-Image

Vote tally

Where the votes landed

GPT Image 1 Mini

75.0%

win rate

Ties

0.0%

Stable Diffusion 3.5 Large

25.0%

win rate

75.0% 0.0% ties 25.0%
Shared challenges 13

Challenge by challenge

The strongest take from each model on every shared challenge, with the AI judge's read.

Geometric Composition

Text-to-Image

“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”

GPT Image 1 Mini
Stable Diffusion 3.5 Large

AI Judge Analysis

GPT Image 1 Mini

  • + Perfect adherence to the spatial requirements of the prompt.
  • + Higher photographic realism with soft, natural lighting.
  • + Clean composition with a clear view of the plant through the glass.
  • The blue sphere appears slightly larger than a 'small' sphere.
  • The book is floating slightly above the glass rim rather than resting flat.

Stable Diffusion 3.5 Large

  • + High clarity and sharp details on the wooden surface and glass edges.
  • + Accurate interpretation of the plant being behind the cube.
  • Failed to place the red book on top of the cube, placing it underneath instead.
  • The lighting is harsh and direct rather than the requested 'soft window light'.

Verdict: GPT Image 1 Mini followed all spatial instructions, correctly placing the book on top of the cube and the sphere inside. Stable Diffusion 3.5 Large failed the primary layout task by placing the book under the sphere and cube, although it produced a very high-resolution image with sharp textures.

Candid Street Photography

Text-to-Image

“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”

GPT Image 1 Mini
Stable Diffusion 3.5 Large
100% wins 0% ties 0% wins

AI Judge Analysis

GPT Image 1 Mini

  • + Excellent shallow depth of field and bokeh
  • + Highly realistic skin texture and facial lighting
  • + Strong cinematic atmosphere with natural, muted colors
  • The white car in the background lacks the requested motion blur
  • Anatomical issues with how the man's hands are interacting with the rear wheel spokes

Stable Diffusion 3.5 Large

  • + Better adherence to the motion blur request for passing vehicles
  • + Captures the scale of a Japanese street with the bus and signage
  • + Vibrant colors and convincing wet pavement reflections
  • The 'rain' looks like static vertical lines rather than realistic droplets
  • The bicycle geometry is broken (seat post missing, frame alignment)
  • Overall image has a slightly AI-processed 'sheen' that ignores the 'no stylization' request

Verdict: GPT Image 1 Mini produces a much more convincing and high-quality portrait with superior skin textures and photographic depth, though it missed the specific request for motion blur. Stable Diffusion 3.5 Large followed more of the prompt instructions regarding the background elements, but failed on technical execution with a poorly rendered bicycle and unrealistic rain effects. GPT Image 1 Mini is the preferred choice for its realism and believable cinematic quality.

Fantasy Warrior

Text-to-Image

“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”

GPT Image 1 Mini
Stable Diffusion 3.5 Large
67% wins 0% ties 33% wins

AI Judge Analysis

GPT Image 1 Mini

  • + Excellent depiction of warm torchlight reflecting off the metal surfaces.
  • + Highly detailed skin texture with convincing dirt and aging.
  • + The ornate engraving on the plate armor is complex and aesthetically pleasing.
  • Missed the request for small beads in the braided hair.
  • The armor engraving lacks some of the physical depth/relief found in the competitor.

Stable Diffusion 3.5 Large

  • + Very crisp skin texture and striking, lifelike eyes.
  • + Excellent implementation of braided hair as requested.
  • + The 'battle-worn' aesthetic is strong with visible dirt and high-contrast armor detailing.
  • The 'warm torchlight' lighting is much weaker and less atmospheric than Model A.
  • Lacks the requested 'beads' in the hair.
  • The metal of the armor looks slightly flat or overly bright in some areas despite being battle-worn.

Verdict: Both models captured the essence of the prompt well, but GPT Image 1 Mini took a superior approach to lighting and atmosphere, creating a much more convincing 'torchlight' effect. Stable Diffusion 3.5 Large produced a sharper image with better hair braids and lifelike eyes, but the lighting felt more like generic daylight, and both models failed to include the requested beads in the braids.

Modern Clean Menu

Text-to-Image

“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”

GPT Image 1 Mini
Stable Diffusion 3.5 Large

AI Judge Analysis

GPT Image 1 Mini

  • + Perfect text rendering of the requested headers
  • + Clean and highly professional grid layout that matches the 'modern minimalist' prompt
  • + High-quality, distinct food photos that correspond to the sections
  • Lack of menu items or descriptions under the section headers
  • Slightly less 'vibrant' than a typical casual dining menu

Stable Diffusion 3.5 Large

  • + More comprehensive layout featuring actual menu pricing and item lines
  • + Includes a wider variety of food imagery in the grid
  • + Good use of white space and vertical design
  • Poor typography with significant spelling errors (e.g., 'APPETIZRS', 'MAIMAES')
  • The grid layout feels a bit cluttered and overlaps the main menu column
  • Fails to accurately follow the 'Pizza' section header requirement

Verdict: GPT Image 1 Mini produced a cleaner, more legible design that perfectly followed the text instructions for the headers, though it left the menu sections empty. Stable Diffusion 3.5 Large attempted a more detailed menu with prices, but suffered from significant spelling errors and a more chaotic layout. GPT Image 1 Mini is the better choice for a professional template due to its crisp typography.

Magic Burger Explosion: Fiery Photorealism Challenge

Text-to-Image

“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”

GPT Image 1 Mini
Stable Diffusion 3.5 Large

AI Judge Analysis

GPT Image 1 Mini

  • + Excellent integration of all three requested text elements with the correct glowing effect.
  • + Clear 'exploded' layout showing distinct mid-air suspension of layers.
  • + Very high photorealistic texture on the bun and patty.
  • The lighting on the burger is slightly flat compared to the intense background.
  • The splash of sauce is a bit minimal.

Stable Diffusion 3.5 Large

  • + Intense, dynamic lighting and high-quality fire effects.
  • + Rich, glistening textures on the melted cheese and patties.
  • Failed to include any of the requested text or the starburst element.
  • Does not show an 'exploded' view; the burger is mostly assembled.

Verdict: GPT Image 1 Mini is the superior choice because it fully adhered to the complex prompt instructions, including all specific text strings and the exploded composition. Stable Diffusion 3.5 Large produced a visually striking image with impressive fire effects, but it completely ignored the text requirements and the mid-air suspension of individual components.

Chalkboard Menu

Text-to-Image

“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”

GPT Image 1 Mini
Stable Diffusion 3.5 Large

AI Judge Analysis

GPT Image 1 Mini

  • + Excellent text rendering with perfect spelling of the entire requested menu.
  • + Highly realistic chalk texture with dusty, hand-drawn edges.
  • + Strict adherence to the 2026 date and specific price points.
  • The Title is in print capitals rather than the requested 'elegant cursive'.
  • The composition is a tight crop rather than showing the full 'cozy café' environment.

Stable Diffusion 3.5 Large

  • + Better environmental context showing the café interior and furniture.
  • + Includes decorative elements like chalk borders and frames.
  • Severe text artifacts and multiple spelling errors (e.g., 'TODAAY', 'Cholcalte', 'Cockies').
  • Failed the date requirement, showing 2024 instead of 2026.
  • The text style looks like a digital brush rather than authentic hand-lettered chalk.

Verdict: GPT Image 1 Mini is the clear winner due to its superior text generation and prompt adherence, correctly spelling every item on the menu and including the specific date requested. While Stable Diffusion 3.5 Large provides a better sense of a café atmosphere, it fails significantly on text legibility, spelling, and adherence to the specific content of the prompt.

The Reversed Rodeo

Text-to-Image

“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”

GPT Image 1 Mini
Stable Diffusion 3.5 Large

AI Judge Analysis

GPT Image 1 Mini

  • + Excellent anatomical detail on the horse
  • + High visual clarity and cinematic lighting
  • + Balanced composition with a clear focal point
  • Failed the negative constraint/logic check as the astronaut is on top, not the horse

Stable Diffusion 3.5 Large

  • + Dynamic sense of movement and scale with the planet background
  • + Highly detailed spacesuit and horse tack textures
  • + Good use of color and atmospheric effects
  • Failed the specific positional instruction that the horse should be on top
  • The horse's muzzle and bridle area are slightly garbled

Verdict: Both GPT Image 1 Mini and Stable Diffusion 3.5 Large completely failed the complex spatial instruction to put the horse on top of the astronaut, defaulting instead to the standard 'astronaut riding a horse' trope. Stable Diffusion 3.5 Large is the better image overall due to its superior scale, vibrant color palette, and more complex cinematic environment compared to the simpler composition of GPT Image 1 Mini.

The Capybara Taxi Driver

Text-to-Image

“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”

GPT Image 1 Mini
Stable Diffusion 3.5 Large

AI Judge Analysis

GPT Image 1 Mini

  • + Successfully includes the human passenger in the back seat looking at a phone.
  • + The capybara's expression and posture perfectly match the 'professional driver' request.
  • + Excellent cinematic lighting that feels photorealistic for a night scene.
  • The capybara only has one paw clearly on the steering wheel, while the prompt requested both.

Stable Diffusion 3.5 Large

  • + The foreground capybara has very sharp, detailed fur texture.
  • + Vibrant colors on the clothing and car exterior.
  • Completely failed to include the human businesswoman in the back seat.
  • The anatomy of the capybara's hands/paws is distorted and unrealistic.
  • The capybara appears to be wearing human pants, which was not requested.

Verdict: GPT Image 1 Mini followed the prompt instructions much more accurately, successfully capturing all requested elements including the passenger and the relative mood. Stable Diffusion 3.5 Large failed to include the secondary subject entirely and produced anatomical irregularities in the capybara's limbs.

The Halloween Invitation

Text-to-Image

“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”

GPT Image 1 Mini
Stable Diffusion 3.5 Large

AI Judge Analysis

GPT Image 1 Mini

  • + Perfect text rendering for all requested details.
  • + Clean and cohesive vintage grainy aesthetic.
  • + Excellent layout balance with strong central focus.
  • Lighting is a bit flat compared to the requested 'cinematic' style.
  • The border is very subtle, almost blending into the darkness.

Stable Diffusion 3.5 Large

  • + Highly detailed parchment texture with torn edges.
  • + Dynamic composition with a bright moon and twisted trees.
  • + Atmospheric lighting and color contrast.
  • Failed to include the date, time, and location details entirely.
  • The text in the scroll banner contains gibberish characters.
  • Incorrect layout with the jack-o-lantern not being the central focus.

Verdict: GPT Image 1 Mini is the superior choice because it followed all text-related instructions perfectly, including the specific date and location. While Stable Diffusion 3.5 Large has a more visually striking artistic style, it failed significantly on prompt adherence by omitting the event details and rendering garbled text.

Isometric Miniature Diorama Scenes

Text-to-Image

“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”

GPT Image 1 Mini
Stable Diffusion 3.5 Large

AI Judge Analysis

GPT Image 1 Mini

  • + Excellent adherence to the 'cartoon' and 'soft refined texture' style requirements.
  • + Perfectly rendered text and flag exactly as requested in the prompt.
  • + Ultra-clean composition with the specified solid light blue background and centered layout.
  • The textures are perhaps too simplified or 'plastic' to be considered realistic PBR materials.

Stable Diffusion 3.5 Large

  • + Higher detail in sushi modeling and more complex arrangement.
  • + Uses a more sophisticated dark diorama base with realistic surface textures.
  • + Includes more variety of sushi types.
  • Failed to place the text 'at top-center' as requested, opting for a 3D sign in the scene.
  • The composition feels cluttered and lacks the 'minimal' aesthetic requested.
  • Chopsticks are floating awkwardly and the background has a textured grain rather than being 'solid'.

Verdict: GPT Image 1 Mini followed the layout and text instructions perfectly, providing a clean, aesthetic isometric scene that matches the 'cartoon' and 'minimal' keywords. Stable Diffusion 3.5 Large struggled with the text placement and the requirement for a solid background, resulting in a cluttered scene with some floating artifacts.

Adorable Baby Animals in Sunny Meadow

Text-to-Image

“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”

GPT Image 1 Mini
Stable Diffusion 3.5 Large
67% wins 0% ties 33% wins

AI Judge Analysis

GPT Image 1 Mini

  • + Excellent anatomical accuracy for all four animals.
  • + Rich, tactile fur texture and clear, expressive eyes.
  • + Clearer rendering of the 'god rays' and sunrise lighting mentioned in the prompt.
  • Composition feels a bit crowded towards the edges.
  • Butterflies appear slightly flat compared to the animals.

Stable Diffusion 3.5 Large

  • + Dynamic composition with a nice sense of movement and 'tumbling'.
  • + Good use of bokeh and depth of field in the foreground/background.
  • + Inclusion of plenty of butterflies to match the 'playfully chasing' prompt.
  • The kitten has anatomically incorrect large, pointed fox-like ears.
  • Lower overall sharpness and fine detail in the fur textures.
  • Lighting feels a bit washed out in the center.

Verdict: GPT Image 1 Mini is the winner due to its superior anatomical accuracy and high-fidelity textures, whereas Stable Diffusion 3.5 Large struggled with the kitten's anatomy, giving it fox-like features. Both models followed the prompt well, but GPT Image 1 Mini's lighting and clarity felt more like the requested '8K masterpiece'.

Vintage Cafe Logo

Text-to-Image

“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”

GPT Image 1 Mini
Stable Diffusion 3.5 Large

AI Judge Analysis

GPT Image 1 Mini

  • + Excellent typography and spelling of the main brand name.
  • + Precise execution of the cloche dome illustration with clean lines.
  • + High-contrast vector style that would work well for physical branding.
  • Ignored the request for a light background, providing a black one instead.
  • The texture is very subtle, almost appearing as noise.

Stable Diffusion 3.5 Large

  • + Followed the background and color palette instructions perfectly.
  • + Includes the subtle paper-like texture and corner flourishes requested in the vintage prompt.
  • + Captures a more authentic 'logo on menu' aesthetic.
  • Spelling error in the brand name ('Cafféé' instead of 'Caffè').
  • The cloche dome illustration is slightly messy and disconnected.
  • The banner contains the name instead of the date, reversing the prompt request.

Verdict: GPT Image 1 Mini produced a much cleaner and more professional graphic with correct spelling, but it completely failed the instruction for a light background. Stable Diffusion 3.5 Large captured the requested aesthetic, colors, and texture perfectly, but it introduced a spelling error and failed to follow the specific banner placement instructions. GPT Image 1 Mini is the likely winner as spelling and graphic clarity are typically more critical for a logo identity.

Apollo 11: Journey to Tranquility

Text-to-Image

“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”

GPT Image 1 Mini
Stable Diffusion 3.5 Large

AI Judge Analysis

GPT Image 1 Mini

  • + Excellent text rendering with no spelling errors
  • + Accurately follows the specific 6-step logical sequence requested
  • + Clean, consistent vector style that matches the modern infographic prompt
  • The 'Translunar' icon is just a scribble and doesn't clearly represent a trajectory
  • The composition is missing the 'crew' details at the bottom of the frame

Stable Diffusion 3.5 Large

  • + Strong NASA-inspired color palette and atmospheric feel
  • + Good use of space with detailed lunar surface illustration
  • Fails to follow the 6-step instruction entirely
  • Contains significant text gibberish and artifacts
  • Depicts a space shuttle instead of the requested Saturn V rocket

Verdict: GPT Image 1 Mini is the clear winner as it successfully interprets the multi-step technical instructions and renders legible, accurate text for all steps. Stable Diffusion 3.5 Large fails on the prompt adherence, generating a confusing layout with nonsensical text and an incorrect spacecraft (a space shuttle instead of the Apollo 11 Saturn V).

Next steps

Explore each model