Head to head
Esc

Models · slot A

to navigate to pick

GPT Image 1 Mini OpenAI Stable Diffusion 3.5 Medium Stability AI

Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.

GPT Image 1 Mini

24.9 arena score

#13 of 62 in Text-to-Image

Skill signature · Text-to-Image

Stable Diffusion 3.5 Medium

16.8 arena score

#56 of 62 in Text-to-Image

Vote tally

Where the votes landed

GPT Image 1 Mini

100.0%

win rate

Ties

0.0%

Stable Diffusion 3.5 Medium

0.0%

win rate

100.0% 0.0% ties 0.0%
Shared challenges 13

Challenge by challenge

The strongest take from each model on every shared challenge, with the AI judge's read.

Geometric Composition

Text-to-Image

“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”

GPT Image 1 Mini
Stable Diffusion 3.5 Medium

AI Judge Analysis

GPT Image 1 Mini

  • + Excellent adherence to the perspective and lighting instructions.
  • + The glass cube has realistic thickness and reflections.
  • + High visual quality with a pleasing shallow depth of field.
  • The plant is behind the cube but not clearly visible through the glass itself, more so above and around it.

Stable Diffusion 3.5 Medium

  • + The sphere has a nice glass-like texture.
  • + The plant is clearly visible through the glass as requested.
  • The sphere is floating unnaturally in the center of the cube without support.
  • The red book looks more like a flat felt pad or a very thin board than a realistic book.
  • The lighting is somewhat harsh and overexposed in the background.

Verdict: GPT Image 1 Mini produced a much more realistic and aesthetically pleasing image with a solid understanding of physics and object properties. While Stable Diffusion 3.5 Medium followed the instruction to show the plant through the glass better, it failed on the rendering of the book and the logic of the sphere's placement.

Candid Street Photography

Text-to-Image

“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”

GPT Image 1 Mini
Stable Diffusion 3.5 Medium

AI Judge Analysis

GPT Image 1 Mini

  • + Excellent natural skin texture and fine detail on the face and hands.
  • + Realistic rendering of water droplets on the bicycle seat.
  • + High quality 50mm look with professional-grade shallow depth of field.
  • The bike anatomy is slightly awkward where his hand meets the frame.
  • Lacks significant motion blur from cars mentioned in the prompt.

Stable Diffusion 3.5 Medium

  • + Captures the 'candid street photo' aesthetic with wider framing.
  • + Better reflections on the wet pavement.
  • + The colors across the scene are more vibrant and street-like.
  • Anatomical failure with the hands merged into the handlebars.
  • Stronger AI artifacts on the face and clothing textures.
  • Bicycle geometry is nonsensical, particularly the pedals and rear hub area.

Verdict: GPT Image 1 Mini is the clear winner due to its superior technical quality and anatomical correctness, particularly in the rendering of the man's face and hands. While Stable Diffusion 3.5 Medium captured the environmental reflections and framing better, the severe distortions in the man's hands and the bicycle's structure make it a less successful image.

Fantasy Warrior

Text-to-Image

“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”

GPT Image 1 Mini
Stable Diffusion 3.5 Medium

AI Judge Analysis

GPT Image 1 Mini

  • + Excellent depiction of ornate engraved plate armor with weathering.
  • + Beautiful warm torchlight lighting and natural skin textures.
  • + Great composition with a focused, battle-worn expression.
  • Missed the request for beads in the hair.
  • Cloth and leather details are less prominent due to the framing.

Stable Diffusion 3.5 Medium

  • + Stronger adherence to the braided hair and beads request.
  • + High-intensity lighting and color contrast create a dramatic scene.
  • + Visible detail in the leather and quilted underlayer.
  • The face has a slightly plastic or oil-slicked sheen rather than natural skin.
  • Anatomical anomalies in the way the hair connects to the scalp and face.
  • The bokeh sparks appear as flat yellow dots rather than integrated light.

Verdict: GPT Image 1 Mini produced a more cohesive and realistic image with superior armor texture and natural lighting, while Stable Diffusion 3.5 Medium followed the specific detail prompts for beads and braids more closely but failed in anatomical realism. Overall, GPT Image 1 Mini is the better image due to its believable skin texture and professional cinematic quality.

Modern Clean Menu

Text-to-Image

“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”

GPT Image 1 Mini
Stable Diffusion 3.5 Medium

AI Judge Analysis

GPT Image 1 Mini

  • + Excellent typography with clean, readable, bold sans-serif fonts.
  • + Perfectly organized grid layout that aligns with the requested sections.
  • + High-quality food photography with vibrant colors and professional lighting.
  • Lack of actual menu items/text descriptions under the section headers.
  • The orange accent bar at the bottom is a bit simplistic compared to a full professional design.

Stable Diffusion 3.5 Medium

  • + Includes realistic pricing and item descriptions, making it feel like a complete document.
  • + Good use of multiple images spread across the layout.
  • + Captures a casual dining aesthetic well.
  • Text consists of illegible gibberish characters.
  • Visual quality of the text is blurry and poorly rendered.
  • The layout feels cluttered and lacks the 'modern minimalist' requested.

Verdict: GPT Image 1 Mini produced a much cleaner and more professional design that perfectly adheres to the 'modern minimalist' aesthetic, although it failed to populate the menu with specific items. Stable Diffusion 3.5 Medium attempted a more complex layout, but the text is completely illegible and the overall coherence of the design is poor. GPT Image 1 Mini is the clear winner for its superior visual quality and adherence to the design brief.

Magic Burger Explosion: Fiery Photorealism Challenge

Text-to-Image

“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”

GPT Image 1 Mini
Stable Diffusion 3.5 Medium

AI Judge Analysis

GPT Image 1 Mini

  • + Excellent adherence to the 'exploded' layout request.
  • + Perfect text rendering with high readability and the requested glowing effect.
  • + Consistent artistic style across all elements including the price starburst.
  • Lighting on the bottom bun feels slightly flat compared to the top elements.

Stable Diffusion 3.5 Medium

  • + High resolution texture on the bun and patty.
  • + Dynamic fire background provides a strong sense of heat.
  • Failed the 'exploded burger' layout, showing a mostly assembled burger.
  • Text is small, plain, and lacks the requested glowing fiery effect.
  • Text 'LIMITED TIME ONLY' has a minor character overlap/artifact.

Verdict: GPT Image 1 Mini followed the prompt instructions precisely, delivering a clear 'exploded' view of the burger and high-quality, glowing text that looks like a finished advertisement. Stable Diffusion 3.5 Medium produced a high-quality image of a burger, but failed on the specific layout and text styling requirements of the prompt.

Chalkboard Menu

Text-to-Image

“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”

GPT Image 1 Mini
Stable Diffusion 3.5 Medium

AI Judge Analysis

GPT Image 1 Mini

  • + Excellent text rendering with perfect spelling and adherence to the prompt's specific menu items.
  • + Realistic chalk texture on the letters and accurate menu board framing.
  • + Consistent handwriting style that feels authentic to a café environment.
  • The 'elegant cursive' requirement for the title was not fully met, as it is a slanted print style.
  • The handwriting looks slightly too uniform, bordering on a digital font appearance.

Stable Diffusion 3.5 Medium

  • + Successfully captured a more artistic, stylized chalk aesthetic with flourishes.
  • + Good representation of a physical café environment with the wood panel background.
  • Severe spelling errors and illegible text across the entire board.
  • Completely failed to follow the specific date and price details in the prompt.
  • Layout is messy and cluttered compared to the clean request.

Verdict: GPT Image 1 Mini is the clear winner as it virtually follows every text instruction perfectly, rendering the specific menu items and prices without a single typo. While it struggled slightly to produce a true 'cursive' title, Stable Diffusion 3.5 Medium produced mostly illegible gibberish that failed to adhere to the prompt's specific content.

The Reversed Rodeo

Text-to-Image

“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”

GPT Image 1 Mini
Stable Diffusion 3.5 Medium

AI Judge Analysis

GPT Image 1 Mini

  • + Excellent cinematic lighting and atmospheric depth.
  • + High level of texture detail on both the spacesuit and horse's coat.
  • Completely failed the negative constraint to put the horse on top.
  • Standard composition lacks the requested surrealism.

Stable Diffusion 3.5 Medium

  • + Successfully placed the planet below to give a sense of scale.
  • + Crisp colors and good contrast between the subject and space.
  • Completely failed the negative constraint to put the horse on top.
  • Anatomical issues with the horse's legs and the astronaut's leg placement.

Verdict: Both GPT Image 1 Mini and Stable Diffusion 3.5 Medium failed the specific negative constraint of placing the horse on top of the astronaut, providing a standard 'astronaut riding a horse' image instead. GPT Image 1 Mini is the winner because it offers significantly higher visual quality, better textures, and a more coherent, cinematic composition compared to the anatomical errors present in the Stable Diffusion output.

The Capybara Taxi Driver

Text-to-Image

“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”

GPT Image 1 Mini
Stable Diffusion 3.5 Medium

AI Judge Analysis

GPT Image 1 Mini

  • + Excellent photorealistic lighting and texture on the capybara's fur.
  • + Strict adherence to the businesswoman's interaction with her phone and her bored expression.
  • + Stronger cinematic composition that feels like a real film still.
  • The capybara's 'front paws' look slightly like human hands covered in fur.

Stable Diffusion 3.5 Medium

  • + Accurately depicts both paws on the steering wheel as requested.
  • + Bright, clear colors that match the 'yellow taxi' theme well.
  • The businesswoman is looking at the camera rather than being bored with her phone.
  • Lower photorealistic quality; the woman's face looks slightly AI-generated and waxy.
  • The capybara's anatomical placement in the seat looks a bit awkward.

Verdict: GPT Image 1 Mini captured the intended mood and specific prompt details, such as the businesswoman's bored interaction with her phone, much more effectively than Stable Diffusion 3.5 Medium. While Stable Diffusion 3.5 Medium followed the 'both paws' instruction more literally, GPT Image 1 Mini's superior photorealism and character directing make it the more compelling image.

The Halloween Invitation

Text-to-Image

“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”

GPT Image 1 Mini
Stable Diffusion 3.5 Medium

AI Judge Analysis

GPT Image 1 Mini

  • + Perfect text rendering for all requested copy
  • + Atmospheric and cinematic lighting
  • + Cleanly integrated gothic border and scroll banner
  • Grainy texture might be too heavy for some professional prints
  • Minimal variation in color palette

Stable Diffusion 3.5 Medium

  • + Vibrant colors and distinct parchment texture
  • + Good composition with multiple jack-o-lanterns and trees
  • Significant spelling errors throughout the text
  • Failed to include the specific banner scroll requested
  • Missing the requested date and time accuracy

Verdict: GPT Image 1 Mini followed the prompt perfectly, delivering accurate text and a cohesive, moody aesthetic suitable for an invitation. Stable Diffusion 3.5 Medium struggled with spelling and failed to include several key details from the prompt, resulting in a less professional output.

Isometric Miniature Diorama Scenes

Text-to-Image

“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”

GPT Image 1 Mini
Stable Diffusion 3.5 Medium

AI Judge Analysis

GPT Image 1 Mini

  • + Excellent typography with perfect spelling and placement.
  • + Clean 3D isometric perspective with high-quality PBR-style textures.
  • + Includes all requested elements like the flag icon and custom diorama base.
  • The textures are slightly more toy-like than realistic for PBR.

Stable Diffusion 3.5 Medium

  • + Successfully captures a 45-degree isometric angle for the plate.
  • + Good depiction of textures on the nori and roe toppings.
  • Failed to include the requested flag icon.
  • Significant text errors with the 'SUSHI' lettering appearing distorted.
  • Missing the raised diorama base requested in the prompt.

Verdict: GPT Image 1 Mini followed the prompt instructions near-perfectly, delivering clean typography, the specific flag icon, and a high-quality 3D diorama aesthetic. Stable Diffusion 3.5 Medium struggled with the text rendering and omitted several key details like the flag and the raised base, resulting in a less polished image.

Adorable Baby Animals in Sunny Meadow

Text-to-Image

“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”

GPT Image 1 Mini
Stable Diffusion 3.5 Medium

AI Judge Analysis

GPT Image 1 Mini

  • + Successfully included all four requested animals (dog, cat, rabbit, fox).
  • + Each animal has distinct, realistic textures and anatomical features.
  • + Dynamic posing captures the 'playfully chasing' and 'tumbling' aspect of the prompt.

Stable Diffusion 3.5 Medium

  • + Vibrant colors with high saturation that feel very 'wholesome'.
  • + Nice butterfly details and clear lighting effects.
  • Failed to include the requested baby bunny.
  • The kitten has odd, fox-like ears and unnatural proportions.
  • The fur texture looks more like a digital painting than a photorealistic 8K masterpiece.

Verdict: GPT Image 1 Mini adhered perfectly to the prompt by including all four animals with realistic fur textures and dynamic action poses. Stable Diffusion 3.5 Medium failed to include the rabbit and produced a more stylized, less photorealistic result with anatomical inconsistencies on the kitten.

Vintage Cafe Logo

Text-to-Image

“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”

GPT Image 1 Mini
Stable Diffusion 3.5 Medium
100% wins 0% ties 0% wins

AI Judge Analysis

GPT Image 1 Mini

  • + Excellent typography with correct spelling and accent marks.
  • + Clean vector-style execution with accurate inclusion of all prompt elements.
  • + Clear depiction of the steam and cloche dome as requested.
  • Failed to provide a 'light background' as explicitly requested.
  • The texture is very subtle, almost appearing as noise rather than paper texture.

Stable Diffusion 3.5 Medium

  • + Beautifully followed the 'light background' and 'warm cream tones' part of the prompt.
  • + Elegant illustration style that fits the 'vintage' aesthetic perfectly.
  • + Strong texture and artistic rendering.
  • Significant spelling errors in both the restaurant name and the establishment date.
  • The 'steam' is rendered as a strange tuft or swirl on top of the dome rather than vapor.
  • Text alignment is slightly cluttered within the circular frame.

Verdict: GPT Image 1 Mini is the superior choice for a logo because it maintains perfect spelling and a clean, usable vector layout, even though it ignored the light background instruction. Stable Diffusion 3.5 Medium captured the requested color palette and texture much better, but suffers from disqualifying typos in the brand name ('Florrian') and the date ('170' instead of '1720').

Apollo 11: Journey to Tranquility

Text-to-Image

“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”

GPT Image 1 Mini
Stable Diffusion 3.5 Medium

AI Judge Analysis

GPT Image 1 Mini

  • + Excellent typography with perfect spelling of all technical terms.
  • + Clear, clean vector aesthetic that perfectly matches the requested flat-vector style.
  • + Accurate and distinct iconography for each phase of the mission.
  • The 'Translunar' trajectory path is a bit messy and overlaps its own text.
  • The image is cropped at the bottom, cutting off the crew and lunar base information.

Stable Diffusion 3.5 Medium

  • + Stronger use of gradients and cosmic background textures for an 'outer space' feel.
  • + The composition uses the space well to suggest a grand scale.
  • Nonsensical text and spelling errors throughout the infographic.
  • Failed to follow the specific iconography requests for several steps.
  • The sequential flow of the mission is disorganized and difficult to follow.

Verdict: GPT Image 1 Mini followed the technical requirements of the prompt with high precision, delivering legible text and clear iconography in the requested flat-vector style. In contrast, Stable Diffusion 3.5 Medium struggled with text rendering and the logical flow of the infographic, resulting in a visually messy and uninformative poster. Despite much of the bottom being cropped, GPT Image 1 Mini is the clear winner for its adherence to the core instructional steps.

Next steps

Explore each model