Stability AI's 8.1-billion parameter Multimodal Diffusion Transformer (MMDiT) text-to-image model featuring improved image quality, typography, complex prompt understanding, and resource-efficiency
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
Stable Diffusion 3.5 Large
#30 of 62 in Text-to-Image
Stable Diffusion 3.5 Medium
#56 of 62 in Text-to-Image
Where the votes landed
Stable Diffusion 3.5 Large
0%
win rate
Ties
0%
Stable Diffusion 3.5 Medium
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
Stable Diffusion 3.5 Large
- + Excellent photo-realistic lighting and reflections on the glass and wood.
- + Sharp, high-resolution textures on the book and table.
- + Coherent geometric structure and shadows.
- − Failed the spatial positioning: the red book is inside the cube rather than on top of it.
- − The blue sphere is on top of the book instead of inside the cube directly.
Stable Diffusion 3.5 Medium
- + Perfect adherence to spatial instructions including 'book on top' and 'sphere inside'.
- + Good implementation of the requested lighting and background plant.
- + Clean, minimalist composition.
- − Lower overall visual fidelity compared to the larger model.
- − The sphere appears to be floating without a physical support.
- − The cube has slightly warped edges on the bottom left.
Verdict: Stable Diffusion 3.5 Large (Image A) produces a much more realistic and aesthetically pleasing image, but it fails significantly on the spatial logic of the prompt. Stable Diffusion 3.5 Medium (Image B) correctly interprets all prepositional relationships ('book on top', 'sphere inside'), making it the winner for prompt adherence despite its slightly lower visual quality.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
Stable Diffusion 3.5 Large
- + Excellent anatomical detail in the subject's skin texture and hands
- + Realistic lighting and convincing wet pavement reflections
- + High technical clarity with a consistent shallow depth of field
- − The cars in the background are static rather than having the requested motion blur
- − Composition feels a bit centered and tidy for an 'imperfect framing' prompt
Stable Diffusion 3.5 Medium
- + Successfully captured the 'imperfect framing' and candid street photography vibe
- + Rich color palette and more atmospheric background lighting
- + Good sense of wetness on the pavement and surroundings
- − Significant anatomy issues with the hands appearing distorted
- − The bicycle geometry is illogical, particularly the frame and handlebars
- − Lacks the requested motion blur from passing cars
Verdict: Stable Diffusion 3.5 Large is the winner due to its superior anatomical and technical rendering; the man and the bicycle are coherent and detailed. While Stable Diffusion 3.5 Medium captured a better 'street photography' composition, it failed significantly on the structural details of the hands and the bicycle.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
Stable Diffusion 3.5 Large
- + Excellent intricate engraving on the plate armor
- + More realistic skin texture and subtle grime
- + Clearer depiction of cloth underlayers and chainmail
- − Missed the request for beads in the hair braids
- − Lighting feels a bit flat compared to the requested warm torchlight
Stable Diffusion 3.5 Medium
- + Dynamic warm lighting that strongly reflects off the metallic surfaces
- + Successfully included numerous small beads in the complex braids
- + More intense and lifelike eye rendering
- − Armor engravings are less sharp and slightly muddy in appearance
- − Overall image has a more 'digital' and high-contrast saturation that looks less natural than A
Verdict: Stable Diffusion 3.5 Large (Image A) produces a more grounded and anatomically coherent portrait with superior armor detail, though it missed the specific detail of beads in the hair. Stable Diffusion 3.5 Medium (Image B) captured the lighting, beads, and intensity of the prompt better, but suffers from slightly less refined textures on the metal and clothing.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
Stable Diffusion 3.5 Large
- + Excellent typography style with bold, professional sans-serif fonts
- + High-quality, vibrant food photography
- + Clean minimalist layout with strong vertical balance
- − Text is largely gibberish or misspelled (e.g., 'MAIMAES', 'APPETIZRS')
- − The food grid is cropped on the sides rather than integrated into a full single-page layouts
Stable Diffusion 3.5 Medium
- + Better logic in layout with prices aligned next to items
- + Includes more variety in dish types and photos
- + Comprehensive single-page menu structure
- − Font choice is a strange, semi-stylized sans-serif that looks less professional
- − Visual fidelity of the food images is slightly lower/noisier than Model A
- − The text rendering is very poor with several illegible characters
Verdict: Stable Diffusion 3.5 Large (Image A) produces a much more visually striking and professional-looking aesthetic that fits the 'modern minimalist' prompt perfectly, despite the nonsense text. Stable Diffusion 3.5 Medium (Image B) has a more functional layout for an actual menu, but the font choices and overall clarity are significantly weaker.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
Stable Diffusion 3.5 Large
- + Excellent photorealistic texture on the meat and bun.
- + Dynamic lighting with realistic fire and ember effects.
- − Completely failed to include any of the requested text.
- − Components are stacked rather than 'exploded' or suspended as separate mid-air pieces.
Stable Diffusion 3.5 Medium
- + Successfully integrated all requested text with mostly correct spelling.
- + Includes the starburst element for the price as specified.
- + Displays high-quality food photography aesthetics.
- − Failed the 'exploded' instruction, showing a mostly assembled burger.
- − Text is plain white/black rather than the requested 'fiery, glowing effect'.
Verdict: Stable Diffusion 3.5 Large produced a more visually striking image with superior lighting and texture, but it completely ignored the text requirements. Stable Diffusion 3.5 Medium is the preferred choice for this specific task because it followed the complex prompt instructions to include specific titles, prices, and starburst elements, despite the burger itself being less 'exploded' than requested.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
Stable Diffusion 3.5 Large
- + Excellent composition and environmental context, showing the café setting.
- + Consistent chalk aesthetic with realistic textures.
- + Correct spelling of 'APRIL' and mostly legible text layout.
- − Spells 'TODAY' as 'TODAAY'.
- − Failed the year prompt, rendering 2024 instead of 2026.
- − The specific menu items requested are scrambled or truncated.
Stable Diffusion 3.5 Medium
- + Captures a very artistic chalk texture with smudges and flourishes.
- + Attempted the cursive style mentioned in the prompt.
- − Text is largely illegible/gibberish (e.g., 'Trodalu', 'Oocclutels').
- − The date is a mess of overlapping characters.
- − Poor layout with text overlapping and inconsistent sizing.
Verdict: Stable Diffusion 3.5 Large produced a much more coherent and realistic café scene with better overall text legibility, despite minor spelling errors and getting the year wrong. Stable Diffusion 3.5 Medium struggled significantly with the text rendering, producing mostly nonsensical words and a cluttered layout.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
Stable Diffusion 3.5 Large
- + Excellent cinematic lighting and atmospheric detail
- + Natural composition with a dynamic sense of motion
- + High-quality texture on the space suit and horse hair
- − Failed to follow the spatial instruction of placing the horse on top of the astronaut
Stable Diffusion 3.5 Medium
- + Strong image clarity and sharp focus
- + Clean, high-contrast colors against the black space background
- − Failed to follow the spatial instruction of placing the horse on top of the astronaut
- − Anatomical issues with the horse's legs appearing elongated and disjointed
- − The astronaut's legs and the horse's saddle area have significant clipping and merging errors
Verdict: Both Stable Diffusion 3.5 Large and Stable Diffusion 3.5 Medium failed the negative constraint of placing the horse on top of the astronaut, defaulting to the common trope of an astronaut riding a horse. Stable Diffusion 3.5 Large is the superior image due to its cinematic quality and realistic rendering, whereas Stable Diffusion 3.5 Medium suffers from significant anatomical distortions and clipping artifacts.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
Stable Diffusion 3.5 Large
- + Excellent texture on the fur and leather jacket
- + Strong color saturation and clear night-time bokeh
- + Accurate depiction of a capybara's facial features
- − Completely failed to include the human businesswoman in the back seat
- − Anatomical issues with the capybara's body and paws
Stable Diffusion 3.5 Medium
- + Followed complex prompt instructions regarding the scene composition
- + Successfully included the bored businesswoman in the background
- + Great capybara expression and pose with paws on the wheel
- − Slightly lower resolution and clarity compared to model A
- − The businesswoman is not looking at her phone as requested
Verdict: Stable Diffusion 3.5 Medium is the winner because it successfully captured the primary narrative element of the prompt: the contrast between the capybara driver and the bored passenger. While Stable Diffusion 3.5 Large has better individual texture and lighting, it completely ignored the instruction to include a person in the back seat.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
Stable Diffusion 3.5 Large
- + Excellent layout with a cinematic, moody atmosphere.
- + The typography for 'Halloween Party' and the banner is highly legible and follows the aesthetic instructions.
- + Impressive detailing on the weathered parchment edges and delicate border illustration.
- − Completely failed to include the event details at the bottom (Date, Time, Location).
- − The jack-o-lanterns are secondary to the moon, despite the prompt asking for a central one.
- − Text at the very bottom of the scroll degrades into gibberish.
Stable Diffusion 3.5 Medium
- + Successfully attempted to include all required text elements, including date and location.
- + Effective use of the square format and bold jack-o-lantern placement.
- + Good color palette that matches the 'moody night sky' requirement.
- − Significant spelling errors throughout the text (e.g., 'Halloweeen Inviloween', 'Timme', 'Loccation').
- − The font choice for the body text feels a bit modern compared to the gothic request.
- − The scroll banner is absent, with text just placed on the main parchment background.
Verdict: Stable Diffusion 3.5 Large (Image A) produces a much more professional and aesthetically pleasing design with superior lighting and atmospheric details, but it ignores the specific event details requested. Stable Diffusion 3.5 Medium (Image B) is more diligent about including all the prompt's data points but fails on basic spelling and overall artistic composition. Image A is the preferred base for a designer, as its aesthetic quality is far higher even with missing text.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
Stable Diffusion 3.5 Large
- + Excellent texture rendering for the rice and fish materials
- + Successfully includes almost all prompt elements like the flag icon and multiple sushi types
- + High clarity and detail in the isometric diorama style
- − Text is on a sign rather than top-center of the overall canvas as requested
- − Includes more garnish and items than the 'minimal' request
Stable Diffusion 3.5 Medium
- + Correctly places text at the top-center of the image
- + Simple and clean composition that adheres to the minimal request
- + Accurate 45-degree angle
- − Missed the flag icon requirement
- − Rice texture appears slightly blurry or lower resolution compared to Model A
- − Text rendering has slight artifacts around the edges
Verdict: Stable Diffusion 3.5 Large (Model A) provides much better material quality and detail for the sushi, capturing the 'PBR materials' request effectively, though it missed the text placement. Stable Diffusion 3.5 Medium (Model B) followed the layout and text placement instructions better but failed to include the flag and yielded a softer, less detailed image.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
Stable Diffusion 3.5 Large
- + Successfully includes all four requested animals: puppy, kitten, bunny, and fox kit.
- + Dynamic composition with a sense of motion and 'tumbling' as requested.
- + Beautiful soft lighting and bokeh that matches the 'dew sparkles' and sunrise atmosphere.
- − The fox and kitten look very similar in facial structure.
- − Anatomical issues with the animals' paws in the foreground.
Stable Diffusion 3.5 Medium
- + Sharper texture on the fur and flowers.
- + Vibrant lighting with clear god rays and a sunrise feel.
- − Failed to include the requested baby bunny animal.
- − The characters are mostly static rather than 'tumbling' or 'playfully chasing'.
- − Oversaturated colors feel less photorealistic and more like a greeting card.
Verdict: Stable Diffusion 3.5 Large followed the complex prompt much more accurately by including all four distinct animals and capturing the action of them tumbling through the field. While Stable Diffusion 3.5 Medium has higher relative sharpness, it failed to generate the bunny and resulted in a more static, less imaginative composition.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
Stable Diffusion 3.5 Large
- + Excellent typography adherence with almost perfect spelling and style.
- + Captures the vector emblem aesthetic perfectly with clean lines.
- + Properly includes all requested elements like the 'Est. 1720' text and cloche handle/steam.
- − Spelling error 'Cafféé' with an extra 'e'.
- − The cloche dome is stylized to the point of being a bit abstract.
Stable Diffusion 3.5 Medium
- + Strong hand-drawn vintage illustration style.
- + Compelling layout with the cloche integrated as the main container.
- − Multiple spelling errors including 'Florrian' and 'Est 170'.
- − Lacks the 'minimalist' and 'vector' style requested in the prompt.
- − Steam element is missing or poorly represented compared to the prompt requirements.
Verdict: Stable Diffusion 3.5 Large followed the prompt much more closely, delivering the requested vector emblem aesthetic with nearly perfect text. Stable Diffusion 3.5 Medium failed significantly on the text rendering ('Florrian' and '170') and captured a detailed illustrative style rather than the requested minimalist vector style.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
Stable Diffusion 3.5 Large
- + Features a complex layout that mimics high-end technical posters.
- + Incorporates a detailed spacecraft graphic that fits the Apollo theme.
- + Uses a consistent and professional navy and muted red color palette.
- − Fails to follow the sequential 1-6 step-by-step technical requirement.
- − Includes a Space Shuttle-style craft which is historically inaccurate for Apollo 11.
- − The text is completely garbled and functions only as a visual texture.
Stable Diffusion 3.5 Medium
- + Successfully lists several numbered steps as requested in the prompt.
- + Follows the 'flat-vector' style and iconography requirements more closely.
- + Text is more legible and cleaner, even with slight misspellings.
- − The composition feels unbalanced with a massive empty white space on the right.
- − Iconography for the steps is inconsistent and doesn't always match the text.
- − Incorrectly identifies 'Landing' with an image that looks like Earth.
Verdict: Stable Diffusion 3.5 Large (Image A) produces a more visually impressive and aesthetically pleasing poster, but it fails significantly on prompt adherence by including the wrong spacecraft and ignoring the step-by-step structure. Stable Diffusion 3.5 Medium (Image B) attempts to follow the numbered sequence and vector style much more accurately, despite its awkward composition and minor spelling errors. Stable Diffusion 3.5 Medium is the likely winner for better following the logical structure of the instructional prompt.
Explore each model
Stability AI's 2.5-billion parameter Multimodal Diffusion Transformer with improvements (MMDiT-X) text-to-image model optimized for consumer hardware, featuring improved image quality, typography, and complex prompt understanding