Stability AI's 2.5-billion parameter Multimodal Diffusion Transformer with improvements (MMDiT-X) text-to-image model optimized for consumer hardware, featuring improved image quality, typography, and complex prompt understanding
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
Stable Diffusion 3.5 Medium
#57 of 62 in Text-to-Image
Wan 2.6
#28 of 62 in Text-to-Image
Where the votes landed
Stable Diffusion 3.5 Medium
0%
win rate
Ties
0%
Wan 2.6
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
Stable Diffusion 3.5 Medium
- + Excellent soft lighting consistent with the prompt
- + Clean glass textures and realistic light refraction
- − The sphere is floating unrealistically in the center of the cube
- − The red book looks like a flat red board rather than a detailed book
- − The plant is mostly above rather than behind the cube
Wan 2.6
- + Highly realistic book texture and aging
- + Excellent plant placement that interacts with the glass as requested
- + Grounded physics with the sphere resting on the bottom of the cube
- − The sphere is quite large relative to the cube, whereas the prompt asked for a small sphere
Verdict: Wan 2.6 followed the spatial instructions much better than Stable Diffusion 3.5 Medium, particularly regarding the plant placement and the realism of the objects. While Stable Diffusion 3.5 Medium captured the 'soft light' well, Wan 2.6 produced a far more convincing composition with superior textures.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
Stable Diffusion 3.5 Medium
- + Successfully captures a filmic, candid quality with 'imperfect framing'.
- + Excellent representation of wet pavement and urban atmosphere.
- + Good sense of 'motion blur from passing cars' in the background.
- − The man's face is largely obscured and lacks detail.
- − The bicycle anatomy is somewhat nonsensical, especially the frame and handlebar connection.
- − The man's hands appear melted or poorly defined.
Wan 2.6
- + Exceptional detail in skin texture and realistic age progression.
- + Much higher anatomical accuracy for the red bicycle and the act of repairing it.
- + Captures the lighting and rain droplets on the jacket with high realism.
- − The framing feels a bit too 'composed' for a requested 'candid/imperfect' shot.
- − Less noticeable motion blur on the passing car compared to the prompt requirement.
Verdict: Wan 2.6 is the clear winner due to its superior anatomical accuracy and much higher level of detail in the subject's face and hands. While Stable Diffusion 3.5 Medium captured the requested 'imperfect framing' and motion blur more effectively, it suffered from significant artifacts and lack of definition in the central elements.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
Stable Diffusion 3.5 Medium
- + Excellent depiction of ornate engraved plate armor with high metallic shine
- + Impressive 'bokeh sparks' effect integrated with the lighting
- + Strong intensity in the character's expression and eyes
- − Missed the request for beads in the hair braids
- − The hair textures look somewhat artificial and overly sharpened
Wan 2.6
- + Perfect adherence to all prompt details including the beads in the braids
- + Superior rendering of leather straps and the 'cloth underlayer' mentioned in the prompt
- + More realistic skin texture with distinct 'faint scars and dirt'
- − The torch in the background is a bit distracting and less 'bokeh' compared to Model A
- − The armor engraving is slightly less intricate than in Image A
Verdict: While Stable Diffusion 3.5 Medium produces a more punchy, cinematic image with beautiful armor reflections, Wan 2.6 is the clear winner for its superior prompt adherence. Wan 2.6 correctly included the beads in the braids and provided much better texture for the secondary requested elements like the leather straps and the frayed cloth underlayer.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
Stable Diffusion 3.5 Medium
- + Features a large amount of food photography
- + Includes distinct sections for different menu items
- − Text is largely illegible and contains gibberish
- − Fonts are chunky and uneven, lacking a professional sans-serif look
- − Layout feels cluttered and lacks whitespace
Wan 2.6
- + Excellent adherence to the 'bold sans-serif' and 'vibrant accents' request
- + Clear, professional layout with legible section headings
- + Higher photographic quality and better grid alignment
- − One photo is cut off by a layout block
- − Some menu item text contains character artifacts
Verdict: Wan 2.6 is the clear winner as it successfully interprets the design aesthetic of a modern restaurant menu, utilizing clean typography and vibrant color blocks. Stable Diffusion 3.5 Medium fails to produce a professional design, resulting in distorted text and a disorganized layout.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
Stable Diffusion 3.5 Medium
- + Excellent photorealistic rendering of food textures
- + Clean typography and clear layout
- + Accurate rendering of embers and fire colors
- − Failed to provide an 'exploded' burger, showing a mostly assembled floating burger instead
- − Text does not have the requested fiery/glowing effect
- − Starburst design is a basic black and white vector
Wan 2.6
- + Perfect adherence to the 'exploded' burger concept
- + High-quality fiery and glowing text effects as requested
- + Dynamic composition with liquid sauce splashes and smoke
- − The 'LIMITED TIME ONLY' text is slightly clipped at the bottom
- − The starburst graphic is a bit clunky and less integrated than the main title text
Verdict: Wan 2.6 followed the prompt instructions much more closely, successfully delivering the 'exploded' view of the burger and the specific fiery glowing text effects. While Stable Diffusion 3.5 Medium produced a highly realistic burger, it failed on the core compositional requirement of an exploded view and used plain white text instead of the requested fire effect.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
Stable Diffusion 3.5 Medium
- + Features a clear wooden frame
- + Captures an artistic chalk texture
- − Failed significantly on text rendering with numerous spelling errors
- − Included printed/outlined fonts instead of pure handwriting
- − Layout is cluttered with non-requested elements like boxes
Wan 2.6
- + Excellent text accuracy with perfect spelling and formatting
- + Realistic chalk texture including dust and smudges
- + Accurately followed the cursive title and specific dates requested
- − The '30' in the date is slightly larger than the rest of the text
Verdict: Stable Diffusion 3.5 Medium struggled with the text-heavy prompt, resulting in gibberish and several stylistic inconsistencies. In contrast, Wan 2.6 executed the prompt near-flawlessly, rendering all requested menu items with perfect spelling and a highly realistic handwritten chalk aesthetic.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
Stable Diffusion 3.5 Medium
- + Successfully captures a cinematic sense of scale above a planet.
- + The horse and astronaut are rendered with high clarity.
- − Failed the specific trick instruction: the astronaut is riding the horse, not the horse on top of the astronaut.
Wan 2.6
- + Greater detail in the horse's anatomy and the astronaut's gear.
- + Rich, vibrant colors with a very cinematic and surreal lighting effect.
- − Failed the specific trick instruction: also shows the astronaut riding the horse.
- − The horse has anatomical floating artifacts behind its rear legs.
Verdict: Both Stable Diffusion 3.5 Medium and Wan 2.6 failed the negative-logic trick prompt ('horse on top, not vice versa'), instead providing the standard 'astronaut on horse' image. Wan 2.6 is the preferred choice as it offers significantly better texture, lighting, and detail, capturing the 'cinematic' and 'surreal' keywords more effectively than Stable Diffusion.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
Stable Diffusion 3.5 Medium
- + Excellent close-up detail on the capybara's fur and hat hardware.
- + Captures the professional, calm expression on the capybara very well.
- + The lighting on the subject's face is cinematic and clear.
- − The passenger is not looking at her phone and is blurred in the background.
- − The composition is a bit tight, losing the 'taxi' environment outside the front dashboard.
- − The capybara's paws are positioned awkwardly on a dashboard/steering wheel that is mostly out of frame.
Wan 2.6
- + Perfectly depicts the passenger looking at her phone as requested.
- + Better composition showing both characters clearly in their respective seats.
- + Highly realistic rainy Manhattan backdrop with accurate taxi exterior details.
- − The passenger appears to be in the front passenger seat rather than the back seat.
- − The capybara's paws look slightly like human hands in gloves.
Verdict: Wan 2.6 is the clear winner as it successfully incorporates all elements of the prompt, including the passenger's expression and specific action of looking at her phone. While Wan 2.6 placed the passenger in the front seat instead of the back, Stable Diffusion 3.5 Medium failed the 'phone' part of the prompt entirely and suffered from a composition that felt less like a full scene and more like a character portrait.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
Stable Diffusion 3.5 Medium
- + Includes all the requested elements in a clear layout.
- + The torn parchment aesthetic fits the vintage theme well.
- − Significant spelling errors throughout the text (e.g., 'Halloweeen', 'Timme', 'Loccation').
- − The composition feels like clip-art rather than a polished cinematic poster.
- − Failed to generate a central jack-o-lantern, instead placing two on the sides.
Wan 2.6
- + Excellent text rendering with near-perfect spelling and elegant gothic typography.
- + Strong cinematic lighting and atmosphere that feels high-quality and spooky.
- + Accurately places the jack-o-lantern, scroll banner, and thorn/web border as requested.
- − The 'Arches' text is slightly condensed, though still very readable.
- − The thorn border is a bit messy in the bottom corner transitions.
Verdict: Wan 2.6 is the clear winner as it successfully follows the complex layout instructions and renders nearly perfect text. Stable Diffusion 3.5 Medium struggles significantly with spelling and fails to place the jack-o-lantern in the requested central position, resulting in a less professional aesthetic.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
Stable Diffusion 3.5 Medium
- + Features realistic textures on the sushi ingredients
- + Clean and simple composition
- − Missed the flag icon requirement
- − Text has slight rendering artifacts on the first and last letters
- − Failed to include the raised diorama base
Wan 2.6
- + Perfect adherence to all prompt elements including the flag icon and diorama base
- + Excellent text rendering with clean, bold typography
- + Captured the 3D miniature/cartoon aesthetic perfectly
- − Colors are slightly more muted compared to the vibrancy of Model A
Verdict: Wan 2.6 followed the prompt instructions near-perfectly, including specific details like the raised diorama base and the flag icon which Stable Diffusion 3.5 Medium missed. Wan 2.6 also produced much cleaner typography and better captured the 'miniature 3D cartoon scene' aesthetic requested.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
Stable Diffusion 3.5 Medium
- + Saturated and vibrant colors
- + Clear facial features for the three animals present
- − Failed to include the requested baby bunny
- − The kitten resembles a cross between a cat and a fox
- − The fur rendering looks somewhat digital and illustrative rather than hyper-photorealistic
Wan 2.6
- + Successfully included all four requested animals: puppy, kitten, bunny, and fox
- + Excellent capture of the 'god rays' and 'dew sparkles' mention in the prompt
- + Dynamic posing that captures the 'tumbling' and 'playful' aspect of the prompt
- − The fox has an extra black paw appearing near its chest
- − Some of the floating seeds/particles are a bit distracting
Verdict: Wan 2.6 is the clear winner as it fulfilled all parts of the prompt, including the specific list of four animals, whereas Stable Diffusion 3.5 Medium missed the bunny entirely. Wan 2.6 also achieved a much more realistic lighting effect and captured the active, playful energy requested.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
Stable Diffusion 3.5 Medium
- + Excellent hand-drawn illustrative style
- + Intricate detail in the banner and shading
- + Solid adherence to color palette
- − Significant text errors including 'Florrian' and 'Est 170'
- − Overly complex for the 'minimalist' requirement
- − The cloche is stylized to the point of being unrecognizable as a dome
Wan 2.6
- + Perfect text rendering of 'Caffè Florian' and 'Est. 1720'
- + Successfully follows the 'minimalist' and 'vector emblem' instructions
- + Clear representation of all prompt elements including steam and cloche
- − Background texture is a bit generic/distressing
- − Composition of the banner is slightly small and off-center
Verdict: Stable Diffusion 3.5 Medium produces a beautiful illustration but fails significantly on text accuracy and the 'minimalist' keyword. Wan 2.6 perfectly captures the brand name and the requested vector style, making it a much more functional logo design.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
Stable Diffusion 3.5 Medium
- + Attempts the sequential infographic layout requested
- + Uses the specified color palette effectively
- + Captures the vector/space aesthetic
- − Text is largely gibberish or heavily misspelled
- − The icons do not accurately represent the specific stages (e.g., Saturn V icon is missing)
- − The logic of the steps is disorganized
Wan 2.6
- + Text rendering for names is legible
- + Follows the color palette accurately
- + Clean, minimal composition
- − Completely fails to include the 6 requested steps/icons
- − The image looks like a towel or fabric print rather than a vector poster
- − Lacks any of the required mission details like orbit rings or the lunar module
Verdict: Stable Diffusion 3.5 Medium followed the complex prompt structure much better by attempting a sequential infographic with icons and a trajectory, despite the garbled text. Wan 2.6 failed almost every specific instruction regarding the mission steps and layout, producing a very basic design that looks like a textile product. Stable Diffusion 3.5 Medium is the clear winner for effort and thematic adherence.
Explore each model
Alibaba's multimodal generation model from the Wan AI suite, supporting text-to-video, image-to-video, reference-to-video with audio, and text-to-image, in both Chinese and English