Stability AI's 8.1-billion parameter Multimodal Diffusion Transformer (MMDiT) text-to-image model featuring improved image quality, typography, complex prompt understanding, and resource-efficiency
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
Stable Diffusion 3.5 Large
#29 of 62 in Text-to-Image
Wan 2.5 (Preview)
#27 of 62 in Text-to-Image
Where the votes landed
Stable Diffusion 3.5 Large
0%
win rate
Ties
0%
Wan 2.5 (Preview)
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
Stable Diffusion 3.5 Large
- + Excellent photorealism with realistic dust and scratches on the cube surface
- + High-quality lighting and shadow integration on the wooden table
- − Failed the spatial constraint of the book, placing it inside/under the sphere instead of on top
- − The blue sphere is sitting on the book rather than being inside the cube on its own
Wan 2.5 (Preview)
- + Perfect adherence to all spatial instructions, including the book on top and sphere inside
- + Realistic caustic reflections of the blue sphere on the glass bottom
- + Atmospheric lighting with visible dust motes that match the window light
- − Minor distortion where the book's shadow meets the top edge of the glass cube
Verdict: Wan 2.5 (Preview) followed the complex spatial instructions perfectly, placing the sphere inside the cube and the book on top, whereas Stable Diffusion 3.5 Large failed and placed the book inside the cube. Wan 2.5 also provided a more aesthetically pleasing composition with better background depth and realistic glass refractions.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
Stable Diffusion 3.5 Large
- + Captures the movement blur of a passing car well
- + Strong atmospheric rain effect
- + Good use of reflections on the wet pavement
- − Anatomy and pose are awkward with mismatched limb connections
- − Bicycle geometry is simplified and slightly warped
- − The background vehicle is semi-merged with the environment
Wan 2.5 (Preview)
- + Highly realistic skin textures and facial details
- + Accurate and intricate bicycle components with visible tools
- + Clear adherence to the shallow depth of field request
- − Missing the motion blur from cars mentioned in the prompt
- − The bike stands on a kickstand while the man is working elsewhere, slightly reducing the 'repairing' action coherence
Verdict: Wan 2.5 (Preview) produces a much more realistic and technically sound image, particularly in the rendering of the face, hands, and bicycle mechanics. While Stable Diffusion 3.5 Large follows the 'motion blur' instruction better, its anatomical errors and messy composition make it a less successful image overall.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
Stable Diffusion 3.5 Large
- + Extremely intricate engraving on the plate armor
- + More realistic skin texture and subtle scarring
- + Excellent framing that conveys the scale of a battlefield
- − Missed the request for beads in the braided hair
- − The lighting feels more like daylight than flickering torchlight
Wan 2.5 (Preview)
- + Perfectly depicts the small beads in the braided hair as requested
- + Atmospheric warm torchlight with clear secondary light sources
- + Outstanding texture on leather straps and frayed cloth underlayer
- − The character looks slightly younger/less battle-worn than expected
- − Minor skin texture smoothing compared to Model A
Verdict: Wan 2.5 (Preview) adhered better to the specific details of the prompt, including the hair beads and the interplay of torchlight and multiple material textures. While Stable Diffusion 3.5 Large produced stunningly detailed armor engravings and a more rugged facial appearance, it missed several key decorative elements requested.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
Stable Diffusion 3.5 Large
- + Features a bold and clear 'Menu' header with professional typography
- + Produces very high-quality food photography with realistic textures
- + Follows the white background and grid layout request perfectly
- − The text content is largely gibberish with frequent spelling errors
- − The layout feels more like a social media collage than an actual functional menu
Wan 2.5 (Preview)
- + Excellent logical structure with clearly defined sections for Appetizers, Pizza, and Mains
- + Successfully incorporates vibrant color accents through text and dividers
- + Captures the 'casual dining' aesthetic with a functional, easy-to-read layout
- − The food items are repetitive and lack variety between sections
- − Minor spelling issues in headers like 'Menue' and 'Mainns'
Verdict: Wan 2.5 (Preview) produced a far more functional and professional menu layout that truly looks like a design for a restaurant, featuring clever use of color-coded sections. While Stable Diffusion 3.5 Large has superior image quality for the food photography itself, its layout is less practical and the text is more heavily corrupted.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
Stable Diffusion 3.5 Large
- + Excellent photorealistic texture on the meat patties and vegetables.
- + Strong atmospheric lighting with realistic fire and charcoal.
- − Failed to include any of the requested text elements.
- − Did not render an 'exploded' view; the burger is fully assembled.
Wan 2.5 (Preview)
- + Perfect adherence to all text requirements including 'MAGIC BURGER', 'LIMITED TIME ONLY', and the price starburst.
- + Captures the 'exploded' motion perfectly with components suspended in mid-air.
- + Creative fiery dripping effect on the typography.
- − The meat patty texture is slightly less realistic compared to Model A.
Verdict: Stable Diffusion 3.5 Large produced a high-quality image of a burger, but it completely ignored the core layout instructions (exploded view) and all text prompts. Wan 2.5 (Preview) followed the prompt perfectly, delivering a professional-looking advertisement with accurate text rendering and the requested dynamic composition.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
Stable Diffusion 3.5 Large
- + Features a wider composition that shows the cafe interior environment.
- + Excellent chalk texture on the board surface and surrounding smears.
- − Significant spelling errors in every line of text.
- − Failed to follow the specific date provided in the prompt.
- − Failed to render the specific menu items correctly, hallucinating strange words.
Wan 2.5 (Preview)
- + Exceptional text rendering with perfect spelling of the requested complex menu items.
- + Accurately captured the requested date of April 30, 2026.
- + Followed the instruction for specific price points correctly.
- − The handwriting style is a bit too clean, bordering on a 'chalk font' appearance despite instructions.
- − The composition is a tight crop, showing less of the surrounding café.
Verdict: Wan 2.5 (Preview) is the clear winner as it perfectly rendered the complex text and specific menu items requested, including the far-future date. Stable Diffusion 3.5 Large failed significantly on the text-rendering task, producing numerous spelling errors and incorrect items despite having a nice environmental composition.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
Stable Diffusion 3.5 Large
- + Excellent cinematic atmosphere with ethereal lighting and nebula-like dust clouds.
- + Strong surrealist quality that integrates the subjects into the environment naturally.
- + Dynamic composition with a sense of high-speed motion across the planetary horizon.
- − Failed the specific spatial instruction; the astronaut is riding the horse, not vice versa.
- − The horse's front legs and hooves have anatomical inconsistencies.
Wan 2.5 (Preview)
- + Very high clarity and sharpness on the astronaut and horse textures.
- + Vibrant color palette with a clear celestial background featuring a galaxy and Earth.
- + Realistic horse anatomy and detailed tack/harness.
- − Failed the specific spatial instruction; the astronaut is riding the horse.
- − The dust trailing at the bottom appears more like ground-based sand than a space-based effect.
Verdict: Both models failed the specific prompt constraint to have the horse on top of the astronaut, instead defaulting to the common trope of an astronaut riding a horse. Stable Diffusion 3.5 Large is visually superior in terms of its cinematic, surreal atmosphere and lighting, whereas Wan 2.5 (Preview) provides a cleaner but more generic digital art look. Stable Diffusion 3.5 Large is the preferred choice for its more creative interpretation of 'cinematic' space aesthetics.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
Stable Diffusion 3.5 Large
- + Excellent texture on the capybara's fur and the leather jacket
- + Strong lighting and color saturation consistent with a night scene
- − Completely failed to include the passenger mentioned in the prompt
- − The capybara's hands are anatomically confused and not correctly placed on the steering wheel
Wan 2.5 (Preview)
- + Followed all prompt instructions including the businesswoman in the back seat
- + Better composition that shows the 'view through the windows' and the external taxi signs
- + Accurate depiction of a professional-style driver's cap
- − The capybara's hands/paws are slightly distorted
- − Image clarity is slightly lower than model A with some soft edges on the background passenger
Verdict: Wan 2.5 (Preview) is the clear winner as it followed every part of the complex prompt, successfully including the bored businesswoman and the Manhattan street view. Stable Diffusion 3.5 Large produced a high-quality close-up of a capybara but entirely ignored the second character and the 'view through the windows' requirement.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
Stable Diffusion 3.5 Large
- + Features a cool burnt-edge parchment effect that suits the vintage gothic theme.
- + Includes the requested scroll banner with legible cursive text.
- − Failed to include the specific event details (Date, Time, Location) at the bottom.
- − The layout is crowded, and the 'central' jack-o-lantern is pushed to the side for a moon.
Wan 2.5 (Preview)
- + Followed all text instructions perfectly, including the specific event details.
- + Excellent interpretation of the 'spooky border with webs and thorns' and 'central glowing jack-o-lantern' prompts.
- + Clean, professional composition that fits the requested square format.
- − The font for the bottom details is a bit plain compared to the gothic heading.
Verdict: Wan 2.5 (Preview) provided a much more accurate response to the prompt, successfully including all requested text and event details where Stable Diffusion 3.5 Large failed. Wan 2.5 also followed the compositional cues better, creating a balanced and polished invitation with a clear central subject.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
Stable Diffusion 3.5 Large
- + Excellent variety and realistic textures of the sushi pieces.
- + Detailed isometric diorama base with realistic lighting and shadows.
- + Accurate and sharp text rendering on the little sign.
- − The text is on a sign in the scene rather than being 'at top-center' as a graphic element.
- − Includes many cluttered extra elements like the bowl of salt and loose chopsticks not requested.
Wan 2.5 (Preview)
- + Perfectly adheres to the text placement and graphic design requirements.
- + Clean minimalist aesthetic that matches the 'cartoon scene' and 'miniature' prompt keywords.
- + Very clean soft-shading and 3D render look that feels professional.
- − The sushi itself is very simple and could benefit from more detailed texture.
- − The camera is more at a side-on eye level than the requested 45° top-down isometric angle.
Verdict: Stable Diffusion 3.5 Large creates a much more detailed and realistic set of sushi, but it misses the layout instructions regarding the text placement. Wan 2.5 (Preview) follows the composition and text layout instructions perfectly, delivering a clean, graphic diorama that matches the 'top-center' text and 'small diorama' requirements exactly.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
Stable Diffusion 3.5 Large
- + Excellent dynamic lighting with soft bokeh and realistic backlighting
- + Cohesive artistic style that feels very warm and joyful
- + Good sense of motion in the puppy's pose
- − The kitten looks like a miniature fox/cat hybrid rather than a tabby
- − Anatomical issues with the rabbit's ears and head placement
- − Butterflies lack realistic detail and look like orange smudges
Wan 2.5 (Preview)
- + Correctly identifies all four animals including the tabby pattern and red fox features
- + Clearer depiction of 'god rays' and dew sparkles as requested
- + Sharper overall focus on all subject characters
- − The fox's eyes have unnatural, saturated blue/orange circular artifacts
- − The dew drops appear as floating glass spheres rather than moisture on plants
- − The composition feels slightly more digital and less 'photorealistic' than Model A
Verdict: Stable Diffusion 3.5 Large creates a more atmospheric and aesthetically pleasing image with superior lighting, but it fails to accurately render the specific 'tabby' kitten requested. Wan 2.5 (Preview) follows the prompt more accurately by including all specific animal types and features, though it suffers from uncanny eye artifacts and less natural-looking environmental details.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
Stable Diffusion 3.5 Large
- + Strong vector emblem style with clean lines
- + Accurate brown and cream color palette
- + Minimalist layout that fits a logo format well
- − Includes a spelling error in the word 'Cafféé'
- − The steam effect coming out of the top of the cloche is slightly confusing
Wan 2.5 (Preview)
- + Perfect text rendering for 'CAFFÈ FLORIAN'
- + Excellent execution of the cloche dome with inside steam
- + High aesthetic appeal with the crumpled paper texture background
- − The 'Est. 1720' text is placed on a second ribbon rather than the same banner, though it still looks good
Verdict: Wan 2.5 (Preview) is the winner as it correctly spelled the restaurant name and provided a more sophisticated illustration of the cloche dome. Stable Diffusion 3.5 Large produced a solid minimalist logo, but the double 'é' in the name is a significant flaw for a branding task.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
Stable Diffusion 3.5 Large
- + Captures a sophisticated, highly detailed infographic aesthetic
- + Excellent color palette adherence with a vintage NASA feel
- − Fails to follow the specific requested steps for the mission
- − Contains nonsensical garbled text and confusing planet/rocket combinations like a space shuttle on a Saturn V base
Wan 2.5 (Preview)
- + Follows the specific mission steps and labels remarkably well
- + Produces legible, accurate, and relevant text
- + Includes clever supporting details like the crew names and Tranquility Base marker
- − The rocket design is a hybrid rather than a true Saturn V
- − The illustration of Michael Collins is wearing a helmet that looks more like a military helmet than a space helmet
Verdict: Wan 2.5 (Preview) significantly outperforms Stable Diffusion 3.5 Large by correctly following the sequence of steps and providing legible, accurate text. While Stable Diffusion 3.5 Large creates a visually interesting layout, it ignores the specific content requirements and produces gibberish, whereas Wan 2.5 (Preview) creates a functional, educational infographic.
Explore each model
Alibaba's text-to-image and image-to-image generation model from the Wan AI suite, offering high-quality visual generation capabilities