Stability AI's 8.1-billion parameter Multimodal Diffusion Transformer (MMDiT) text-to-image model featuring improved image quality, typography, complex prompt understanding, and resource-efficiency
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
Stable Diffusion 3.5 Large
#30 of 62 in Text-to-Image
Wan 2.7
#39 of 62 in Text-to-Image
Where the votes landed
Stable Diffusion 3.5 Large
0%
win rate
Ties
0%
Wan 2.7
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
Stable Diffusion 3.5 Large
- + Excellent handling of glass transparency and realistic surface scratches
- + Correct lighting direction from the left with sharp, realistic shadows
- − Failed the spatial request by putting the red book inside the cube and the sphere on top of the book
- − The plant is not clearly visible through the glass as requested
Wan 2.7
- + Perfectly followed the complex spatial instructions
- + High-quality textures on the wooden table and the red book
- + Clear and accurate glass refraction and reflections
- − The glass cube has a strange vertical internal pane that wasn't requested
Verdict: Wan 2.7 followed all spatial instructions perfectly, placing the sphere inside the cube and the red book on top, whereas Stable Diffusion 3.5 Large swapped these elements. Wan 2.7 also better integrated the green plant behind the glass, making it the superior choice for prompt adherence.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
Stable Diffusion 3.5 Large
- + Excellent skin texture and realistic age-related details on the man's arms and face.
- + Beautiful lighting and reflections that convey a cinematic, rainy atmosphere.
- + The red bicycle color is vibrant and correctly executed.
- − Physical interaction between the hands and the bike is messy with some structural clipping.
- − The rain effect looks a bit like static or a digital overlay in some areas.
Wan 2.7
- + Very convincing street photography aesthetic with 'imperfect' framing as requested.
- + The environment feels authentic to a Japanese city street with believable background elements.
- + Good source preservation of the 50mm lens look with natural background figures.
- − The man's skin looks somewhat waxy and lacks the fine texture requested in the prompt.
- − The bicycles structure is distorted, particularly around the handlebars and front fork.
Verdict: Stable Diffusion 3.5 Large wins on pure visual quality and skin texture, though both models struggle with the complex geometry of the bicycle and hands. Wan 2.7 captures the 'candid' street photography vibe and imperfect framing more accurately, but it fails to deliver the high-fidelity 'natural skin texture' specified in the prompt.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
Stable Diffusion 3.5 Large
- + Excellent intricate engraving on the plate armor.
- + Sharp, lifelike eyes and facial expression.
- + Captures the sense of a large-scale battlefield in the background.
- − Missed the request for beads in the braided hair.
- − Armor appears almost too clean and reflective despite the battle-worn description.
Wan 2.7
- + Perfect adherence to the 'beads in hair' prompt detail.
- + Highly realistic skin textures, scars, and dirt application.
- + Strong implementation of warm torchlight, bokeh sparks, and textured leather straps.
- − The armor engraving is slightly less detailed compared to the other image.
- − The background wall is a bit generic.
Verdict: Wan 2.7 followed the prompt more comprehensively by including the specific 'small beads' in the hair and showing clear examples of leather straps and cloth underlayers. While Stable Diffusion 3.5 Large produced stunning armor engravings, Wan 2.7 better captured the gritty, battle-worn atmosphere and the specific lighting effects requested.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
Stable Diffusion 3.5 Large
- + Features a bold, experimental layout with high-quality food photography
- + Strong adherence to the 'grid' request for photos
- + Excellent rendering of the word 'Menu' in a bold sans-serif font
- − Most of the secondary text is illegible gibberish
- − The layout feels more like a poster than a functional restaurant menu
Wan 2.7
- + Highly realistic and functional menu layout with logical sections
- + Impressive text rendering for dishes, prices, and even social media handles
- + Cleverly includes professional elements like a QR code and restaurant info
- − The grid of photos is slightly smaller and less 'modern minimalist' than Image A
- − Adds environmental props (pen, rosemary) not explicitly requested in the design prompt
Verdict: Wan 2.7 is the clear winner as it produces a fully functional and legible menu that follows all instructions, including specific sections for appetizers, pizza, and mains. While Stable Diffusion 3.5 Large has a striking visual style with its large photo grid, it fails to produce readable text or a coherent hierarchy of information, whereas Wan 2.7 delivers a professional, print-ready aesthetic.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
Stable Diffusion 3.5 Large
- + Excellent photorealistic texture on the meat and bun
- + Very coregent lighting with the fire background
- + Powerful, high-action atmospheric effect
- − Completely failed to include any of the requested text
- − The burger is stacked rather than being the requested 'exploded' view
Wan 2.7
- + Perfect adherence to all requested text elements with accurate spelling
- + Accurately represents the 'exploded' motion with suspended components
- + Creative fiery render on the typography
- − Image has a slightly more 'digital illustration' feel compared to the realism request
- − Some floating sesame seeds look a bit disconnected
Verdict: Stable Diffusion 3.5 Large produced a more realistic-looking burger with impressive textures, but completely ignored the core layout and typography requirements. Wan 2.7 successfully integrated all complex text elements, symbols, and the specific 'exploded' layout requested, making it the superior advertisement overall.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
Stable Diffusion 3.5 Large
- + Excellent contextual composition showing the menu within a realistic café environment.
- + Captures a very authentic chalk texture with dust and smudging on the board.
- − Significant spelling errors throughout the board (e.g., 'TODAAY', 'Ottpups', 'Chocoolalfe').
- − Incorrect date rendered as 2024 instead of the requested 2026.
- − Font style looks more digital/aligned than the requested organic handwriting style.
Wan 2.7
- + Perfect text accuracy and spelling for all requested menu items and numbers.
- + Successfully rendered the correct date (April 30, 2026) as specified in the prompt.
- + Strong adherence to the 'handwritten cursive' style with natural-looking ligatures.
- − The text has a slight 'glow' or drop-shadow effect that makes it look like a digital overlay rather than physical chalk.
- − The composition is a tight crop on the board, offering less environmental context than Model A.
Verdict: Stable Diffusion 3.5 Large creates a more believable café atmosphere, but fails significantly on the prompt's specific text requirements and spelling. Wan 2.7 follows the prompt instructions meticulously, delivering perfect text rendering and the correct date, despite the text looking slightly more like a digital font than actual physical chalk.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
Stable Diffusion 3.5 Large
- + Excellent cinematic atmosphere with realistic lighting and nebular dust.
- + High level of detail in the astronaut suit and horse's texture.
- + Strong composition that feels integrated into the space environment.
- − The astronaut's leg position on the saddle is slightly anatomicaly awkward.
Wan 2.7
- + Clear, sharp focus on all elements including the background planets.
- + Well-defined anatomy for both the horse and the astronaut's gear.
- + Vibrant colors and a clean, surreal aesthetic.
- − The horse appears to be floating staticly rather than moving dynamically.
- − Lighting on the subjects feels slightly detached from the background environment.
Verdict: Both models successfully interpreted the prompt, but Stable Diffusion 3.5 Large is the winner due to its superior cinematic quality and the way it blends the horse and rider into the space environment with dust and atmospheric light. Wan 1.3 is very clean and detailed, but it feels more like a composite of two separate images compared to the cohesive scene in Image A.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
Stable Diffusion 3.5 Large
- + Excellent texture on the capybara's fur and clothing.
- + Vibrant color palette that captures a night aesthetic.
- + High resolution with sharp details on the interior hardware.
- − Completely failed to include the human businesswoman in the back seat.
- − The capybara's anatomy is slightly distorted with human-like legs/torso.
- − Only one paw is visible on the wheel despite the prompt's request for both.
Wan 2.7
- + Successfully included all elements of the prompt including the businesswoman and the capybara driver.
- + Perfectly captured the requested bored expression on the passenger's face.
- + Accurately depicted both paws on the steering wheel.
- − Lighting is a bit flat compared to Model A.
- − The textures are slightly less refined, especially the hair of the human passenger.
- − The capybara's face is a bit more static and less expressive.
Verdict: Stable Diffusion 3.5 Large produced a more visually striking image with superior textures, but it failed a major part of the prompt by omitting the passenger. Wan 2.7 followed every instruction perfectly, handling the complex scene composition and multiple characters with high accuracy while maintaining a realistic style.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
Stable Diffusion 3.5 Large
- + Successfully captures a moody, high-contrast cinematic lighting
- + Includes the gothic title and scroll banner with mostly legible text
- + Excellent use of the 'twisted trees' and 'night sky' thematic elements
- − Missed all specific event details (Date, Time, Location) at the bottom
- − Text at the bottom of the scroll degrades into gibberish
- − The 'central' jack-o-lantern is offset to the side rather than the focal point
Wan 2.7
- + Flawless adherence to all text requirements, including specific event details
- + Superior composition with the jack-o-lantern appearing as the central focus
- + Intricate border featuring webs, thorns, and skulls that perfectly fits the gothic theme
- − Lighting is more illustrative than 'cinematic' or 'moody'
- − The color palette is slightly brighter than requested by 'dark parchment' and 'night sky'
Verdict: While Stable Diffusion 3.5 Large creates a more atmospheric and 'spooky' aesthetic, it fails to include the essential event details requested in the prompt. Wan 2.7 followsทุก single instruction perfectly, providing clear, elegant typography for the title, scroll, and the specific event details, while maintaining a very high level of illustrative detail in the gothic borders.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
Stable Diffusion 3.5 Large
- + High level of realistic detail in the sushi textures.
- + Coherent spatial arrangement of the objects.
- + Good lighting and shadows.
- − Failed to place text at top-center, opting for a physical card instead.
- − Composition is a bit cluttered with extra side dishes and garnishes.
- − Camera angle is too low for a requested 45 degree top-down view.
Wan 2.7
- + Perfect adherence to the text placement and layout instructions.
- + Clean, minimalist 3D cartoon aesthetic that matches 'miniature diorama' request.
- + Accurate 45 degree isometric perspective.
- − The shrimp sushi has some slight anatomical irregularities in the tail.
- − The sushi rolls are very simplistic in texture.
Verdict: Wan 2.7 followed the prompt instructions much more accurately, placing the text at the top-center and capturing the clean isometric diorama feel perfectly. While Stable Diffusion 3.5 Large has better material realism, it failed the layout and text placement requirements by placing the text on a small sign within the scene.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
Stable Diffusion 3.5 Large
- + Excellent depiction of dynamic movement through blurred backgrounds and paws off the ground.
- + Strong lighting effects with prominent bokeh and soft, hazy god rays.
- + Captures a very joyful and whimsical expression on the animals' faces.
- − The kitten looks like a miniature fox or generic canine hybrid rather than a tabby kitten.
- − The environment is quite blurry, losing some of the 'lush wildflower' detail.
- − The animals overlap in a way that creates some anatomical confusion in the midground.
Wan 2.7
- + Perfect prompt adherence for all four distinct animals, including a clear tabby pattern on the kitten.
- + Excellent 'god rays' and dew sparkle effects that feel integrated into the physics of the scene.
- + Sharp detail across the entire frame, showing individual petals and grass blades.
- − The fox's face appears a bit stiff or grumpy compared to the 'joyful' prompt.
- − The butterfly sizes are slightly inconsistent relative to each other.
Verdict: Wan 2.7 is the clear winner as it successfully rendered all four specific animals requested, including a distinct tabby kitten which Stable Diffusion 3.5 Large failed to differentiate from the fox. Wan 2.7 also provided a much more detailed and vibrant meadow while perfectly executing the requested lighting effects.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
Stable Diffusion 3.5 Large
- + Strong minimalist vector aesthetic
- + Includes the cloche, steam, and banner as requested
- + Applies a nice subtle paper texture to the background
- − Spelling error in primary text: 'Cafféé' instead of 'Caffè'
- − The cloche illustration is a bit abstract and lacks clarity
Wan 2.7
- + Excellent typography with correct primary spelling 'Caffè'
- + Very clean vector execution with a classic circular emblem layout
- + Detailed and recognizable cloche dome illustration
- − Slight misspelling in the name as 'Florion'
- − The cloche is placed below the main text rather than having the text on a banner below it
Verdict: Wan 2.7 provides a much more polished and professional vector emblem with superior clarity and a more 'classic' feel. While Stable Diffusion 3.5 Large follows the requested layout of placing the banner at the bottom, it suffers from a significant spelling error in 'Cafféé' and a less refined illustration style.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
Stable Diffusion 3.5 Large
- + Features complex, detailed planetary illustrations
- + Adheres well to the requested navy, white, and red color palette
- − Includes a Space Shuttle instead of the requested Saturn V rocket
- − Text is completely illegible and nonsensical
- − Fails to follow the sequential numbered steps requested in the prompt
Wan 2.7
- + Accurately depicts the six requested steps with relevant icons for each
- + Text and labels are highly legible and contextually accurate for the mission
- + Strictly follows the flat-vector design style and color palette
- − Minor spelling errors in text (e.g., 'Descript' instead of 'Descent' and 'Tranquiliry')
- − The background stars are somewhat cluttered compared to the clean vector aesthetic
Verdict: Wan 2.7 significantly outperformed Stable Diffusion 3.5 Large by correctly interpreting the sequential infographic structure and providing readable, relevant text. Stable Diffusion 3.5 Large failed on content accuracy by including a Space Shuttle (which never went to the moon) and produced gibberish text in a chaotic layout.
Explore each model
Alibaba's Wan 2.7 image generation and editing model for text-to-image, reference-guided generation, and instruction-based image edits