Distilled version of SD 3.5 Large that generates high-quality images in just 4 steps, offering faster inference and reduced costs
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
Stable Diffusion 3.5 Large Turbo
#61 of 62 in Text-to-Image
Wan 2.7
#38 of 62 in Text-to-Image
Where the votes landed
Stable Diffusion 3.5 Large Turbo
0%
win rate
Ties
0%
Wan 2.7
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
Stable Diffusion 3.5 Large Turbo
- + Clean, minimalist aesthetic with sharp rendering
- + Clear soft lighting coming from the left as requested
- − Failed the spatial positioning of the red book, placing it inside/under the sphere instead of on top of the cube
- − The back of the 'cube' is missing its frame, appearing like an open-faced display box
Wan 2.7
- + Perfect prompt adherence, placing all objects in their correct relative positions
- + Highly realistic textures on the wooden table and the book's binding
- + Convincing glass reflections and refractive effects showing the plant behind
- − The sphere's reflection on the bottom glass panel is slightly misaligned
- − The glass cube has double-thick edges that look slightly more like a terrarium or heavy tank than a simple glass cube
Verdict: Wan 2.7 followed every spatial instruction perfectly, correctly placing the book on top of the cube and the sphere inside it. Stable Diffusion 3.5 Large Turbo failed the basic spatial logic by placing the book inside the cube, serving as a base for the sphere. Wan 2.7 also features significantly higher levels of photorealism and texture detail.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
Stable Diffusion 3.5 Large Turbo
- + Strong red color on the bicycle
- + Good lighting contrast
- − Anatomical errors with hands and hair texture
- − Bicycle geometry is nonsensical with missing fork and frame parts
- − Poor background coherence
Wan 2.7
- + Exceptional realism and natural skin textures
- + Perfectly captures the 'candid street photo' aesthetic with believable reflections
- + Mechanically accurate bicycle design and lifelike rain effects
- − Very subtle motion blur on the cars despite the request
- − Shallow depth of field is present but could be slightly more pronounced
Verdict: Wan 2.7 significantly outperforms Stable Diffusion 3.5 Large Turbo by delivering a highly realistic, photographically believable scene with accurate proportions and textures. While Wan 2.7 looks like a real 50mm street photograph, Stable Diffusion 3.5 Large Turbo suffers from significant anatomical distortions and a lack of understanding of bicycle mechanics.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
Stable Diffusion 3.5 Large Turbo
- + Ornate engraving on the armor is very intricate and artistic.
- + Distinct bokeh lights in the background align with the prompt.
- − The character's face has a plastic, overly-smoothed texture typical of early AI models.
- − The blood/scars look like face paint rather than physical injuries.
- − Failed to include beads in the hair braids.
Wan 2.7
- + Excellent skin texture with realistic pores and believable battle scars.
- + Highly accurate adherence to hair details, including both braids and small beads.
- + Superior armor lighting that convincingly reflects the warm torchlight in the scene.
- − The bokeh sparks are a bit sparse compared to the main light sources.
Verdict: Wan 2.7 is the clear winner as it demonstrates significantly higher realism, particularly in skin texture and lighting. While Stable Diffusion 3.5 Large Turbo produces a stylized, almost digital-art look with missing details like hair beads, Wan 2.7 perfectly captures the 'battle-worn' aesthetic with realistic scars, detailed leather straps, and convincing torchlight reflections.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
Stable Diffusion 3.5 Large Turbo
- + Features a grid layout for food photos as requested
- + Includes specific sections like 'Pizza' and 'Mians'
- − Confusing composition that looks like a collage of items rather than a finished menu page
- − Text is largely illegible gibberish
- − Food images look overly saturated and plasticky
Wan 2.7
- + Highly professional and realistic menu layout
- + Excellent text rendering for headers, prices, and descriptions
- + Food photography is realistic and fits the 'casual dining' vibe perfectly
- − Uses a single photo grid for all items rather than distinct visual sections for categories
- − Minor typo in 'Calanrfri' and 'Bisge'
Verdict: Stable Diffusion 3.5 Large Turbo produces a disjointed and abstract layout that fails the 'professional menu' requirement, despite following the category prompt. Wan 2.7 delivers a near-perfect commercial-grade menu design with readable fonts, realistic food photography, and a cohesive aesthetic that aligns perfectly with the minimalist and casual dining prompts.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
Stable Diffusion 3.5 Large Turbo
- + Strong atmospheric lighting with a fiery, dramatic glow.
- + High visual impact with the heat effects at the base of the burger.
- − Failed to include any of the requested text.
- − The burger is slightly stacked and not truly 'exploded' into individual components.
- − Includes an illogical white stick/spit through the top.
Wan 2.7
- + Excellent adherence to the 'exploded' concept with clearly separated ingredients.
- + Perfect text rendering of all requested phrases including the price in a starburst.
- + Highly creative fiery font effect that matches the prompt.
- − Individual ingredients (like tomatoes and pickles) feel slightly more illustrative than photorealistic.
- − The composition is a bit crowded with many floating lettuce leaves.
Verdict: Wan 2.7 followed every instruction in the prompt, including complex text integration and the specific 'exploded' layout requested. Stable Diffusion 3.5 Large Turbo produced a visually striking image but completely ignored the text requirements and the instruction to separate the burger components.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
Stable Diffusion 3.5 Large Turbo
- + Clean and modern café aesthetic
- + Bright lighting and pleasant composition
- − Significant spelling errors throughout the text
- − Failed to render the specific menu items and date requested
- − Text looks more like a digital font than handwriting
Wan 2.7
- + Excellent prompt adherence with nearly perfect spelling of all complex menu items
- + Authentic chalkboard texture with realistic smudges and chalk dust
- + Consistent handwriting style across the entire board
- − Text lacks the requested 'elegant cursive' for the title, using a stylized print instead
Verdict: Wan 2.7 followed the prompt with high precision, correctly spelling the specific menu items and prices while maintaining a realistic chalk texture. In contrast, Stable Diffusion 3.5 Large Turbo struggled with the specific text requirements, resulting in numerous spelling errors and a generic layout that did not match the prompt's content.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
Stable Diffusion 3.5 Large Turbo
- + Strong cinematic lighting with high contrast
- + Clean, smooth textures on the spacesuit and horse
- + Dynamic composition with a dramatic perspective of Earth
- − Anatomical issues with the horse's legs and hooves
- − The rider's leg is positioned through the horse rather than in a stirrup
- − Failed the negative constraint (astronaut is on top, not the horse)
Wan 2.7
- + Excellent fine detail in the horse's coat and spacesuit textiles
- + Better background complexity with galaxies and celestial bodies
- + Clearer rendering of the astronaut's face behind the visor
- − Failed the negative constraint (astronaut is on top, not the horse)
- − The horse's front-right leg has an anatomical disconnect at the shoulder
- − Somewhat generic 'AI art' composition compared to the more stylized Model A
Verdict: Both models completely failed the negative constraint/spatial instruction to place the horse on top of the astronaut, instead providing the standard astronaut-on-horse trope. Wan 2.1 is the slight winner due to superior detail in textures and a more expansive celestial background, whereas Stable Diffusion 3.5 Large Turbo suffered from more significant anatomical distortions in the horse's limbs.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
Stable Diffusion 3.5 Large Turbo
- + Excellent shallow depth of field effect on the city lights
- + Sharp, cinematic lighting on the capybara's fur
- + Good interpretation of the 'professional' capybara expression
- − The passenger in the back is extremely blurry and barely visible
- − The capybara's paws do not clearly wrap around the steering wheel
- − Lacks the wider context of the taxi exterior
Wan 2.7
- + Perfect adherence to the passenger details, including her expression, coat, and phone
- + Dynamic and clear composition showing both characters effectively
- + Realistic depiction of the capybara's paws on the steering wheel
- − The capybara's head is slightly oversized for the body
- − Background buildings are somewhat less 'blurred' than the prompt requested
Verdict: Wan 2.7 is the clear winner as it successfully captures all elements of the prompt, specifically the businesswoman's bored expression and her interaction with her phone in the back seat. Stable Diffusion 3.5 Large Turbo focuses too much on the capybara, resulting in a passenger that is out of focus and lacking the requested details.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
Stable Diffusion 3.5 Large Turbo
- + Clean silhouettes of twisted trees and bats.
- + The central jack-o-lantern has high-contrast, glowing lighting.
- − Failed to include almost all requested text, including date, time, and location.
- − Missing the scroll banner element.
- − Lacks the 'vintage gothic' parchment aesthetic, looking more like a modern digital graphic.
Wan 2.7
- + Excellent adherence to all text requirements with perfect spelling.
- + Rich, detailed composition featuring the banner, thorns, moody sky, and twisted trees.
- + Authentic vintage parchment illustration style that fits the 'gothic invitation' theme.
- − The 'Halloween Party Invitation' text's 'y' in Party has slightly awkward kerning.
- − A bit more busy/crowded than a minimalist poster.
Verdict: Stable Diffusion 3.5 Large Turbo failed significantly on the prompt adherence, missing the specific event details and the banner while using a generic vector-style aesthetic. Wan 2.7 followed the complex instructions flawlessly, rendering all text perfectly and capturing the specific vintage gothic atmosphere requested.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
Stable Diffusion 3.5 Large Turbo
- + Excellent 3D miniature textures and lighting.
- + Good isometric perspective and depth of field.
- − Spelling error in the text 'SIIHI' instead of 'SUSHI'.
- − The flag icon is generic and does not accurately represent the Japanese flag.
- − Text is placed on signs rather than top-center as requested.
Wan 2.7
- + Perfect text rendering for both 'JAPAN' and 'SUSHI'.
- + Accurately included the requested flag icon in the correct position.
- + Followed the layout instructions precisely, including the top-center text placement.
- − The composition is slightly crowded compared to 'minimal garnish' request.
- − One piece of sushi involves a shrimp with eyes, which leans more into cartoonish than the 'refined textures' requested.
Verdict: Wan 2.7 followed every instruction in the prompt, including the specific text content, placement, and the inclusion of the correct flag icon. Stable Diffusion 3.5 Large Turbo produced high-quality objects but failed on the text spelling and placement, and provided an incorrect flag.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
Stable Diffusion 3.5 Large Turbo
- + Excellent fur texture and backlighting effect.
- + High level of polish and color vibrancy.
- − Failed to include all four requested animals, missing the bunny and fox.
- − The animals have a very stylized, 'AI-look' rather than hyper-photorealistic.
- − Anatomy issues with multiple floating/extra paws on the kitten.
Wan 2.7
- + Accurately included all four requested animals (dog, kitten, bunny, fox).
- + High adherence to environmental cues like god rays, dew sparkles, and butterflies.
- + Better sense of action and 'tumbling' as requested in the prompt.
- − The kitten's facial anatomy is slightly distorted.
- − The butterfly wings are somewhat generic in detail.
Verdict: Wan 2.7 is the clear winner as it successfully follows the complex prompt requiring four specific animal types, whereas Stable Diffusion 3.5 Large Turbo only generates two. Wan 2.7 also better captures the requested photographic style and environmental effects like dew and sunbeams.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
Stable Diffusion 3.5 Large Turbo
- + Strong vector emblem style with bold shading
- + Correct color palette with a nice aged paper texture background
- + Creative integration of the cloche with a coffee mug shape
- − Spelling error in the main name ('Caffee Florin')
- − The 'Est. 1720' banner is slightly cluttered and has some minor rendering artifacts
Wan 2.7
- + Excellent text rendering with correct spelling and accents
- + Clean, balanced composition that fits the 'minimalist' prompt
- + Very professional vector-style execution with high clarity
- − Small spelling error in the name ('Florion' instead of 'Florian')
- − Cloche steam icons are slightly detached from the object
Verdict: Wan 2.7 is the stronger choice because it better captures the 'minimalist' and 'vector' elements of the prompt with much cleaner lines and superior typography. While both models had minor spelling errors, Wan 2.7's layout is more professional and balanced compared to the slightly messy banner in Stable Diffusion 3.5 Large Turbo.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
Stable Diffusion 3.5 Large Turbo
- + Sophisticated retro-futuristic artistic style
- + Excellent use of the NASA-inspired color palette
- + Dynamic composition with attractive illustration work
- − Text rendering is poor with many gibberish characters
- − Fails to follow the specific 6-step infographic structure requested
- − Iconography is not consistent with the prompt instructions
Wan 2.7
- + Outstanding adherence to the requested 6-step structure and logic
- + Highly legible and accurate text rendering
- + Clean, consistent iconography in a professional flat-vector style
- − Minor typographical errors ('DESCRIPT', 'Tranquiliry')
- − The layout is more functional than artistic
Verdict: Wan 2.7 is the clear winner as it followed every detail of the complex instructional prompt, including the specific sequence of mission steps and clear text rendering. While Stable Diffusion 3.5 Large Turbo produced a visually stunning piece of art, it failed as an infographic by ignoring the requested content structure and failing to produce legible data.
Explore each model
Alibaba's Wan 2.7 image generation and editing model for text-to-image, reference-guided generation, and instruction-based image edits