Head to head
Esc

Models · slot A

to navigate to pick

Stable Diffusion 3.5 Large Stability AI Wan 2.6 Alibaba

Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.

Stable Diffusion 3.5 Large

22.9 arena score

#29 of 62 in Text-to-Image

Skill signature · Text-to-Image

Wan 2.6

23.3 arena score

#28 of 62 in Text-to-Image

Top 2 in Image-to-Video
Vote tally

Where the votes landed

Stable Diffusion 3.5 Large

71.4%

win rate

Ties

0.0%

Wan 2.6

28.6%

win rate

71.4% 0.0% ties 28.6%
Shared challenges 13

Challenge by challenge

The strongest take from each model on every shared challenge, with the AI judge's read.

Geometric Composition

Text-to-Image

“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”

Stable Diffusion 3.5 Large
Wan 2.6

AI Judge Analysis

Stable Diffusion 3.5 Large

  • + Excellent photorealistic lighting and shadows
  • + Highly sharp and clear details on the glass and book paper
  • Failed spatial reasoning: placed the book and sphere inside the cube rather than the book on top
  • The blue sphere is resting on the book, which was not the prompt instruction

Wan 2.6

  • + Perfect adherence to complex spatial instructions
  • + Realistic vintage texture on the red book
  • + Correctly positioned plant behind the cube visible through glass
  • The glass cube has slightly irregular, wobbly edges
  • The internal reflections in the glass are a bit chaotic

Verdict: Stable Diffusion 3.5 Large produced a higher quality image in terms of clarity and light, but failed significantly on the spatial arrangement by putting the red book inside the cube. Wan 2.6 followed every detail of the prompt perfectly, placing the sphere inside and the book on top, making it the better response to the specific request.

Candid Street Photography

Text-to-Image

“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”

Stable Diffusion 3.5 Large
Wan 2.6

AI Judge Analysis

Stable Diffusion 3.5 Large

  • + Excellent depiction of heavy rain and wet surface reflections
  • + Dynamic and interesting rainy street composition
  • + Strong color contrast with the red bicycle frame
  • Physical interaction with the bike is vague and anatomically confusing near the hands
  • The bike design is structurally nonsensical around the pedals and chain area
  • Character looks slightly more like a general 3D render than a realistic photograph

Wan 2.6

  • + Superb skin textures and realistic aging details on the man's hands and face
  • + The act of 'repairing' is much more explicit and grounded with tools on the ground
  • + Excellent shallow depth of field and bokeh that feels like a 50mm lens
  • Raindrops on the jacket appear as static gel-like beads rather than falling water
  • The bicycle handle orientation and front fork geometry are surreal/incorrect

Verdict: Wan 2.6 provides a much more convincing 'candid' feel with highly realistic skin and textile textures, making the repair scene feel grounded and authentic despite some geometric issues with the bike. Stable Diffusion 3.5 Large creates a beautiful cinematic atmosphere but fails on the specific action of 'repairing' and the physical structure of the bicycle.

Fantasy Warrior

Text-to-Image

“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”

Stable Diffusion 3.5 Large
Wan 2.6

AI Judge Analysis

Stable Diffusion 3.5 Large

  • + Ornate engraved plate armor is highly detailed with consistent patterns.
  • + Braided hair matches the prompt structure well.
  • + The face has a clean, sharp resolution with realistic skin texture.
  • Missed the request for beads in the hair braids.
  • The lighting feels more like daylight than specific warm torchlight.
  • The background contains some distracting, undefined blurry shapes that look like duplicate helmets.

Wan 2.6

  • + Excellent adherence to the 'beads in hair' and 'warm torchlight' prompts.
  • + Outstanding texture on the leather straps and frayed cloth underlayer.
  • + Very convincing 'battle-worn' appearance with realistic dirt and grime.
  • The eyes, while detailed, look slightly moist/glassy in a way that feels a bit digital.
  • The metal engraving is slightly less intricate than in the competing model.

Verdict: Stable Diffusion 3.5 Large produced a very clean and noble-looking paladin, but it missed several key descriptive details like the beads in the hair and the specific warm torchlight atmosphere. Wan 2.6 followed the prompt much more closely, delivering a gritier, more atmospheric image with excellent material textures on the leather and cloth.

Modern Clean Menu

Text-to-Image

“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”

Stable Diffusion 3.5 Large
Wan 2.6
100% wins 0% ties 0% wins

AI Judge Analysis

Stable Diffusion 3.5 Large

  • + Excellent photographic quality and appetizing food presentation
  • + Strong implementation of the requested food grid layout
  • + High-impact bold typography for the main title
  • The 'Mains' and 'Appetizers' sections are poorly structured and mostly unintelligible
  • Food grid occupies the borders rather than being integrated into the functional menu design

Wan 2.6

  • + Excellent structure with clearly defined Appetizers/Pizza/Mains sections
  • + Effective use of vibrant color accents as requested
  • + More realistic menu pricing and itemized structure compared to Model A
  • Food photography is slightly lower in resolution and less vibrant than Model A
  • Text contains several gibberish character artifacts in the body font

Verdict: Stable Diffusion 3.5 Large produced beautiful food photography, but the layout feels more like a poster than a functional menu, with text that is largely unreadable. Wan 2.6 followed the structural instructions much better, creating an organized, three-section menu with vibrant accents and a logical grid, though his food photos were slightly less polished. Wan 2.6 is the winner for better adhering to the design requirements of a restaurant menu.

Magic Burger Explosion: Fiery Photorealism Challenge

Text-to-Image

“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”

Stable Diffusion 3.5 Large
Wan 2.6

AI Judge Analysis

Stable Diffusion 3.5 Large

  • + Excellent photorealistic texture and lighting details on the food.
  • + Vibrant and dynamic fire effects that blend well with the subject.
  • Failed to include any of the requested text.
  • Failed to create an 'exploded' view where components are suspended separately.

Wan 2.6

  • + Perfect adherence to text requirements including 'MAGIC BURGER', 'LIMITED TIME ONLY', and the price starburst.
  • + Accurate interpretation of the 'exploded' layout with components suspended in mid-air.
  • + Strong sense of motion with splashing sauce and floating ingredients.
  • The starburst graphic looks like a flat 2D sticker rather than being integrated into the 3D scene.
  • Slightly less realistic rendering of the burger patties compared to the competitor.

Verdict: Wan 2.6 is the clear winner as it followed every instruction in the prompt, including complex text rendering and the specific layout of an exploded burger. Stable Diffusion 3.5 Large produced a high-quality food image but completely ignored the text and the 'exploded' structural requirement.

Chalkboard Menu

Text-to-Image

“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”

Stable Diffusion 3.5 Large
Wan 2.6

AI Judge Analysis

Stable Diffusion 3.5 Large

  • + The café environment has a clear and realistic composition with balanced lighting.
  • + The chalk board follows a structured design suitable for a real restaurant.
  • Significant spelling errors in the title ('TODAAY') and menu items.
  • Incorrect date year (2024 instead of 2026).
  • The text style looks more like a digital font than natural chalk handwriting.

Wan 2.6

  • + Perfect text rendering with zero spelling errors and correct date.
  • + Excellent chalk texture with realistic smudges, dust, and varying stroke pressure.
  • + Strictly adheres to the 'handwritten' aesthetic without appearing like a font.
  • The composition is a close-up, showing less of the 'cozy café' atmosphere than the other model.
  • Slight blurring on the very edges of the board due to shallow depth of field.

Verdict: Wan 2.6 is the clear winner as it perfectly captured the complex text requirements and the specific chalk texture requested, whereas Stable Diffusion 3.5 Large suffered from numerous spelling errors and a font-like appearance. While Stable Diffusion 3.5 Large provided a better view of the surrounding café, Wan 2.6's superior prompt adherence regarding the handwriting and text content makes it far more successful.

The Reversed Rodeo

Text-to-Image

“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”

Stable Diffusion 3.5 Large
Wan 2.6

AI Judge Analysis

Stable Diffusion 3.5 Large

  • + Excellent cinematic atmosphere and lighting
  • + Dynamic and fluid composition that feels more surreal
  • + Beautifully integrated dust and cloud textures that blend the subjects into space
  • The astronaut's leg position looks physically awkward relative to the horse's back
  • The horse's hind legs blur into the background a bit too much

Wan 2.6

  • + Extremely high detail on the astronaut's suit and horse's fur
  • + Sharp, clear rendering with vibrant colors
  • + Stronger anatomical accuracy for both the horse and the rider's posture
  • The lighting on the horse feels a bit disconnected from the background nebulae
  • The composition is less unique, looking like a standard studio-lit portrait

Verdict: Both models failed the negative constraint to put the horse on top of the astronaut, instead providing the standard rider configuration. Stable Diffusion 3.5 Large offers a more cinematic and dreamy interpretation of the prompt, while Wan 2.6 provides superior technical detail and sharpness. Stable Diffusion 3.5 Large is slightly preferred for its overall artistic composition and cohesive atmosphere.

The Capybara Taxi Driver

Text-to-Image

“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”

Stable Diffusion 3.5 Large
Wan 2.6

AI Judge Analysis

Stable Diffusion 3.5 Large

  • + Excellent high-resolution textures on the fur and clothing
  • + Crisp, vibrant lighting and bokeh effect in the background
  • Completely missed the secondary character in the back seat
  • The capybara's anatomy is a bit strange with human-like legs and jeans

Wan 2.6

  • + Perfect adherence to all prompt elements including the bored businesswoman with a phone
  • + Excellent cinematic composition showing both characters and the exterior night setting
  • + Realistic integration of the capybara's paws on the steering wheel
  • Slightly lower resolution/clarity compared to Model A
  • The capybara's hat looks more like a police hat than a standard taxi driver cap

Verdict: Wan 2.6 is the clear winner as it successfully captured the entire narrative requested, including the bored passenger in the back seat which Stable Diffusion 3.5 Large completely ignored. While Stable Diffusion 3.5 Large had sharper textures and better lighting, its failure to include the second subject and its odd choice to give the animal human legs make it less effective.

The Halloween Invitation

Text-to-Image

“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”

Stable Diffusion 3.5 Large
Wan 2.6

AI Judge Analysis

Stable Diffusion 3.5 Large

  • + Strong parchment effect on the paper edges.
  • + Creative use of a large moon as a light source.
  • Failed to include the specific event details (Date, Time, Location) requested at the bottom.
  • The font style is modern sans-serif rather than the requested elegant gothic style.
  • Contains gibberish text at the bottom of the scroll.

Wan 2.6

  • + Excellent adherence to all text requirements, including the specific date and location.
  • + High-quality rendering of the gothic font with a gold texture.
  • + Beautifully detailed border featuring the requested thorns and webs.
  • The parchment texture is limited to the border and banner rather than the whole layout.

Verdict: Wan 2.6 is the clear winner as it followed every instruction, including the specific event details which Stable Diffusion 3.5 Large completely omitted. Wan 2.6 also better captured the 'elegant gothic' aesthetic requested for the typography and border elements.

Isometric Miniature Diorama Scenes

Text-to-Image

“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”

Stable Diffusion 3.5 Large
Wan 2.6
75% wins 0% ties 25% wins

AI Judge Analysis

Stable Diffusion 3.5 Large

  • + Excellent 3D toy-like aesthetic with appealing textures
  • + Great variety of sushi types on the diorama
  • + Good rendering of shadows and highlights
  • Failed to place text at top-center as requested
  • Text is rendered on a sign within the scene rather than overhead
  • Includes extra cluttered decorative elements contrary to the 'minimal' request

Wan 2.6

  • + Followed text placement and formatting instructions perfectly
  • + Captures the 'minimal' and 'clean' aesthetic requested
  • + Accurate 45-degree isometric perspective and centered diorama
  • The sushi models are relatively simple compared to Model A
  • Texturing on the rice is a bit repetitive

Verdict: Wan 2.6 is the clear winner as it followed every layout instruction, including placing the specified text and flag icon at the top-center of a clean composition. Stable Diffusion 3.5 Large produced a more visually intricate 3D model, but failed the specific positioning instructions for the text, integrating it into the scene on a physical sign instead.

Adorable Baby Animals in Sunny Meadow

Text-to-Image

“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”

Stable Diffusion 3.5 Large
Wan 2.6

AI Judge Analysis

Stable Diffusion 3.5 Large

  • + Excellent soft lighting with a warm atmosphere
  • + Good dynamic sense of movement
  • + Correctly included all requested animal types
  • The fox and kitten look very similar in facial structure
  • Lower level of fine fur detail compared to Model B
  • The butterflies have slightly unnatural placements on heads

Wan 2.6

  • + Exceptional fur texture and detail on all animals
  • + Perfect execution of 'god rays' and dew sparkles
  • + Distinct and accurate physical characteristics for the kitten and fox kit
  • The kitten's pose is a bit awkward with its paws in the air
  • Composition is slightly crowded compared to the more airy Model A

Verdict: Wan 2.6 is the clear winner due to its superior photographic quality and adherence to the 'hyper-photorealistic' part of the prompt. While Stable Diffusion 3.5 Large creates a beautiful scene, Wan 2.6 manages much better distinctness between the animals (specifically the tabby kitten and fox) and provides more detailed textures and more effective 'god rays' and dew drop effects.

Vintage Cafe Logo

Text-to-Image

“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”

Stable Diffusion 3.5 Large
Wan 2.6
0% wins 0% ties 100% wins

AI Judge Analysis

Stable Diffusion 3.5 Large

  • + Excellent command of vector emblem aesthetics with clean lines and balanced spacing.
  • + Accurate rendering of nearly all text including 'Est. 1720' and 'Florian'.
  • + Effective use of subtle parchment-like background texture.
  • Spelled the main text 'Cafféé' incorrectly by doubling the 'e'.
  • The floating cloche top and steam elements underneath create a slightly messy central silhouette.

Wan 2.6

  • + Successfully rendered the complex 'Caffè' accent and overall spelling perfectly.
  • + Stronger 3D shading on the cloche dome giving it more dimension.
  • + Followed the request for a banner containing the establishment date more integrated with the icon.
  • The 'Est. 1720' text is slightly warped following the curve of the banner.
  • The steam iconography is a bit thin and less stylistically cohesive than the rest of the logo.
  • Lack of frame/border elements makes the composition feel more standard than vintage.

Verdict: Wan 2.6 is the winner primarily due to its perfect spelling of 'Caffè Florian', including the specific accent required. While Stable Diffusion 3.5 Large has a slightly more sophisticated vintage border and layout, its double-vowel spelling error and awkward floating cloche design make it less viable for a professional logo task.

Apollo 11: Journey to Tranquility

Text-to-Image

“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”

Stable Diffusion 3.5 Large
Wan 2.6
100% wins 0% ties 0% wins

AI Judge Analysis

Stable Diffusion 3.5 Large

  • + Successfully integrated multiple infographic elements and technical diagrams.
  • + Captured the requested NASA-inspired color palette and vector illustration style.
  • + Includes a wide variety of space imagery including the Moon, Earth, and spacecraft.
  • The main spacecraft is a Space Shuttle, which is historically incorrect for the Apollo 11 mission.
  • The text is largely illegible gibberish.
  • The layout is cluttered and lacks the requested 6-step logical flow.

Wan 2.6

  • + Excellent text rendering with clear, legible names and title.
  • + Clean, minimalist composition that feels modern and professional.
  • + Follows the requested color palette perfectly.
  • Completely failed to include the requested 6-step infographic content (Launch, Orbit, etc.).
  • Lacks the icons requested in the prompt, such as the Saturn V or Lunar Module.
  • Very simplistic interpretation that misses the core technical nature of the prompt.

Verdict: Stable Diffusion 3.5 Large attempted the complex technical nature of the prompt, creating a dense infographic feel, but it failed historically by showing a Space Shuttle instead of the Apollo hardware and lacked a coherent step-by-step flow. Wan 2.6 produced a very clean and aesthetically pleasing graphic with perfect text, but it ignored almost all the specific infographic steps and instructions in the prompt. Stable Diffusion 3.5 Large is the winner for better following the intent of an 'infographic poster' despite its factual errors and messy text.

Next steps

Explore each model