Head to head
Esc

Models · slot A

to navigate to pick

Stable Diffusion 3.5 Large Turbo Stability AI Wan 2.5 (Preview) Alibaba

Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.

Stable Diffusion 3.5 Large Turbo

10.7 arena score

#61 of 62 in Text-to-Image

Skill signature · Text-to-Image

Wan 2.5 (Preview)

23.4 arena score

#28 of 62 in Text-to-Image

Vote tally

Where the votes landed

Stable Diffusion 3.5 Large Turbo

0%

win rate

Ties

0%

Wan 2.5 (Preview)

0%

win rate

Shared challenges 13

Challenge by challenge

The strongest take from each model on every shared challenge, with the AI judge's read.

Geometric Composition

Text-to-Image

“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”

Stable Diffusion 3.5 Large Turbo
Wan 2.5 (Preview)

AI Judge Analysis

Stable Diffusion 3.5 Large Turbo

  • + Clean, sharp vector-like aesthetic
  • + Good lighting direction with clear shadows
  • Failed the spatial logical prompt by placing the book inside the cube
  • Plant is to the left rather than behind the cube
  • Cube edges are black frames rather than pure glass

Wan 2.5 (Preview)

  • + Excellent adherence to spatial relationships, placing the book on top and the sphere inside
  • + Realistic materials including glass reflections and paper texture
  • + Accurate placement of the plant behind the cube visible through the glass
  • Slightly dusty/noisy particles in the air might be distracting for some

Verdict: Wan 2.5 (Preview) correctly followed all spatial instructions, placing the book on top of the cube and the sphere inside, whereas Stable Diffusion 3.5 Large Turbo incorrectly placed the book inside with the sphere. Wan 2.5 (Preview) also achieved a much higher level of photorealism and accurate material interaction.

Candid Street Photography

Text-to-Image

“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”

Stable Diffusion 3.5 Large Turbo
Wan 2.5 (Preview)

AI Judge Analysis

Stable Diffusion 3.5 Large Turbo

  • + Clean representation of the red bicycle
  • + Adheres to the color requirements
  • Anatomical failures including a fused hand and a missing leg
  • Artificial, plastic-like skin and clothing texture
  • Lack of wet pavement reflections and raindrops
  • Cars in background are static and not motion blurred

Wan 2.5 (Preview)

  • + Highly realistic skin textures and believable elderly Japanese subject
  • + Excellent execution of wet pavement reflections and rainy atmosphere
  • + Accurate shallow depth of field and street photography layout
  • + Complex environmental details like scattered tools and a kickstand
  • The 'motion blur from passing cars' is present but subtle due to the shallow depth of field
  • Slightly less 'imperfect' framing than requested

Verdict: Wan 2.5 (Preview) significantly outperforms Stable Diffusion 3.5 Large Turbo by delivering a highly realistic, cinematic street photograph with detailed skin textures and convincing rain effects. Stable Diffusion 3.5 Large Turbo fails on multiple anatomical levels, resulting in a surreal, fused human figure that lacks the atmosphere and realism requested in the prompt.

Fantasy Warrior

Text-to-Image

“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”

Stable Diffusion 3.5 Large Turbo
Wan 2.5 (Preview)

AI Judge Analysis

Stable Diffusion 3.5 Large Turbo

  • + Strong, high-contrast colors and sharp details
  • + Ornate engraving on the armor is clearly visible
  • + Good interpretation of the brave, heroic paladin aesthetic
  • Has a distinct digital/airbrushed look on the skin
  • The hair beads mentioned in the prompt are largely missing
  • The beard and stray hairs look somewhat unnaturally sharp and repetitive

Wan 2.5 (Preview)

  • + Excellent adherence to the 'beads in hair' and 'leather straps and cloth' details
  • + Very lifelike skin texture with realistic dirt and scars
  • + Masterful use of warm torchlight and bokeh sparks in the background
  • The character looks slightly younger/less physically imposing than a typical paladin
  • The armor engraving is slightly less intricate than model a

Verdict: Wan 2.5 (Preview) produced a far more realistic and lifelike image that strictly followed the prompt's specific details, such as the beads in the hair and the texture of the underlayers. Stable Diffusion 3.5 Large Turbo created a polished but more 'illustrative' image that felt more like a video game render, missing several specific prompt requirements.

Modern Clean Menu

Text-to-Image

“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”

Stable Diffusion 3.5 Large Turbo
Wan 2.5 (Preview)

AI Judge Analysis

Stable Diffusion 3.5 Large Turbo

  • + Excellent photographic quality and lighting in food images
  • + High contrast and vibrant color palette
  • + Professional usage of bold sans-serif typefaces
  • Layout is fragmented into separate cards rather than a unified menu design
  • Typo in 'Mains' ('Mians')
  • Included unrealistic 'plastic-looking' food illustrations mixed with photography

Wan 2.5 (Preview)

  • + Perfect adherence to the 'grid' and 'unified menu' layout request
  • + Clean inclusion of all requested sections: Appetizers, Pizza, and Mains
  • + Logical and professional use of whitespace and dividers
  • Text contains several spelling errors ('Restormalit Menue')
  • Food images look slightly repetitive in styling across sections
  • Lower resolution/clarity in the small body text

Verdict: Stable Diffusion 3.5 Large Turbo provides superior individual image quality but fails to compose a cohesive menu, offering instead a collection of floating UI elements and cards with a major typo. Wan 2.5 (Preview) perfectly captures the layout and design intent of a modern restaurant menu with a structured grid and clear categorization, making it much more functional for the specific brief despite minor spelling errors.

Magic Burger Explosion: Fiery Photorealism Challenge

Text-to-Image

“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”

Stable Diffusion 3.5 Large Turbo
Wan 2.5 (Preview)

AI Judge Analysis

Stable Diffusion 3.5 Large Turbo

  • + Strong fiery atmosphere with embers and smoke
  • + Vibrant colors and high contrast lightning
  • Failed to include any of the requested text
  • Missing the 'exploded' aspect of the burger, keeping it mostly stacked
  • Lower level of photorealism compared to the competitor

Wan 2.5 (Preview)

  • + Perfect adherence to text requirements including 'MAGIC BURGER', 'LIMITED TIME ONLY', and the price
  • + Excellent 'exploded' layout with components suspended in mid-air
  • + Highly photorealistic textures on the patty, buns, and vegetables
  • The '6.99' text is a bit flat compared to the glowing effect of the main title
  • The fiery background is slightly more subdued than Model A's

Verdict: Wan 2.5 (Preview) is the clear winner as it followed every part of the complex prompt, specifically the integration of multiple text elements and the 'exploded' burger layout. Stable Diffusion 3.5 Large Turbo completely failed to render the requested text and did not separate the burger components as instructed.

Chalkboard Menu

Text-to-Image

“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”

Stable Diffusion 3.5 Large Turbo
Wan 2.5 (Preview)

AI Judge Analysis

Stable Diffusion 3.5 Large Turbo

  • + Clean and modern café environment.
  • + Correct wooden frame and lighting.
  • Terrible prompt adherence regarding text accuracy, including many misspellings like 'trulale' and 'risctto'.
  • The text looks like a digital font rather than natural chalk handwriting.
  • The date and specific menu items are incomplete or incorrect.

Wan 2.5 (Preview)

  • + Exceptional text rendering with almost 100% accuracy to the complex prompt.
  • + Highly realistic chalk texture including authentic smudges and dust on the board.
  • + Perfect execution of the requested handwriting style with natural variations.
  • The pricing for the cookies is repeated ($9 and -$9) which was not in the prompt.
  • The composition is a bit tight on the left side of the frame.

Verdict: Wan 2.5 (Preview) significantly outperformed Stable Diffusion 3.5 Large Turbo by delivering highly accurate text that followed the specific menu items and date requested. While Stable Diffusion 3.5 Large Turbo struggled with spelling and produced text that looked like a digital font, Wan 2.5 (Preview) created a convincing, hand-drawn look with grit and texture that perfectly matched the prompt's aesthetic requirements.

The Reversed Rodeo

Text-to-Image

“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”

Stable Diffusion 3.5 Large Turbo
Wan 2.5 (Preview)

AI Judge Analysis

Stable Diffusion 3.5 Large Turbo

  • + Clean, sharp lighting on the astronaut suit.
  • + Effective use of negative space for a cinematic feel.
  • Anatomical failure in the front legs of the horse.
  • The composition feels a bit flat compared to the requested cinematic style.

Wan 2.5 (Preview)

  • + Excellent texture and details on the horse's coat and mane.
  • + Vibrant galactic background with a dynamic sense of motion.
  • + Superior anatomical rendering of both the astronaut and the horse.
  • The reins are not physically connected to the horse's mouth correctly.

Verdict: Wan 2.5 (Preview) provided a much more detailed and visually interesting image with superior anatomy and a more 'cinematic' feel. Stable Diffusion 3.5 Large Turbo struggled with the horse's front legs, resulting in a distorted pose that detracted from the overall quality.

The Capybara Taxi Driver

Text-to-Image

“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”

Stable Diffusion 3.5 Large Turbo
Wan 2.5 (Preview)

AI Judge Analysis

Stable Diffusion 3.5 Large Turbo

  • + Excellent fur texture and realistic lighting on the capybara
  • + High quality bokeh in the background highlights the city lights
  • + The capybara's paw placement on the wheel looks natural
  • The passenger is severely out of focus and their action is unclear
  • The 'TAXI' text on the hat is slightly garbled
  • Perspective makes it feel more like a close-up profile than a scene showing the whole interior

Wan 2.5 (Preview)

  • + Strict adherence to all prompt elements including the woman looking at her phone
  • + Great composition showing the exterior city, the driver, and the passenger clearly
  • + Realistic rain effects and taxi-specific details like the meter and top light
  • The capybara's hands/claws look a bit strange and human-like
  • The text on the taxi sign is nonsensical
  • Lighting on the capybara's face is a bit flat compared to the surrounding environment

Verdict: Wan 2.5 (Preview) provided a much better overall composition that captured all narrative elements of the prompt, specifically the businesswoman on her phone in the back seat. While Stable Diffusion 3.5 Large Turbo had superior texture and lighting on the capybara itself, it failed to clearly depict the passenger and the 'bored' narrative requested.

The Halloween Invitation

Text-to-Image

“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”

Stable Diffusion 3.5 Large Turbo
Wan 2.5 (Preview)

AI Judge Analysis

Stable Diffusion 3.5 Large Turbo

  • + Strong composition with a clean, graphic border
  • + Good adherence to the 'twisted trees' and 'webs' prompt elements
  • Failed to include major prompt text like the scroll banner and specific event details
  • The lighting is flat and lacks the requested cinematic quality
  • Text rendering is incomplete and simple

Wan 2.5 (Preview)

  • + Excellent adherence to all text requirements including date, time, and location
  • + Highly detailed rendering with cinematic lighting and textures
  • + Includes complex elements like the thorn and web border and the scroll banner perfectly
  • The 'Halloween' text is slightly warped following the curve
  • The background sky is a bit contained within a circle rather than a full page texture

Verdict: Wan 2.5 (Preview) significantly outperformed Stable Diffusion 3.5 Large Turbo by successfully following every detail of the prompt, including complex text and specific layout elements like the scroll banner. While Stable Diffusion 3.5 Large Turbo created a visually pleasing graphic, it ignored almost all of the specific party details and the gothic title requirements.

Isometric Miniature Diorama Scenes

Text-to-Image

“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”

Stable Diffusion 3.5 Large Turbo
Wan 2.5 (Preview)

AI Judge Analysis

Stable Diffusion 3.5 Large Turbo

  • + Excellent 3D isometric diorama perspective.
  • + Creative use of physical signage within the scene.
  • + Effective PBR-like material rendering on the wooden base and plate.
  • Significant spelling error with 'SIIHI' instead of 'SUSHI'.
  • The flag icon is incorrect/unrecognizable.
  • Text placement is on signage rather than floating at top-center as requested.

Wan 2.5 (Preview)

  • + Perfect text rendering for both 'JAPAN' and 'SUSHI'.
  • + Accurate Japanese flag icon included in the layout.
  • + Very clean, soft 3D cartoon aesthetics as requested.
  • Perspective is more eye-level/cinematic than 45° isometric.
  • Diorama base is a simple cylinder rather than a complex miniature scene.

Verdict: While Stable Diffusion 3.5 Large Turbo captures the 'isometric diorama' style more effectively, it fails on basic text accuracy with a glaring typo. Wan 2.5 (Preview) ignores the isometric constraint but delivers a professional-grade graphic with perfect text, a correct flag, and high-quality 3D shading.

Adorable Baby Animals in Sunny Meadow

Text-to-Image

“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”

Stable Diffusion 3.5 Large Turbo
Wan 2.5 (Preview)

AI Judge Analysis

Stable Diffusion 3.5 Large Turbo

  • + Strong backlighting effect with vibrant glowing edges on the fur
  • + Good character consistency across the animals
  • Failed to include a rabbit as requested
  • Animals appear fused together with anatomical clipping and shared fur blocks
  • Style leans towards digital illustration rather than hyper-photorealism

Wan 2.5 (Preview)

  • + Successfully included all four requested animals (dog, cat, rabbit, fox)
  • + Excellent interpretation of 'tumbling together' with a dynamic, playful action poses
  • + Faithful rendering of 'god rays' and 'dew sparkles' in the morning light
  • Fox eyes have an unnatural, glowing blue artifact
  • Floating water droplets look somewhat artificial
  • Fox kit's front paw is slightly distorted

Verdict: Wan 2.5 (Preview) is the clear winner for its superior prompt adherence, successfully including all four requested animals while Stable Diffusion 3.5 Large Turbo omitted the rabbit. Wan 2.5 also captured the dynamic energy of the scene and more realistically interpreted the 'hyper-photorealistic' instruction compared to the more illustrative and physically merged figures in the SD3.5 output.

Vintage Cafe Logo

Text-to-Image

“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”

Stable Diffusion 3.5 Large Turbo
Wan 2.5 (Preview)

AI Judge Analysis

Stable Diffusion 3.5 Large Turbo

  • + Strong vector aesthetic with bold lines
  • + Good use of high-contrast warm brown tones
  • + Includes a creative hybrid elements of a cloche and a coffee mug
  • Spelling error in the main text ('Caffeé Florin')
  • The cloche looks more like a jar or mug lid than a traditional restaurant cloche

Wan 2.5 (Preview)

  • + Perfect text accuracy including the accent mark
  • + Closer adherence to the 'minimalist' and 'cloche dome' prompt elements
  • + Authentic vintage texture on the background
  • The steam icon is slightly off-center from the cloche handle

Verdict: Wan 2.5 (Preview) produced a superior logo by following the cloche imagery more accurately and maintaining perfect spelling. While Stable Diffusion 3.5 Large Turbo had a nice vector style, it failed on the text rendering despite the simpler layout.

Apollo 11: Journey to Tranquility

Text-to-Image

“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”

Stable Diffusion 3.5 Large Turbo
Wan 2.5 (Preview)

AI Judge Analysis

Stable Diffusion 3.5 Large Turbo

  • + Strong retro artistic aesthetic reminiscent of vintage science fiction
  • + Uses the requested NASA-inspired color palette effectively across the layout
  • + Excellent use of negative space and composition for a poster
  • Text is largely gibberish or misspelled ('Apoll.o', 'Laurch')
  • Fails to provide the specific 6-step iconography requested in the prompt
  • Does not follow the structured list of steps

Wan 2.5 (Preview)

  • + Accurately maps out almost all steps requested in the prompt (Launch, Earth Orbit, Translunar, etc.)
  • + Significantly better text legibility and accuracy in naming mission phases and crew members
  • + Follows the 'flat-vector' style with crisp icons and clear diagrammatic flow
  • The lunar module on the surface is a bit cluttered compared to the other icons
  • The inclusion of a space shuttle-style craft for 'Lunar Orbit' is historically inaccurate for the Apollo mission

Verdict: Wan 2.5 (Preview) is the clear winner as it successfully interprets the prompt's request for a structured, multi-step infographic with readable text and specific iconography. While Stable Diffusion 3.5 Large Turbo creates a visually striking artistic poster, it fails almost entirely on the informational requirements and text accuracy of the challenge.

Next steps

Explore each model