Stability AI's 2.5-billion parameter Multimodal Diffusion Transformer with improvements (MMDiT-X) text-to-image model optimized for consumer hardware, featuring improved image quality, typography, and complex prompt understanding
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
Stable Diffusion 3.5 Medium
#56 of 62 in Text-to-Image
Wan 2.7
#39 of 62 in Text-to-Image
Where the votes landed
Stable Diffusion 3.5 Medium
50.0%
win rate
Ties
50.0%
Wan 2.7
0.0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
Stable Diffusion 3.5 Medium
- + Excellent depiction of the plant seen through the glass medium
- + Accurate soft window lighting from the left
- + Vibrant colors and high clarity
- − The sphere is floating unnaturally
- − The cube looks more like a frame or open box rather than a solid object
- − The book is very thin and lacks realistic book textures
Wan 2.7
- + Highly realistic textures on the wooden table and red book
- + Complex and accurate glass reflections including the sphere and plant
- + Solid physical composition with the sphere resting on the bottom
- − The green plant is mostly beside the cube rather than behind it as requested
- − Minor glass internal geometry artifacts on the right side
Verdict: Stable Diffusion 3.5 Medium creates a more artistic and cleaner image, but Wan 2.7 provides significantly better photorealism and physical grounding. Wan 2.7 captures the sophisticated reflections and textures of the glass, wood, and book, making it the superior image despite slightly trailing on the 'plant behind the cube' spatial instruction.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
Stable Diffusion 3.5 Medium
- + Excellent filmic aesthetic with rich colors and high-contrast lighting
- + Strong adherence to the 'imperfect framing' and 'shallow depth of field' requirements
- + Convincing street photography atmosphere with cinematic grain
- − Anatomical errors where the man's hand blends into the bicycle basket
- − The bicycle geometry is slightly warped and unrealistic
Wan 2.7
- + Natural skin textures and realistic clothing details
- + Clearer depiction of the man actually interacting with the bicycle parts
- + Accurate representation of a Japanese street setting
- − Lacks the requested 'motion blur' for passing cars
- − The lighting is somewhat flat and lacks the 'cinematic' mood requested in the prompt
Verdict: Stable Diffusion 3.5 Medium captures the specified 'candid street photo' mood much better with its cinematic lighting and intentional imperfect framing, despite some anatomical merging between the man and the bike. Wan 2.7 is more technically grounded in its anatomy and environmental detail but misses the specific stylistic prompts like motion blur and cinematic depth. Stable Diffusion 3.5 Medium is preferred for its superior adherence to the artistic tone of the prompt.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
Stable Diffusion 3.5 Medium
- + Excellent fine engraving details on the plate armor
- + Subjective intensity in the eyes is very lifelike
- + Warm lighting highlights the metallic surfaces well
- − Missed the request for beads in the hair braids
- − The hair and skin textures look slightly softened or airbrushed
- − The depth of field is a bit flat compared to the other model
Wan 2.7
- + Perfect adherence to specific details like hair beads and facial scars
- + Highly realistic leather and cloth textures as requested
- + Stronger sense of atmosphere with the visible torch and background bokeh
- − The facial expression is a bit more static
- − Light artifacts on the background wall appear slightly digital
Verdict: Wan 2.7 is the clear winner as it followed every specific detail of the prompt, including the small beads in the braids and the leather strap textures, whereas Stable Diffusion 3.5 Medium ignored the beads entirely. Wan 2.7 also achieved a more convincing 'battle-worn' look with distinct scars and a more complex lighting environment.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
Stable Diffusion 3.5 Medium
- + Features a true white background as requested
- + Includes a high number of food photos arranged in a grid
- − Text is illegible garbled characters
- − Food images are blurry and look unappetizing
- − Layout is messy with uneven spacing
Wan 2.7
- + Excellent text legibility and realistic fonts
- + Professional graphic design layout with vibrant orange accents
- + Followed the specific section naming instructions perfectly
- − Background is a stylized tabletop rather than a plain white background
- − Food photos have rounded corners instead of a strict minimalist grid
Verdict: Wan 2.7 produced a commercially viable menu with legible text, logical sections, and high-quality photography, whereas Stable Diffusion 3.5 Medium failed to generate readable words or clear imagery. Although Stable Diffusion adhered closer to the 'white background' instruction, Wan 2.7's superior composition and professional aesthetics make it the clear winner.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
Stable Diffusion 3.5 Medium
- + Excellent photorealistic texture on the bun and patty
- + Consistent and accurate text rendering
- − Failed to produce an 'exploded' view; the burger is mostly assembled
- − Text is plain white/black and lacks the requested fiery glowing effect
Wan 2.7
- + Perfect adherence to the 'exploded' burger concept with suspended components
- + Highly creative and consistent fiery glowing effect on all text elements
- + Excellent dynamic composition with sauces and embers throughout
- − The cucumber slice was not explicitly requested but included
- − Individual food items have a slightly more stylized look compared to the hyper-realism of Model A
Verdict: Wan 2.7 is the clear winner as it followed every complex instruction in the prompt, including the exploded burger layout and the specific fiery glowing text effects. Stable Diffusion 3.5 Medium produced a higher quality texture on the food itself but failed to deconstruct the burger or apply the requested artistic styles to the typography.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
Stable Diffusion 3.5 Medium
- + Strong chalk texture and realistic smudge effects.
- + Accurate handheld framing with a wooden background.
- − Severely failed text rendering with numerous spelling errors and gibberish.
- − Inconsistent layout that creates a cluttered and confusing menu structure.
Wan 2.7
- + Perfect text accuracy, matching every word of the complex menu prompt.
- + Excellent legibility and clean composition.
- + Successfully rendered the specific requested date and prices.
- − The text looks slightly too uniform, leaning towards a 'chalk font' rather than organic handwriting.
- − The chalk texture is a bit too clean and lacks the physical grit seen in real chalkboards.
Verdict: Wan 2.7 is the clear winner because it correctly rendered every specific word and quantity requested in the prompt, including the date and the full description of the chocolate chip cookies. Stable Diffusion 3.5 Medium struggled significantly with text legibility, producing garbled words like 'Ooccluter' and 'Musrroom' which rendered the menu unreadable.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
Stable Diffusion 3.5 Medium
- + Excellent visual realism with cinematic lighting
- + Complex interaction between the horse's legs and the planetary horizon creates a unique surrealist effect
- − Failed the negative constraint entirely by placing the astronaut on top of the horse
- − The astronaut appears to have three legs or an extremely distorted limb
Wan 2.7
- + Clean, high-resolution textures on both the spacesuit and the horse
- + Creative addition of multiple small planets and satellites to enhance the space theme
- − Failed the core specific instruction to place the horse on top of the astronaut
- − The composition feels a bit cluttered with repetitive planet assets
Verdict: Both Stable Diffusion 3.5 Medium and Wan 2.7 failed the challenging semantic constraint of placing the 'horse on top' of the astronaut, both defaulting to the common trope of an astronaut riding a horse. Wan 2.7 is slightly better in terms of technical image quality and anatomy, whereas Stable Diffusion 3.5 Medium has significant anatomical glitches in the astronaut's legs.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
Stable Diffusion 3.5 Medium
- + Excellent photographic lighting and aesthetic.
- + High-quality fur texture and realistic facial expression on the capybara.
- + Strong cinematic composition with beautiful bokeh in the background.
- − Failed to include the phone in the passenger's hands.
- − The passenger is looking forward/at the driver rather than appearing bored and distracted.
Wan 2.7
- + Perfect adherence to the passenger's posture, bored expression, and phone usage.
- + Higher fidelity regarding the specific requested action of paws on the steering wheel.
- + Clearer representation of a New York city street through the window.
- − The capybara's fur has a slightly artificial, brush-like texture compared to Model A.
- − The passenger is sitting in the front seat instead of the back seat as requested.
Verdict: Both models struggled with the spatial instruction of placing the woman in the back seat, as both placed her in the front (either adjacent to or behind the dashboard). While Stable Diffusion 3.5 Medium produces a more cinematically beautiful image with superior textures, Wan 2.7 is the winner for prompt adherence as it successfully captured the 'bored businesswoman on her phone' detail which is central to the prompt's narrative.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
Stable Diffusion 3.5 Medium
- + Features a classic spooky aesthetic with vibrant jack-o-lanterns.
- + Good contrast between the parchment and the dark background.
- − Numerous spelling errors including Halloweeen and Inviloween.
- − The layout is cluttered and the text rendering is distorted.
- − Failed to include the specific scrolls and banner design requested.
Wan 2.7
- + Excellent text rendering with near-perfect spelling for all requested details.
- + Superior composition with a clear central focal point and elegant gothic framing.
- + Followed all complex instructions including the small scroll banner and specific date/time.
- − The lighting on the central pumpkin is slightly flatter than Model A's version.
- − The text 'You are invited...' is slightly off-center on its banner.
Verdict: Wan 2.7 significantly outperforms Stable Diffusion 3.5 Medium by correctly rendering all the requested text and event details with a high degree of legibility. While Stable Diffusion 3.5 Medium captures a moody atmosphere, its failure to spell basic words and follow the layout instructions makes it unusable as an invitation. Wan 2.7 produced a polished, professional-looking design that perfectly adheres to the gothic vintage prompt.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
Stable Diffusion 3.5 Medium
- + Excellent texture on the ikura (fish roe) and seaweed nori wrapping.
- + Clean lighting and consistent with the requested 'soft refined textures'.
- − Failed to include the requested flag icon.
- − The text 'SUSHI' has minor rendering artifacts on the letter 'S'.
- − Did not include a 'raised diorama base', putting the plate directly on the background.
Wan 2.7
- + Perfect adherence to all prompt elements, including the flag icon and the raised diorama base.
- + Extremely clean text rendering with professional-looking typography.
- + Stronger 3D isometric composition that captures the 'miniature scene' aesthetic.
- − The textures look slightly more plastic-like compared to the realistic PBR rice in the other model.
- − The shrimp tail has a slightly unusual shape.
Verdict: Wan 2.7 is the clear winner as it followed every instruction in the prompt, including the specific diorama base, flag icon, and complex layout. While Stable Diffusion 3.5 Medium produced nice surface textures, it missed several key components of the scene and had minor issues with the text rendering.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
Stable Diffusion 3.5 Medium
- + Vibrant color palette with high saturation
- + Expressive and consistent large eyes across animals
- + Good use of bokeh and foreground floral elements
- − Failed to include the bunny/rabbit completely
- − The kitten looks like a hybrid creature with fox-like ears
- − Anatomical issues with the fox's front leg and paw structure
Wan 2.7
- + Accurately included all four requested animals
- + Beautiful light interaction with visible god rays and dew sparkles
- + Dynamic composition that captures the 'playfully chasing' action
- − The fox's face lacks the 'baby' features requested, looking more like an adult
- − Slightly less 'masterpiece' clarity in the fur textures compared to a hyper-photorealistic style
Verdict: Wan 2.7 is the clear winner as it successfully incorporated all four specific animals requested, whereas Stable Diffusion 3.5 Medium missed the bunny entirely and produced a strange kitten/fox hybrid. Wan 2.7 also better captured the atmospheric elements of the prompt like the dew sparkles and god rays, resulting in a more cohesive and accurate scene.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
Stable Diffusion 3.5 Medium
- + Strong hand-drawn vintage aesthetic
- + Highly detailed engraving style
- + Captures the warm brown and cream tones well
- − Text spelling is incorrect (Florrian instead of Florian)
- − Date is incorrect (Est 170 instead of Est 1720)
- − The cloche dome is stylized to the point of being abstract
Wan 2.7
- + Perfectly accurate text for the name and establishment date
- + Clean vector-style execution following the minimalist prompt
- + Clearly depicts the cloche dome with steam as requested
- − The 'Florion' spelling in the second line is a slight typo from Florian (though the first word is correct)
- − Composition is a bit generic compared to the artistic merit of Model A
Verdict: Wan 2.7 followed the prompt much more accurately, successfully including all requested elements like the 'Est. 1720' banner and the cloche dome in a clean vector style. Stable Diffusion 3.5 Medium produced a beautiful artistic illustration but failed significantly on the text accuracy and the specific date requested.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
Stable Diffusion 3.5 Medium
- + Successfully uses the requested navy and muted red color palette.
- + Achieves a clean, modern vector aesthetic.
- − Text rendering is poor with many garbled or nonsensical words.
- − Fails to follow the chronological order of the requested steps.
- − Iconography is inconsistent and does not match the specific descriptions given.
Wan 2.7
- + Excellent prompt adherence for all six steps with relevant, high-quality icons.
- + Cleverly includes additional details like astronaut names and mission dates while maintaining a clean layout.
- + Superior text legibility and alignment with the 'modern infographic' style.
- − Minor spelling errors in supporting text (e.g., 'Tranquiliry', 'Descript').
Verdict: Wan 2.7 is the clear winner as it perfectly captured the structure and intent of a step-by-step infographic, adhering to all specific icon requests from launch to landing. While Stable Diffusion 3.5 Medium produced a visually pleasing color field, it failed to organize the content into a coherent sequence and had significant issues with text generation.
Explore each model
Alibaba's Wan 2.7 image generation and editing model for text-to-image, reference-guided generation, and instruction-based image edits