Alibaba's multimodal generation model from the Wan AI suite, supporting text-to-video, image-to-video, reference-to-video with audio, and text-to-image, in both Chinese and English
Settled by community votes across 17 shared challenges, with an AI judge weighing in on each.
Wan 2.6
#28 of 62 in Text-to-Image
Z-Image Turbo
#12 of 62 in Text-to-Image
Where the votes landed
Wan 2.6
50.0%
win rate
Ties
0.0%
Z-Image Turbo
50.0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
Wan 2.6
- + Excellent adherence to lighting instructions with a clear source from the left.
- + Highly realistic textures on the distressed book and the wooden table surface.
- + Perfect placement and rendering of the reflection of the blue sphere on the glass base.
- − The plant is more beside the cube than behind it, though it is still partially visible through the glass.
Z-Image Turbo
- + The green plant is clearly positioned behind the glass cube as requested.
- + Clean, modern aesthetic with sharp lines on the glass edges.
- − The lighting is flat and lacks the 'soft window light' characteristic seen in the other image.
- − The reflection of the red book on the top edge of the glass is physically inconsistent and distracting.
Verdict: Wan 2.6 is the clear winner due to its superior photographic realism, particularly in how it handles the interaction of light and textures. While Z-Image Turbo followed the placement of the plant more literally, the overall image quality and lighting of Wan 2.6 feel much more authentic and visually appealing.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
Wan 2.6
- + Excellent adherence to the 'repairing' aspect of the prompt
- + Superior cinematic lighting and atmospheric rain effects
- + Highly detailed natural skin textures and realistic hand details
- − The raindrops on the jacket appear slightly static and oversized in some areas
Z-Image Turbo
- + Good inclusion of background traffic as requested
- + Matches the 'light rain' prompt without overpowering the scene
- − Subject is just holding the bike rather than repairing it
- − Composition feels flat and lacks the requested cinematic quality or shallow depth of field
- − Missing the requested motion blur for passing cars
Verdict: Wan 2.6 is the clear winner as it fully captured the narrative of the prompt, showing the man actively working on the bike with rich, cinematic textures. Z-Image Turbo failed on several technical requirements, including the shallow depth of field, motion blur, and the specific action of 'repairing'.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
Wan 2.6
- + Exceptional texture detail on skin, including dirt, pores, and sweat.
- + Highly intricate engraving on the armor with realistic metal patina.
- + Perfect adherence to 'hair braided with small beads' and 'lifelike eyes'.
- − The torch flame in the background is slightly blurry compared to the foreground.
Z-Image Turbo
- + Good lighting contrast with the torch in the foreground.
- + Effective focus on the protagonist and decent facial features.
- + Captures the 'battle-worn' aesthetic with facial scarring.
- − The beads in the hair are very small and repetitive, looking less organic than Model A.
- − The skin texture lacks the micro-detail and realism found in the competitor.
- − The armor engraving is slightly softer and less defined.
Verdict: Wan 2.6 is the clear winner due to its incredible rendering of textures, particularly the skin, leather straps, and the frayed cloth underlayer. While Z-Image Turbo produces a competent image, it feels more like a CGI render compared to the lifelike realism and superior detail provided by Wan 2.6.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
Wan 2.6
- + Excellent adherence to the grid-based food photo layout requested in the prompt.
- + Professional typography with distinct sections for Appetizers, Pizza, and Mains.
- + Very clean aesthetic with vibrant color-block accents that enhance the modern feel.
- − Gibberish text in the subtitle and item descriptions.
- − Repeats the 'Pizza' section header twice, once in the grid and once in the list.
Z-Image Turbo
- + Bold, clear sans-serif typography that is very easy to read.
- + Consistent and appetizing food photography throughout the grid.
- + Higher contrast layout that feels energetic and professional.
- − Spelling error in a major heading ('PIZZA MANS').
- − Section headers do not perfectly match the prompt requirements (Mains and SE IIIION instead of Appetizers/Pizza/Mains).
Verdict: Both models followed the prompt well, producing clean, minimalist designs with clear food grids. Wan 2.6 provided a more sophisticated layout with better colored accents and correctly identified all three requested sections, despite some garbled text. Z-Image Turbo had cleaner individual text characters and more readable prices, but the 'PIZZA MANS' typo and odd section naming make it less successful as a professional menu design.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
Wan 2.6
- + Excellent adherence to the 'exploded' requirement with dynamic mid-air suspension.
- + High photorealistic texture on the meat, vegetables, and bun.
- + Text elements are well-integrated into the fiery theme with great font choices.
- − The starburst for the '€6.99' looks a bit like a flat clip-art asset compared to the rest of the 3D scene.
Z-Image Turbo
- + Clean typography and layout suitable for a standard advertisement.
- + Vibrant colors and high-quality rendering of the beef patties.
- + Good use of embers and glowing effects on the text.
- − Failed the 'exploded' part of the prompt; the burger is mostly assembled.
- − The starburst is glowing but lacks the requested 'fiery' texture compared to the title.
Verdict: Wan 2.6 is the clear winner as it followed the 'exploded' instruction perfectly, creating a much more dynamic and interesting composition than the static burger in Z-Image Turbo. Wan 2.6 also achieved a better sense of motion through the scattering of ingredients and more realistic integration of smoke and fire.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
Wan 2.6
- + Natural chalk texture with realistic smudges and dust
- + Excellent handle of the 'handwritten' request with authentic slants and strokes
- + Accurately completed the cutoff item from the prompt text
- − Repeats the price on new lines for the first two items, creating clutter
- − Text alignment is a bit messy and crowded at the bottom
Z-Image Turbo
- + Very clean and legible text layout
- + Perfect spelling for most items with a high degree of clarity
- + Consistent font style throughout
- − Handwriting looks slightly more like a digital font than natural chalk
- − Contains a spelling error: 'Mustroom' instead of 'Mushroom'
- − Lacks the authentic chalk smudging and grime seen in the other model
Verdict: Wan 2.6 captures a much more authentic and atmospheric chalk aesthetic with realistic textures and handwriting variations, although it suffers from repetitive price lines. Z-Image Turbo provides a cleaner, more readable layout but has a spelling error and the text looks a bit too much like a clean digital overlay to be truly convincing as hand-drawn chalk.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
Wan 2.6
- + Excellent cinematic lighting with a glowing nebula and light rays.
- + High level of detail on the space suit and the horse's coat.
- + Great sense of movement and dynamic composition.
- − The horse's back legs blend awkwardly into the dust and ground.
Z-Image Turbo
- + Clean representation of the astronaut and horse subjects.
- + Anatomically correct horse proportions and gear.
- − Failed the 'surreal' and 'cinematic' style prompts, looking more like a basic collage.
- − The background is very flat and lacks the 'space' detail requested.
- − The lighting on the subject does not match the environment.
Verdict: Wan 2.6 is the clear winner as it fully embraces the 'surreal' and 'cinematic' aspects of the prompt, creating a vibrant, cohesive scene with impressive lighting. Z-Image Turbo produced a much flatter, less detailed image that lacked the cosmic scale and atmosphere requested.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
Wan 2.6
- + Excellent photorealism and depth of field
- + Superior light bokeh and texture on the car's exterior
- + The capybara's pose and expression feel more natural and professional
- − The passenger's hand holding the phone is slightly mangled
Z-Image Turbo
- + Successfully includes the seatbelt for the capybara
- + Clearer rendering of the passenger's face
- − The background is quite generic and lacks the vibrant 'Manhattan at night' energy requested
- − The lighting on the capybara is flat compared to the environment
- − The capybara's hand/paw anatomy on the steering wheel is awkward
Verdict: Wan 2.6 is the clear winner due to its superior atmosphere, lighting, and textures, which perfectly capture the 'New York at night' aesthetic. While Z-Image Turbo followed the prompt's logical details well (like the seatbelt), it lacked the professional cinematic quality and detailed background found in Wan 2.6.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
Wan 2.6
- + Perfect text rendering for all requested details and phrases.
- + Strong atmospheric lighting with a cohesive gold and deep blue color palette.
- + Detailed and aesthetically pleasing border incorporating both webs and thorns as requested.
- − The glowing texture on the main title text is slightly grainy.
- − The scroll banner is somewhat small Relative to the image size.
Z-Image Turbo
- + Creative use of torn parchment as the central element.
- + Good implementation of the twisted trees and moody sky background.
- + Clear, legible gothic font style.
- − Spelling error in the location text ('The Archves' instead of 'The Arches').
- − Missing the small scroll banner for the specific phrase 'You are invited to a night of frights'.
- − Composition feels a bit cluttered with the multiple overlapping parchment layers.
Verdict: Wan 2.6 is the superior image because it followed every specific text instruction perfectly, including the location name and the second banner phrase. Its composition is more polished and cinematic compared to Z-Image Turbo, which suffered from a spelling error and failed to place the specific secondary text on its own scroll banner.
Bald man challenge
Image Editing“Give the person a full, thick head of natural hair with realistic texture, density, and a natural hairline. Preserve facial features and lighting.”
AI Judge Analysis
Wan 2.6
- + Successfully applied a full, thick head of hair as requested.
- + Maintained facial features and lighting with high fidelity to the original.
- + The hair texture and lighting match the environment and existing beard.
- − The hairline on the left side of the forehead looks slightly merged with the temple area.
Z-Image Turbo
- + Preserved the original facial features and background perfectly.
- − Failed the primary edit instruction by only adding thin stubble/buzz cut instead of 'full, thick head of hair'.
- − The hairline remains identical to the bald original, just with added darkening.
Verdict: Wan 2.6 followed the instructions perfectly, providing a realistic and aesthetically pleasing full head of hair that integrates well with the original person's appearance. Z-Image Turbo failed the prompt, providing only a very thin layer of stubble that does not meet the 'full, thick' requirement.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
Wan 2.6
- + Perfectly renders the requested text and flag icon
- + Follows the isometric 45° perspective accurately
- + High-quality textures for the rice and fish and a clean diorama base
- − The 'JAPAN' text is slightly off-center to the left
Z-Image Turbo
- + Pleasing soft cartoon aesthetic
- + Very clean, balanced composition
- − Displays the flag of China instead of the flag of Japan
- − Text layout is less aligned with the specific prompt instructions
- − Textures are more simple and less 'realistic PBR' than requested
Verdict: Wan 2.6 is the clear winner as it accurately followed several specific prompt details that Z-Image Turbo missed, most notably the flag of Japan. Wan 2.6 also succeeded in providing the requested refined textures and complex miniature scene, whereas Z-Image Turbo produced a generic cartoon sushi on a plate with the wrong national flag.
Over-the-top cartoon caricature
Editing“Create a caricature of me and my job. Make it exaggerated and humorous, incorporating my profession as a tv show anchor and my love for dogs and hockey.”
AI Judge Analysis
Wan 2.6
- + Excellent adherence to all prompt elements, including the caricature style, news set, hockey stick, and multiple dogs.
- + Successfully captures the likeness of the woman within a stylized caricature format.
- + Clear, vibrant composition with high visual appeal and humorous details like the pug in a jersey.
- − Completely replaces the original image background, though this aligns with the caricature request.
Z-Image Turbo
- + Successfully preserves much of the original image's texture and face detail.
- + Adds a small dog in the background.
- − Completely fails to create a caricature or exaggerated style.
- − Fails to incorporate the profession as a TV anchor or the hockey element.
- − Ignores the 'exaggerated and humorous' instruction.
Verdict: Wan 2.6 followed the creative brief perfectly, transforming the source image into a vibrant caricature that included all requested thematic elements (TV anchor, dogs, and hockey). In contrast, Z-Image Turbo failed to apply any significant edits, essentially returning the source image with a tiny, blurry dog added to the background, ignoring the core request for a caricature and the hockey/anchor themes.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
Wan 2.6
- + Excellent depiction of warm golden light and atmospheric 'god rays'
- + Highly expressive and dynamic poses that suggest movement and play
- + Detailed textures on fur and environment with convincing dew sparkles
- − The fox kit has a slightly awkward paw position on the right side
Z-Image Turbo
- + Successfully includes all four animals with clear visibility
- + Clean focal point with the puppy and bunny in the foreground
- + Vibrant colors and cheerful expressions
- − The lighting feels more studio-like rather than a natural sunrise with god rays
- − A strange anatomical error where the puppy's paw appears to be growing out of its neck/chest area
- − The kitten is partially merged into the puppy's side
Verdict: Wan 2.6 provides a much more cohesive and atmospheric scene, successfully capturing the complexity of the lighting and the playful energy requested in the prompt. Z-Image Turbo suffers from significant anatomical merging issues, particularly where the puppy's paw and the kitten overlap, and fails to deliver the cinematic 'god ray' effect.
Studio Ghibli Anime Style
Editing“Transform this photo into a Studio Ghibli–inspired illustration. Use soft pastel colors, hand-painted textures, gentle lighting, dreamy backgrounds, and a warm, nostalgic mood”
AI Judge Analysis
Wan 2.6
- + Excellent adherence to the Studio Ghibli illustration style
- + Preserves the composition and poses of the iconic meme perfectly
- + Beautiful watercolor textures and soft pastel color palette
- − The faces look more like modern shoujo anime than classic Ghibli character designs
- − Added sparkling bokeh dots that weren't specifically requested
Z-Image Turbo
- + Near-perfect preservation of the original street background and clothing details
- − Completely failed the styling instruction by producing a photorealistic image
- − The girl on the right has a neutral expression rather than the shocked/angry look from the original
- − No artistic transformation or 'Ghibli' aesthetic applied
Verdict: Wan 2.6 successfully transformed the image into a high-quality hand-painted illustration as requested, maintaining the original's humor and composition. Z-Image Turbo failed the style transfer entirely, providing what looks like a slightly modified photo that ignores the Ghibli and pastel instructions.
Golden Hour Stroll
Image Editing“Add dynamic motion to this photo: make hair blow in the wind, add leaves flying, energetic and lively feel.”
AI Judge Analysis
Wan 2.6
- + Excellent hair motion that feels natural and dynamic.
- + High fidelity to the original person and dog's appearance.
- + Leaves are large and clearly visible, contributing to the theme.
- − The wind effect on the hair is slightly asymmetrical, looking more like a Photoshop warp than a natural gust.
- − One leaf in the bottom center looks like a flat graphic overlay.
Z-Image Turbo
- + Successfully added both hair motion and flying leaves.
- + Effective background blur/bokeh gives a stronger sense of depth.
- − Significant changes to the woman's facial features, losing the likeness of the source image.
- − The dog's tail and the background foliage show messy artifacts and structural changes.
- − The dog's leash and the woman's left hand are poorly rendered compared to the original.
Verdict: Wan 2.6 is the clear winner as it successfully applies the 'motion' and 'leaves' edits while perfectly preserving the identity of the subjects and the details of the environment. Z-Image Turbo applied the requested edits but failed the 'source preservation' aspect by significantly altering the woman's face and creating messy artifacts in the background and on the dog's fur.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
Wan 2.6
- + Perfect text rendering for both the name and the banner
- + Includes the requested banner element for the year
- + Superior vintage texture on the background and logo
- − The cloche lacks the traditional handle design seen in better vector emblems
Z-Image Turbo
- + Closer to a minimalist vector icon style
- + Cleaner cloche illustration with better symmetry
- − Failed to include a banner for the 'Est. 1720' text
- − The typography on 'Caffè' is slightly inconsistent in weight and spacing
- − Less background texture than requested
Verdict: Wan 2.6 followed the prompt more accurately by including the requested banner and applying a more visible vintage texture to the background. While Z-Image Turbo has a cleaner vector aesthetic, it missed the banner element and had slightly weaker typography.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
Wan 2.6
- + Legible text for astronaut names
- + Clean, minimal aesthetic
- − Failed to include the infographic steps (launch, orbit, etc.)
- − Background texture looks more like a fabric towel than a vector poster
- − Lacks the requested NASA-inspired muted red and light gray palette elements
Z-Image Turbo
- + Successfully included multiple icons representing mission steps
- + Adhered well to the flat-vector style requested
- + Captured the NASA-inspired color palette effectively
- − Typos in text including 'APOLIO E 11' and 'Descenty'
- − Icons are overlapping or disorganized in layout
- − Missing the specific trajectory arc and some icons are slightly distorted
Verdict: Wan 2.6 failed to follow the core instruction of creating a multi-step infographic, instead producing a very simple commemorative poster. Z-Image Turbo followed the complex instructions for specific icons and steps, even though it suffered from significant spelling errors and some layout clutter. Z-Image Turbo is the clear winner for actually attempting the infographic structure and vector style.
Explore each model
Tongyi-MAI's 6-billion parameter distilled text-to-image model optimized for speed, achieving high-quality generation in 8 steps or fewer with support for bilingual text rendering