Alibaba's text-to-image and image-to-image generation model from the Wan AI suite, offering high-quality visual generation capabilities
Settled by community votes across 17 shared challenges, with an AI judge weighing in on each.
Wan 2.5 (Preview)
#27 of 62 in Text-to-Image
Z-Image Turbo
#12 of 62 in Text-to-Image
Where the votes landed
Wan 2.5 (Preview)
0%
win rate
Ties
0%
Z-Image Turbo
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
Wan 2.5 (Preview)
- + Excellent handling of glass physics including refraction and internal reflections
- + Captures the request for light from the left with realistic shadows and highlights
- + Detailed textures on the book and table add to realism
- − The sphere is quite large relative to the cube
- − Small particles or dust motes in the air might be distracting to some users
Z-Image Turbo
- + Successfully includes all prompt elements in a clean layout
- + Accurately represents the sphere as 'small' relative to the cube
- + Logical composition with the plant in the background
- − Lighting is somewhat flat compared to the other model
- − The book seems to float slightly on the left edge of the cube
- − Less realistic refraction through the glass
Verdict: Wan 2.5 (Preview) produces a much more visually striking and realistic image with superior lighting and glass refraction. While Z-Image Turbo captures the scale of the 'small sphere' better, it lacks the photographic depth and detailed texture work found in Wan 2.5.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
Wan 2.5 (Preview)
- + Excellent shallow depth of field and bokeh realism.
- + Superior skin texture and natural facial features.
- + Accurate bike repair interaction with visible tools and convincing physical posture.
- − The visible rain looks a bit like static streaks rather than natural droplets in some areas.
- − Slightly lacks the 'motion blur from passing cars' specifically requested.
Z-Image Turbo
- + Successfully captures the motion blur of passing cars in the background.
- + Good colors and saturated red on the bicycle frame.
- − The subject is walking with the bike rather than 'repairing' it as requested.
- − The skin texture and lighting appear flat and less realistic than its competitor.
- − Poor composition with a car cutting through the base of the frame awkwardly.
Verdict: Wan 2.5 (Preview) produced a much more cinematic and high-quality image that accurately depicted a man repairing a bike with tools and appropriate posture, whereas Z-Image Turbo showed a man simply pushing a bike. Wan 2.5 (Preview) also followed the technical instructions for shallow depth of field and skin texture far better, resulting in a significantly more realistic photograph.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
Wan 2.5 (Preview)
- + Excellent detail on the engraved plate armor and leather texture
- + Realistic and expressive facial features with lifelike skin texture
- + Strong adherence to the bokeh sparks and torchlight reflection prompt
- − The beads in the hair look more like metallic studs or pins rather than woven beads
- − The background torch is slightly distracting in its proximity to the face
Z-Image Turbo
- + Successfully captures a more mature, battle-hardened appearance
- + Good integration of the braided hair and beads into the character design
- + Atmospheric lighting with clear bokeh sparks
- − The facial skin texture is slightly smoother and less detailed compared to Model A
- − The engraving on the chest plate looks slightly flatter and less defined
Verdict: Wan 2.5 (Preview) produces a superior close-up with incredible detail on the armor engravings, leather straps, and realistic skin texture. While Z-Image Turbo captures the character's 'battle-worn' essence well with a more mature look, Wan 2.5 (Preview) is the overall winner due to its exceptional technical fidelity and lighting.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
Wan 2.5 (Preview)
- + Excellent adherence to the grid layout for food photos.
- + Clean, modern use of white space and professional sans-serif typography.
- + Logical compartmentalization of sections (Appetizers, Pizza, Mains) with distinct color-coded separators.
- − Nonsense filler text and some misspellings like 'Menue'.
- − Food images are somewhat repetitive in styling and garnish.
Z-Image Turbo
- + Strong bold typography that feels very modern and high-impact.
- + High quality and variety in the food photography within the grid.
- + Good use of vibrant orange accents to create a cohesive brand identity.
- − Layout is slightly cluttered compared to the requested minimalism.
- − The header 'PIZZA MANS' is a significant typo.
- − Sectioning is a bit confusing with 'SETIIION' and overlapping categories.
Verdict: Wan 2.5 (Preview) produced a more balanced and professional layout that truly captures the 'minimalist' requirement, using clean lines and structured sections. While Z-Image Turbo has more visually appetizing food photography, its layout is slightly more cramped and the text errors are more jarring.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
Wan 2.5 (Preview)
- + Excellent adherence to the 'exploded' layout with clearly separated components.
- + Highly creative typography with a dripping cheese/lava effect that fits the theme.
- + Photorealistic texture on the meat patty and fresh vegetables.
- − The '€6.99' text is slightly less cohesive with the overall artistic style of the upper text.
Z-Image Turbo
- + Strong, clean text rendering for all three requested messages.
- + Warm lighting and vibrant colors create an appetizing look.
- + Good interpretation of the dark, fiery background with glowing embers.
- − Failed the 'exploded' layout requirement, showing a mostly assembled burger.
- − Lower level of fine detail in the burger textures compared to the competitor.
Verdict: Wan 2.5 (Preview) and Z-Image Turbo both handled the text integration well, but Wan 2.5 (Preview) is the superior image for strictly following the 'exploded' composition prompt. Wan 2.5 also featured more realistic ingredient textures, whereas Z-Image Turbo produced a standard stacked burger that missed the dynamic motion requested.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
Wan 2.5 (Preview)
- + Excellent chalk texture with realistic smudging and dusting on the board.
- + Strong perspective and depth of field within the café environment.
- + Very natural-looking cursive handwriting that flows well.
- − Failed to include 'Herbs' in the second item.
- − Duplicated the price '$9' for the cookies and broke the text awkwardly across lines.
Z-Image Turbo
- + Followed the text requirements more accurately, including almost the entire prompt content.
- + Good legibility and layout for a menu board.
- + Realistic letter grain that mimics actual chalk strokes.
- − Spelling error in 'Mustroom' for 'Mushroom'.
- − The handwriting style is a bit too uniform, leaning towards a digital font aesthetic compared to Model A.
- − Limited depth and environmental context compared to the other model.
Verdict: Both models captured the essence of a chalkboard menu, but Wan 2.5 (Preview) produced a more artistically convincing image with superior chalk textures and lighting, despite some text repetition errors. Z-Image Turbo followed the text content of the prompt more closely, including the word 'Herbs', but suffered from a spelling error and a flatter, less immersive composition.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
Wan 2.5 (Preview)
- + Excellent cinematic lighting and background details including a galaxy and planet.
- + High resolution with realistic textures on the space suit and horse hair.
- + Dynamic composition with a sense of motion.
- − Failed the negative constraint to have the horse on top of the astronaut.
Z-Image Turbo
- + Clean rendering of the astronaut and horse equipment.
- + Good anatomical proportions for the horse.
- − Failed the specific spatial instruction for the horse to be on top.
- − The background is quite empty and lacks the cinematic depth requested.
- − Visible anatomical glitch with five legs or a duplicate limb near the rear.
Verdict: Both models failed the specific 'horse on top' spatial instruction, which is a common challenge for T2I models with unusual positional requests. However, Wan 2.5 (Preview) is the superior image due to its stunning cinematic background, superior lighting, and dynamic movement, whereas Z-Image Turbo has a plain background and a significant anatomical error with the horse's legs.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
Wan 2.5 (Preview)
- + Excellent atmospheric lighting with realistic rain droplets on the windshield and vibrant city lights.
- + Clearly captures both the capybara and the passenger in a balanced composition.
- + High level of detail in the fur texture and the taxi hood.
- − The capybara's paws look slightly reptilian or unnatural.
- − The perspective is a bit confusing, showing the taxi's roof sign from the front while being inside the car.
Z-Image Turbo
- + The capybara's hands/paws are more realistically rendered and have a better grip on the steering wheel.
- + The side-profile composition feels very natural and cinematic for an interior car shot.
- + Accurately places the passenger in the back seat with a perfect bored expression.
- − The lighting is a bit flat compared to the requested 'Manhattan at night' feel.
- − The background is somewhat generic and lacks the vibrant, blurred city lights described in the prompt.
Verdict: Wan 2.5 (Preview) produced a more visually stunning image with beautiful rain effects and lighting, though it struggled with the anatomical accuracy of the paws. Z-Image Turbo followed the physical instructions better, providing a more grounded and realistic side-view composition with superior paw rendering, even if the atmosphere was less 'New York at night'.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
Wan 2.5 (Preview)
- + Excellent typography rendering with perfect spelling and coherent fonts.
- + Strong cinematic lighting with a vibrant, glowing jack-o-lantern.
- + Clean and professional composition that feels like a completed graphic design piece.
- − The 'You are invited' text is slightly small for the banner it inhabits.
- − The parchment has a very clean, modern edge rather than a rugged vintage feel.
Z-Image Turbo
- + Features a great weathered parchment texture with torn edges.
- + Captures the 'twisted trees' and 'moody night sky' prompt elements more prominently in the background.
- − Contains a typo in the location ('The Archves' instead of 'The Arches').
- − The additional banner scrolls are floating awkwardly and used for the main title instead of the invite text.
- − The resolution and sharpness are lower than the competitor.
Verdict: Wan 2.5 (Preview) produced a far more polished and usable design with 100% accurate text and beautiful cinematic lighting. While Z-Image Turbo did well with the vintage texture of the paper, its spelling error and awkward placement of the banner elements make it less successful as a functional invitation.
Bald man challenge
Image Editing“Give the person a full, thick head of natural hair with realistic texture, density, and a natural hairline. Preserve facial features and lighting.”
AI Judge Analysis
Wan 2.5 (Preview)
- + Excellent adherence to the 'full, thick head of hair' instruction
- + Perfectly preserves the person's identity, glasses, and facial expression
- + Integrates the new hair seamlessly with the existing Beard and lighting
- − The hairline on the forehead is slightly too sharp/perfect
Z-Image Turbo
- + Maintains the overall composition and color palette
- − Completely failed the primary instruction to add a full head of hair
- − Removed the subject's glasses which was not requested
- − Significantly altered the facial structure and eyes, losing person's identity
Verdict: Wan 2.5 (Preview) successfully fulfilled the edit request by providing a realistic and thick head of hair while perfectly preserving the identity of the person in the source image. Z-Image Turbo failed the prompt entirely, leaving the subject bald, removing their glasses, and changing their facial features. Wan 2.5 (Preview) is the clear winner for both instruction following and source preservation.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
Wan 2.5 (Preview)
- + Excellent typography with clean, bold font and perfect alignment.
- + Accurately represents the Japanese flag as requested.
- + Features premium 3D rendering with soft shadows and refined textures.
- − The lighting in the background has a slight gradient rather than being a perfectly solid flat color.
Z-Image Turbo
- + Clean isometric perspective on a layered diorama base.
- + Good 3D cartoon aesthetic with soft material quality.
- − Incorrect flag icon (shows the flag of China instead of Japan).
- − The text 'SUSHI' is slightly off-center compared to the 'JAPAN' header.
- − The salmon texture has tiny artifacts that look like air bubbles or pores.
Verdict: Wan 2.5 (Preview) provided a significantly more accurate and professional result, following all prompt instructions including the correct flag and perfectly aligned typography. Z-Image Turbo failed the specific geographic requirement by placing a Chinese flag next to text about Japan, and the overall composition felt less polished.
Over-the-top cartoon caricature
Editing“Create a caricature of me and my job. Make it exaggerated and humorous, incorporating my profession as a tv show anchor and my love for dogs and hockey.”
AI Judge Analysis
Wan 2.5 (Preview)
- + Excellent adherence to all prompt requirements including caricature style, hockey, dogs, and TV anchor profession.
- + Strong caricature art style with expressive, exaggerated features that maintain the likeness of the source subject.
- + Creative composition that integrates all story elements into a cohesive scene.
- − The transition to a flat vector illustration style loses the photographic texture of the original face.
- − Minor artifacts in the hockey screen content where the anatomy of the player is slightly warped.
Z-Image Turbo
- + High preservation of the original facial features and photographic quality.
- + Subtle addition of a dog in the background.
- − Completely failed the 'caricature' instruction, providing a realistic photo instead.
- − Missing the hockey and TV anchor elements requested in the prompt.
- − The requested edit was largely ignored except for a small background addition.
Verdict: Wan 2.5 (Preview) followed every detail of the prompt, successfully translating the subject into a humorous caricature that incorporated hockey, a news desk, and dogs. Z-Image Turbo largely failed the task, producing a standard photograph that missed almost all of the requested thematic elements. Wan 2.5 (Preview) is the clear winner for following the complex creative instructions.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
Wan 2.5 (Preview)
- + Dynamic sense of movement and action
- + Vibrant colors and strong 'god rays' lighting
- + High level of detail in the grass and dew sparkles
- − The fox's eyes appear unnaturally large and glowing
- − Floating water droplets look artificial and disconnected from the scene
- − The cat has an extra-long, awkwardly positioned tail
Z-Image Turbo
- + More realistic anatomical proportions for the animals
- + Softer, more natural integration of the animals into the environment
- + Charming, cohesive composition where the animals are interacting
- − The lighting is a bit hazy and lacks the requested '8K masterpiece' sharpness
- − Fewer butterflies and less 'playful chasing' action compared to the other model
Verdict: Both models followed the prompt well, featuring all four requested animals. Wan 2.5 (Preview) offers a more high-energy, fantastical aesthetic with intense lighting, but suffers from anatomical oddities like the fox's eyes. Z-Image Turbo provides a more grounded and realistic interpretation with better species features, making it the more visually pleasing and coherent image despite having slightly less vibrant colors.
Studio Ghibli Anime Style
Editing“Transform this photo into a Studio Ghibli–inspired illustration. Use soft pastel colors, hand-painted textures, gentle lighting, dreamy backgrounds, and a warm, nostalgic mood”
AI Judge Analysis
Wan 2.5 (Preview)
- + Excellent adherence to the Studio Ghibli art style with proper linework and cel-shading.
- + Perfectly captures the requested soft pastel color palette and warm, nostalgic mood.
- + Preserves the composition and character expressions of the original meme accurately.
- − The man's hand on his hip looks slightly anatomically awkward compared to the original.
Z-Image Turbo
- + Successfully preserves the identity and faces of the people in the original photograph.
- + Maintains high structural consistency with the source image.
- − Completely fails the primary instruction to transform the image into a Studio Ghibli–inspired illustration.
- − The image remains a realistic photograph with only minor color adjustments.
- − The background remains blurry photographic bokeh rather than a 'dreamy background' illustration.
Verdict: Wan 2.5 (Preview) followed the prompt perfectly, delivering a high-quality illustration that captures the specific essence of Studio Ghibli while maintaining the humor of the original meme. Z-Image Turbo largely ignored the stylistic instructions, providing what looks like a slightly filtered photograph rather than the requested art style.
Golden Hour Stroll
Image Editing“Add dynamic motion to this photo: make hair blow in the wind, add leaves flying, energetic and lively feel.”
AI Judge Analysis
Wan 2.5 (Preview)
- + Excellent adherence to the 'hair blowing in wind' instruction.
- + Near-perfect source preservation of the subject's face and clothing.
- + Adds motion blur to the dog's tail which enhances the energetic feel.
- − The added leaves look like flat digital stickers rather than natural elements.
- − The color of the added leaves is an oversaturated neon green that doesn't match the scene.
Z-Image Turbo
- + Natural-looking falling leaves that match the colors and lighting of the environment.
- + Good preservation of the overall composition and background.
- − The hair edit is very subtle and doesn't clearly convey strong wind.
- − The model changed the subject's facial features, making her look like a different person.
- − The dog's leash has been completely removed, breaking the logic of the image.
Verdict: Wan 2.5 (Preview) is the winner because it followed the specific motion instructions much more effectively, particularly with the wind-blown hair and motion-blurred tail, while keeping the subject's face identical to the source. Although Z-Image Turbo rendered more realistic leaves, it failed to preserve the subject's identity and unnaturally removed the leash.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
Wan 2.5 (Preview)
- + Excellent typography with correct inclusion of the accent mark in 'Caffè'
- + Superior texture and 'vintage' feel with the paper background and border
- + Perfect adherence to all prompt elements including the banner
- − The cloche design is slightly more complex than a 'minimalist' logo might require
Z-Image Turbo
- + Strong minimalist aesthetic with clean vector lines
- + Good balance between the icon and text elements
- − Missed the 'banner' requirement for the Est. 1720 text
- − Minor typo in 'Caffè' where the accent mark is incorrectly placed/rendered
- − Less visual interest in the background compared to image A
Verdict: Wan 2.5 (Preview) produced a more polished and professional logo that followed all prompt instructions, including the banner and the specific typography of the brand name. Z-Image Turbo captures the minimalist spirit well but fails on technical details like the banner requirement and the correct placement of the accent mark in 'Caffè'.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
Wan 2.5 (Preview)
- + Excellent typography and spelling throughout the infographic.
- + Clear, logical flow of information following the requested steps.
- + Sophisticated vector style that accurately represents the NASA color palette.
- − The 'Descent' and 'Landing' text blocks lack their specific icons and are somewhat floating.
- − The Saturn V rocket interpretation adds a space-shuttle-like craft on top, which is historically inaccurate.
Z-Image Turbo
- + Clean, minimalist layout that adheres to the 'flat-vector' request.
- + Good use of the requested color palette on a white background.
- − Multiple spelling errors in the headers like 'APOLIO E 11' and 'Translurian'.
- − Poor iconography consistency and failed to include all six requested steps.
- − Incorrect rocket anatomy with a red bulbous nose cone.
Verdict: Wan 2.5 (Preview) is the clear winner as it produces a professional, legible, and organized infographic with correct spelling and a logical structure. Z-Image Turbo suffers from significant text errors and fails to include several of the requested stages of the mission.
Explore each model
Tongyi-MAI's 6-billion parameter distilled text-to-image model optimized for speed, achieving high-quality generation in 8 steps or fewer with support for bilingual text rendering