Alibaba's Qwen Image 2.0 model with enhanced text rendering, supporting both Chinese and English prompts with up to 6 images per request
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
Qwen Image 2.0
#34 of 62 in Text-to-Image
Wan 2.6
#27 of 62 in Text-to-Image
Where the votes landed
Qwen Image 2.0
0%
win rate
Ties
0%
Wan 2.6
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent photorealistic texture on the book and table surface.
- + Accurate representation of refraction and reflections within the glass panels.
- − The plant appears to be fused with or immediately behind the glass rather than having clear depth.
- − The sphere is floating unnaturally in the center of the cube without support.
Wan 2.6
- + Natural composition with a potted plant clearly positioned behind the object.
- + Beautiful lighting with soft shadows and realistic caustic reflections on the table.
- + The sphere is realistically resting on the bottom surface of the cube.
- − The glass cube is slightly less sharp in terms of edge definition compared to Model A.
Verdict: Wan 2.6 is the superior image as it creates a more believable and aesthetically pleasing scene. While Qwen Image 2.0 has impressive micro-details on the book texture, Wan 2.6 handles the spatial relationship between the plant and cube much better and provides a more realistic placement for the blue sphere.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent natural skin texture and realistic age spots on the subject
- + Successful capture of 'imperfect framing' with the candid crop
- + Very authentic wet pavement reflections and subtle rain effects
- − The rear wheel of the bicycle is significantly warped and lacks structure
- − The hands and fingers of the subject are anatomically messy while working
Wan 2.6
- + Better overall composition and cinematic atmosphere
- + The red bicycle is rendered with much higher structural accuracy
- + Excellent execution of light rain and rain droplets on surfaces
- − The 'rain' on the man's jacket looks more like glass beads than water droplets
- − The skin texture feels slightly more processed and digital compared to model A
Verdict: Qwen Image 2.0 excels at the 'natural skin texture' and 'imperfect framing' aspects of the prompt, feeling truly like a candid snapshot, though the bicycle's geometry is quite broken. Wan 2.6 provides a much more coherent background and a far better rendered bicycle, though its rain effects on clothing look slightly unnatural. Wan 2.6 is the stronger overall image due to its technical consistency and superior detail in the mechanical objects.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent portrayal of 'battle-worn' with realistic skin texture, deep wrinkles, and scars.
- + Highly detailed engraving on the plate armor and clear representation of beads in the hair.
- + Strong color contrast and sharp focus on the character's facial features.
- − The lighting feels more like a large fire rather than specific 'warm torchlight' as requested.
- − Small anatomical glitches on the hand resting on the hilt.
Wan 2.6
- + Perfect interpretation of torchlight lighting, creating realistic shadows and warm highlights.
- + Superior rendering of detailed leather straps and the frayed cloth underlayer as requested.
- + Highly lifelike, glossy eyes that capture the specified emotion and lighting.
- − The scale mail/cloth textures at the bottom left are slightly blurry compared to the face.
- − The 'battle-worn' aspect is mostly communicated through mud splatters rather than physical scarring/exhaustion.
Verdict: Wan 2.6 is the winner as it followed the technical details of the prompt more precisely, particularly the lighting and the specific textures for the leather and cloth underlayer. While Qwen Image 2.0 did an excellent job with the character's face, Wan 2.6 achieved a more cinematic and cohesive look that perfectly captured the requested 'torchlight' atmosphere.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent photo-realistic food imagery
- + Logical alignment of sections with relevant photos
- + Clean and readable section headers
- − Garbled and unreadable body text
- − Repetitive pricing values
- − Visual layout feels like a catalog rather than a functional menu
Wan 2.6
- + Accurately depicts a menu layout with actual price lists and descriptions
- + Incorporates vibrant geometric accents as requested
- + Better hierarchy between titles, subtitles, and price points
- − Confusing categorization where pizzas are shown under the 'Appetizers' heading
- − Less consistent photo quality compared to Model A
- − Text is still largely gibberish
Verdict: Qwen Image 2.0 produces significantly higher quality food photography and a cleaner grid, but it lacks the functional layout elements of a real menu. Wan 2.6 better understands the structural requirements of a menu design by including price columns and geometric accents, though it fails to correctly categorize the food images under their respective headings. Qwen Image 2.0 is the preferred choice for its visual fidelity and clarity in minimalism.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent text legibility and high-quality font rendering.
- + Professional advert layout with the price starburst visually balanced.
- + Sharp textures on the burger ingredients and glowing fire effects.
- − The burger is less 'exploded' and looks more like a standard stack with a floating bun.
- − Background is slightly more generic with less environmental depth.
Wan 2.6
- + Successfully captures the 'exploded' motion with components scattering realistically in space.
- + Dynamic use of smoke and embers creating a sense of heat and depth.
- + Good 3D effect on the main title text.
- − The price starburst looks like a flat clip-art element that clashes with the photorealistic scene.
- − Small visual glitch with a sauce drip connecting the bun to the burger unnaturally.
Verdict: Both models followed the complex prompt well, but Qwen Image 2.0 produced a more polished final advertisement with superior integration of the text elements and price starburst. While Wan 2.6 achieved a better 'exploded' effect for the burger components, the flat graphic style of its price tag and the odd sauce connector make it feel less cohesive than the Qwen output.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
Qwen Image 2.0
- + Text is perfectly legible and correctly spelled across all items.
- + The background café scene is well-lit and adds to the cozy atmosphere.
- + Successfully interprets the cut-off prompt to complete the third item as logical 'Chocolate Chip Cookies'.
- − The handwriting looks slightly digital/synthetic despite the chalk texture.
- − The chalkboard lacks a physical frame, making it look like a floating green panel.
Wan 2.6
- + Superior chalk texture with realistic dust, smudges, and varying pressure.
- + Authentic cursive title that better matches the 'elegant cursive' prompt instruction.
- + The composition feels more grounded with a visible wooden frame and chalk dust on the ledge.
- − Warped characters in the bottom line of text (e.g., 'options').
- − Minor inconsistencies in common letterforms across different words.
Verdict: Both models followed the complex text requirements with impressive accuracy. While Qwen Image 2.0 produced cleaner and more legible text, Wan 2.1 is the winner due to its superior artistic rendering of chalk textures and the more realistic, framed presentation of the chalkboard. Wan 2.1 also captured the 'elegant cursive' style for the title much more effectively than the simpler handwriting in Qwen Image 2.0.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
Qwen Image 2.0
- + Crisp and clear visual quality with interesting scaly textures on the horse neck.
- + Dynamic composition with floating water droplets that add to the surreal theme.
- − Completely failed the negative constraint to have the horse on top of the astronaut.
Wan 2.6
- + Beautiful cinematic lighting and vibrant nebulae in the background.
- + Detailed rendering of the space suit and saddle leather.
- − Completely failed the negative constraint to have the horse on top of the astronaut.
- − Anatomical oddity with a third leg or leg fragment appearing near the back of the horse.
Verdict: Both models failed the specific spatial logic requested in the prompt ('horse on top, not vice versa'), instead providing standard images of an astronaut riding a horse. Qwen Image 2.0 is the preferred choice as it avoids the major anatomical artifacts found in Wan 2.6, such as the extra leg segments, and provides a cleaner overall image.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent texture on the capybara's fur and the jacket fabric.
- + Realistic phone lighting reflecting on the woman's face.
- + Good sense of depth with the blurred city lights and passing car.
- − The transition between the capybara's paws and the steering wheel looks slightly AI-generated and messy.
- − The perspective makes the capybara look significantly larger than the person.
Wan 2.6
- + Perfect adherence to the prompt regarding the woman's bored expression.
- + Includes great environmental details like raindrops on the windshield and taxi roof sign.
- + Stable hand and paw anatomy on the steering wheel.
- − The capybara's head looks a bit stiff and cutout-like against the background.
- − The interior cabin feels slightly cramped for a New York taxi.
Verdict: Both models followed the complex prompt effectively, but Wan 2.6 provided a more cinematic atmosphere with the addition of rain and a perfectly executed 'bored' expression for the passenger. While Qwen Image 2.0 has superior fur texture, Wan 2.6 feels more like a cohesive movie scene and adheres better to the specific character expressions requested.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent typography with a clean, classic gothic font
- + Clean and highly legible text layout
- + Strong parchment texture and a cohesive vintage aesthetic
- − Lighting is a bit flat compared to the 'cinematic' request
- − The border feels more like a sketch than a physical frame
Wan 2.6
- + Atmospheric cinematic lighting with a strong orange glow and deep blue sky
- + Intricate 3D-looking border with layered thorns and webs
- + Gold-leaf effect on the title text adds an elegant touch
- − Small text at the bottom is slightly less crisp than Model A
- − Text alignment in the banner is slightly off-center
Verdict: Both models followed the prompt exceptionally well, rendering all requested text perfectly. Qwen Image 2.0 offers a cleaner, more graphic 'invitation' feel with superior legibility, while Wan 2.6 provides a more immersive, atmospheric piece with impressive lighting and texture depth. Wan 2.6 is the likely winner for effectively capturing the 'cinematic lighting' and 'polished' aspects of the prompt.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent photographic realism in the sushi textures.
- + Perfect text rendering of 'JAPAN' and 'SUSHI'.
- + Clean, professional composition.
- − Failed the '3D cartoon' and 'miniature isometric' style request, looking like a real photo instead.
- − Perspective is not a true 45-degree isometric projection.
Wan 2.6
- + Perfectly captures the '3D cartoon' and 'isometric' aesthetic requested.
- + Follows the instruction for a 'small raised diorama base' much better than Model A.
- + Accurate text and flag placement.
- − The textures are more simplified compared to the 'realistic PBR' request.
- − Slightly less crisp resolution on the text edges compared to Model A.
Verdict: While Qwen Image 2.0 produces a photorealistic image with perfect black typography, it largely ignores the '3D cartoon' and 'isometric miniature' stylistic requirements. Wan 2.6 accurately interprets the full prompt, delivering a stylized diorama with the correct 45-degree perspective and 3D miniature feel. Wan 2.6 is the preferred choice for following the specific creative direction of the prompt.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent interaction between the animals, with a realistic 'tumbling' dynamic.
- + Vibrant and detailed wildflowers that fill the foreground effectively.
- + Strong adherence to the fox kit and tabby kitten descriptions.
- − The fox's face/mouth looks slightly distorted as it rolls over.
- − The butterfly on top of the kitten's head looks a bit flat and pasted on.
Wan 2.6
- + Exceptional lighting with beautiful god rays and sparkling dew effects as requested.
- + Very clean 'masterpiece' quality with professional-grade bokeh.
- + Highly expressive and cute facial features on all animals.
- − The kitten and fox are positioned slightly repetitively next to each other.
- − The floating seeds/particles in the air are a bit distracting.
Verdict: Both models performed exceptionally well on this complex prompt, capturing all four specific animals and the morning atmosphere. Qwen Image 2.0 excels at the physical interaction and 'tumbling' of the animals, while Wan 2.1 produces a superior aesthetic with its handling of light, dew sparkles, and overall photographic clarity. Wan 2.1 is the likely winner for its better interpretation of the atmospheric 'god rays' and '8K masterpiece' quality.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent typography with proper accent mark and spacing.
- + Clear banner implementation that adheres to the requested layout.
- + Clean vector-style shading and a pleasant warm color palette.
- − The placement of the steam inside the cloche is conceptually strange compared to a standard logo.
- − The line work on the banner scroll is a bit cluttered.
Wan 2.6
- + Elegant minimalist composition that better reflects a modern vector emblem.
- + Beautiful subtle texture application around the edges of the light background.
- + Creative and realistic placement of the steam detail above the cloche.
- − The banner is very small and tucked to the side rather than being a primary structural element.
- − Missing the grave accent on the letter 'e' in 'Caffè'.
Verdict: Qwen Image 2.0 followed the specific layout instructions more accurately, including a large legible banner and perfect spelling of the brand name. However, Wan 2.6 produced a far more professional-looking minimalist logo with superior composition and texture, despite the small spelling error and tiny banner. Qwen Image 2.0 is the winner for strict prompt adherence, particularly regarding the text and banner placement.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent adherence to the sequential steps requested in the prompt.
- + Accurate and legible text rendering for mission phases.
- + Clean vector aesthetic with a correct NASA-inspired color palette.
- − Includes a spelling error in 'Translunjar' (Translunar).
- − Vertical layout is slightly cluttered at the bottom.
Wan 2.6
- + Successfully incorporated the astronaut names with correct spelling.
- + Maintains the requested color palette.
- − Completely failed to include the requested infographic steps and icons.
- − Very Poor composition with large amounts of empty space.
- − Does not meet the 'infographic' requirement for mission phases.
Verdict: Qwen Image 2.0 successfully followed the complex instructions to create a multi-step infographic with specific iconography, despite a minor typo. In contrast, Wan 2.6 failed to include almost all of the critical mission steps and details, resulting in a largely empty and unsuccessful design.
Explore each model
Alibaba's multimodal generation model from the Wan AI suite, supporting text-to-video, image-to-video, reference-to-video with audio, and text-to-image, in both Chinese and English