Head to head
Esc

Models · slot A

to navigate to pick

Qwen Image 2.0 Alibaba Wan 2.7 Alibaba

Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.

Qwen Image 2.0

21.7 arena score

#34 of 62 in Text-to-Image

Skill signature · Text-to-Image

Wan 2.7

20.4 arena score

#38 of 62 in Text-to-Image

Vote tally

Where the votes landed

Qwen Image 2.0

50.0%

win rate

Ties

25.0%

Wan 2.7

25.0%

win rate

50.0% 25.0% ties 25.0%
Shared challenges 13

Challenge by challenge

The strongest take from each model on every shared challenge, with the AI judge's read.

Geometric Composition

Text-to-Image

“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”

Qwen Image 2.0
Wan 2.7

AI Judge Analysis

Qwen Image 2.0

  • + Excellent photorealism in the wood grain and lighting.
  • + The sphere is centered and floating, creating a clean artistic effect.
  • + The glass cube edges are sharp and consistent.
  • The sphere appears to be floating unnaturally without support.
  • The reflection of the sphere on the side glass panels is a bit confusingly rendered.

Wan 2.7

  • + Natural physics with the sphere resting on the bottom of the cube.
  • + Highly detailed book texture with visible spine text and realistic wear.
  • + Excellent composition with a wider field of view showing more of the environment.
  • The bottom of the glass cube has some clipping/blending issues with the wooden table surface.
  • The glass reflection on the left side shows a duplicate sphere that doesn't align perfectly with the main sphere's position.

Verdict: Both models followed the complex spatial prompt accurately. Qwen Image 2.0 produced a cleaner, more minimalist image with a floating sphere, while Wan 2.7 opted for a more grounded, rustic aesthetic with superior texture work on the book. Wan 2.7 is slightly preferred for its realistic depiction of the sphere sitting on the base rather than hovering.

Candid Street Photography

Text-to-Image

“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”

Qwen Image 2.0
Wan 2.7

AI Judge Analysis

Qwen Image 2.0

  • + Excellent skin texture with realistic age spots and details
  • + Convincing shallow depth of field and motion blur on the background car
  • + Authentic framing that feels like a genuine candid street photo
  • The bike chain and pedal junction are physically impossible
  • The person standing in the top right background is partially cut off in an awkward way

Wan 2.7

  • + Better full-body composition and clear Japanese street environment
  • + Accurate depiction of wet pavement and light rain streaks
  • + Good source of the 'light rain' prompt aspect with visible droplets on the jacket
  • The red bicycle is missing several key parts like a seat and a functioning chain
  • Lacks the requested 'motion blur' for passing cars
  • The man's hands are anatomically blurry and poorly defined

Verdict: Qwen Image 2.0 captures the 'candid photo' aesthetic much better, particularly through the use of lens-accurate motion blur and highly detailed skin textures. While Wan 2.7 provides a wider view of the street, it fails to deliver on several technical prompt requirements like motion blur and shallow depth of field, and the bicycle is missing its seat.

Fantasy Warrior

Text-to-Image

“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”

Qwen Image 2.0
Wan 2.7

AI Judge Analysis

Qwen Image 2.0

  • + Excellent depiction of ornate engraving and filigree on the armor breastplate
  • + Highly lifelike eyes with dramatic, glowing reflections
  • + Superior color palette with vibrant contrast between the red cloth and silver metal
  • The character's hand resting on the hilt has some structural irregularities and blurry finger tips
  • The braids look a bit more like contemporary dreadlocks than traditional braids with beads

Wan 2.7

  • + Perfect execution of the hair braiding with small beads as requested
  • + Very realistic skin texture with believable scars and fine dirt
  • + Higher overall technical clarity and cleaner details on the leather straps
  • The armor engravings are slightly more repetitive and less 'ornate' than Model A
  • Lighting is a bit flatter and less cinematic than the warm glow in the first image

Verdict: Both models followed the prompt exceptionally well, but Wan 2.7 produced a cleaner, more technically sound image with better adherence to the specific hair braiding and bead details. Qwen Image 2.0 offered a more cinematic feel and impressive engravings, but was let down by anatomical issues in the hand and a less accurate interpretation of the hair style.

Modern Clean Menu

Text-to-Image

“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”

Qwen Image 2.0
Wan 2.7

AI Judge Analysis

Qwen Image 2.0

  • + Features a very clean, minimalist grid that is aesthetically pleasing
  • + Photographs of the food are high-quality, vibrant, and appetizing
  • + Clearly uses bold sans-serif fonts for section headers as requested
  • Text rendering for dish names is garbled and unreadable
  • Food items do not match the section headers (e.g., pizzas are shown under 'Appetizers' and 'Mains')

Wan 2.7

  • + Excellent layout that resembles a real, functional restaurant menu
  • + Remarkably clear and mostly accurate text, including dish names and descriptions
  • + Successful categorization of food items into relevant sections
  • The 'Bread & Bowl' branding was not specifically requested
  • The composition includes background props (pen, rosemary) rather than being just the graphic design file

Verdict: Wan 2.7 is the clear winner as it produces a professional, functional menu with legible text and logical categorization of food items. While Qwen Image 2.0 has high-quality food photography, its inability to render readable text and the misalignment between headers and images makes it unusable as a design layout.

Magic Burger Explosion: Fiery Photorealism Challenge

Text-to-Image

“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”

Qwen Image 2.0
Wan 2.7

AI Judge Analysis

Qwen Image 2.0

  • + Excellent photorealistic texture on the meat and bun
  • + Very effective fiery text effect for the main title
  • + Clean, high-quality rendering of the starburst price tag
  • The burger is only partially exploded, with most components still stacked
  • The secondary text is relatively small and lacks the 'fiery' effect requested

Wan 2.7

  • + Successfully achieves the 'exploded' look with all components widely separated
  • + All text elements perfectly follow the prompt instructions regarding placement and style
  • + Dynamic use of splashes and floating seeds enhances the sense of motion
  • Lighting on the lettuce looks somewhat artificial and flat
  • The bun texture is less realistic than in model A

Verdict: While Qwen Image 2.0 has superior photorealism in its textures, Wan 2.7 is the winner for its better adherence to the 'exploded' composition and accurate rendering of all requested text elements. Wan 2.7 effectively captures the dynamic motion of the individual ingredients and better integrates the glowing fiery theme across all text components.

Chalkboard Menu

Text-to-Image

“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”

Qwen Image 2.0
Wan 2.7

AI Judge Analysis

Qwen Image 2.0

  • + Excellent adherence to the 'handwritten chalk' requirement with realistic texture and smudges
  • + Perfect spelling and logical line breaks
  • + Contains the most realistic chalk lighting and board texture
  • The slant and size vary significantly, which might reduce legibility for some
  • Composition is a bit tighter on the left side

Wan 2.7

  • + Perfect alignment and centered composition
  • + Clear and highly legible text
  • + Successfully completed the truncated prompt for the third menu item
  • The text looks like a digital font or vector graphic rather than chalk handwriting
  • Missing the requested 'elegant cursive' for the title
  • The text lacks the requested chalk texture and grit, appearing too smooth and plastic-like

Verdict: Qwen Image 2.0 followed the prompt's stylistic instructions much better than Wan 2.1, providing a truly realistic chalk-on-blackboard aesthetic with natural smudges and variations. Wan 2.1 produced very clean and centered text, but it failed the negative constraint by using a style that looks like a digital or printed font rather than realistic handwriting.

The Reversed Rodeo

Text-to-Image

“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”

Qwen Image 2.0
Wan 2.7
50% wins 25% ties 25% wins

AI Judge Analysis

Qwen Image 2.0

  • + Excellent adherence to the 'surreal' instruction with scale-like horse skin and floating droplets.
  • + High cinematic quality with dynamic lighting and a focused composition.
  • + The astronaut's suit has complex metallic and textile details.
  • The horse's facial structure is slightly distorted around the eyes and bridge.
  • The front left leg has an anatomically awkward connection to the shoulder.

Wan 2.7

  • + Naturalistic and realistic horse anatomy and fur texture.
  • + Clear, sharp rendering of the astronaut suit and equipment.
  • + Good depth of field including background galaxies and planets.
  • Less 'surreal' than the prompt requested, feeling more like a standard composite.
  • The horse's rear right leg disappears/blends poorly into the horse's body.

Verdict: Both models failed to follow the specific 'horse on top' instruction, which likely required a reverse-centaur or a horse physically riding on the back of an astronaut. Between the two standard interpretations, Qwen Image 2.0 is preferred for more effectively capturing the 'surreal' and 'cinematic' style requested, whereas Wan 2.7 produced a more conventional but anatomically cleaner horse.

The Capybara Taxi Driver

Text-to-Image

“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”

Qwen Image 2.0
Wan 2.7

AI Judge Analysis

Qwen Image 2.0

  • + Excellent photorealistic texture on the capybara's fur and the paws
  • + Great dynamic lighting reflecting off the taxi window and interior
  • + The composition feels intimate and follows the 'inside the taxi' instruction well
  • The passenger's hands and phone are slightly blurred and look a bit messy
  • The perspective makes it look like the passenger is in the front seat next to the driver

Wan 2.7

  • + Clearly places the passenger in the back seat as requested
  • + Clean and sharp rendering of both subjects
  • + Good execution of the 'bored expression' prompt for the businesswoman
  • The capybara's fur looks somewhat simplified and less realistic compared to Model A
  • The capybara's paws are not correctly positioned on the steering wheel, appearing to float behind or beside it

Verdict: Qwen Image 2.0 captures a much more realistic and atmospheric scene with superior textures and lighting, though it struggles with the spatial placement of the passenger. Wan 2.7 follows the 'back seat' instruction better, but lacks the photorealistic depth and has noticeable errors with the capybara's paws and steering wheel interaction.

The Halloween Invitation

Text-to-Image

“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”

Qwen Image 2.0
Wan 2.7

AI Judge Analysis

Qwen Image 2.0

  • + Excellent atmospheric lighting and cinematic depth
  • + Perfect adherence to the 'dark parchment' and gothic aesthetic
  • + Clean and legible typography that fits the gothic theme
  • The twisted trees and background elements are a bit blurry
  • Composition is slightly cluttered around the center text

Wan 2.7

  • + Crisp and clear illustration style with great detail
  • + Excellent text rendering for both titles and tiny details
  • + Well-organized layout that uses the border effectively
  • The style is more 'cartoon/storybook' than the requested 'cinematic' look
  • The background sky is relatively flat and less moody than Image A

Verdict: Qwen Image 2.0 followed the stylistic cues for a 'cinematic' and 'moody' gothic invitation more effectively, producing a cohesive atmospheric piece. Wan 2.7 produced a cleaner, more readable design with superior text rendering and extra details like skulls and books, but it feels like a flat illustration rather than a vintage poster.

Isometric Miniature Diorama Scenes

Text-to-Image

“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”

Qwen Image 2.0
Wan 2.7

AI Judge Analysis

Qwen Image 2.0

  • + Excellent photographic realism in texturing
  • + Clear and accurate typography
  • + Vibrant colors that pop against the background
  • Missed the 'cartoon' and 'isometric' style requested, opting for realism
  • Perspective is a standard photo angle rather than a technical 45° isometric view

Wan 2.7

  • + Perfectly captures the requested 3D cartoon isometric diorama style
  • + Accurate 45° perspective following the diorama base prompt
  • + Clean, soft textures that match the 'miniature' aesthetic
  • The flag icon is slightly merged into the text box
  • Top of the nigiri rice clusters looks a bit uniform and repetitive

Verdict: While Qwen Image 2.0 produced a high-quality realistic image, it failed to follow the stylistic instructions for an isometric cartoon diorama. Wan 2.1 followed every keyword of the prompt, including the specific perspective, the stylized 3D miniature look, and the layout of the diorama base.

Adorable Baby Animals in Sunny Meadow

Text-to-Image

“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”

Qwen Image 2.0
Wan 2.7

AI Judge Analysis

Qwen Image 2.0

  • + Excellent dynamic interaction showing all four animals tumbling together as requested.
  • + Effective use of god rays and warm sunrise lighting to create a wholesome mood.
  • + High level of detail on fur textures and realistic anatomy.

Wan 2.7

  • + Beautiful backlighting and bokeh effect in the meadow.
  • + Very clear and distinct rendering of each animal's facial features.
  • + Captures the 'dew sparkles' from the prompt effectively through light glints.
  • The animals are standing side-by-side rather than 'tumbling together' as the prompt requested.
  • The kitten's tail appears to be missing or poorly positioned behind its body.
  • The interaction feels static and posed compared to the dynamic energy of Model A.

Verdict: Qwen Image 2.0 is the clear winner as it successfully follows the complex instruction for the animals to be 'tumbling together,' creating an adorable and chaotic pile-on that matches the 'joyful vibe'. Wan 2.7 provides a high-quality image, but the animals are simply walking in a line, missing the primary action described in the prompt.

Vintage Cafe Logo

Text-to-Image

“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”

Qwen Image 2.0
Wan 2.7

AI Judge Analysis

Qwen Image 2.0

  • + Excellent typography with correct spelling and accents
  • + Creative integration of the cloche and steam elements
  • + Clean vector aesthetic and pleasant warm color palette
  • The 'E' in 'Est.' on the banner is slightly deformed
  • The perspective of the steam inside/on the cloche is slightly confusing

Wan 2.7

  • + Classic balanced emblem composition typical of vintage logos
  • + Well-executed cloche illustration and banner integration
  • + Professional border and framing details
  • Spelling error in the main text ('Florion' instead of 'Florian')
  • The steam icon looks more like a musical clef or noodle than steam clouds

Verdict: Qwen Image 2.0 followed the text requirements perfectly, correctly rendering the name 'Caffè Florian' with the appropriate accent. While Wan 2.1 produced a very professional-looking circular emblem, it failed on a fundamental requirement by misspelling the primary brand name.

Apollo 11: Journey to Tranquility

Text-to-Image

“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”

Qwen Image 2.0
Wan 2.7

AI Judge Analysis

Qwen Image 2.0

  • + Excellent adherence to the requested NASA-inspired color palette.
  • + High-quality, detailed illustrations for the lunar module stages.
  • + Clear and legible typography for the main labels.
  • Contained a spelling error in 'Translunjar'.
  • The layout feels a bit crowded and lacks some of the 'pro' infographic finish.

Wan 2.7

  • + Perfect professional infographic layout with structured alignment and supporting data points.
  • + Consistent, minimalist iconography that adheres well to the flat-vector style.
  • + Includes extra details like the astronaut names and mission dates in a clean footer.
  • Multiple spelling errors in small text (e.g., 'DESCRIPT', 'Tranquiliry').
  • The landing icon is significantly less detailed than the one in Model A.

Verdict: Both models followed the complex prompt instructions well, but Wan 2.7 produced a superior composition that feels more like a professional infographic poster. While Qwen Image 2.0 has better individual illustrations and fewer major typos, Wan 2.7's use of balance, consistent iconography, and supporting data makes it more visually effective for the requested format.

Next steps

Explore each model