Alibaba's Qwen image model
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
Qwen Image
#35 of 62 in Text-to-Image
Wan 2.7
#39 of 62 in Text-to-Image
Where the votes landed
Qwen Image
0%
win rate
Ties
0%
Wan 2.7
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
Qwen Image
- + Excellent clean aesthetic with high resolution.
- + Accurate positioning of the sphere and reflections.
- + Soft, realistic window lighting that matches the prompt.
- − The glass cube has internal vertical seams that make it look more like a case than a solid object.
Wan 2.7
- + Highly realistic textures on the book and the rustic wooden table.
- + Complex, realistic reflections on the glass and base.
- + Good adherence to the spatial arrangement of the plant behind the cube.
- − The plant appears to be growing inside the cube as much as it is behind it due to refraction issues.
- − The book is slightly too large for the cube's top surface.
Verdict: Both models followed the prompt near-perfectly regarding object placement and lighting. Qwen Image produced a cleaner, more modern look with a very smooth sphere, while Wan 2.1 excelled at capturing grit and texture in the old book and weathered wood, though it struggled slightly with the refraction logic of the plant behind the glass.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
Qwen Image
- + Excellent handling of shallow depth of field and bokeh
- + Beautiful reflections on the wet pavement
- + Strong sense of motion in the background car
- − Anatomical issues with the hands/fingers
- − Bicycle geometry is slightly nonsensical around the pedals and seat
Wan 2.7
- + Exceptional realism in skin and clothing texture
- + Perfectly captures the 'candid' street photography aesthetic
- + The bicycle design is much more coherent and believable
- − Less motion blur on the passing cars than requested
- − Depth of field is slightly deeper than a typical 50mm f/1.8 look
Verdict: Wan 2.7 produced a significantly more realistic and cinematic image that feels like a genuine candid photograph, particularly with the textures on the man's face and jacket. While Qwen Image handled the background motion blur and depth of field better, it suffered from AI artifacts in the man's hands and the bicycle structure.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
Qwen Image
- + Strong bokeh sparks effect
- + High contrast lighting with clear torchlight reflection
- + Very intricate engraving details on the shoulder plate
- − The beads in the hair look like modern colorful plastic
- − The scars appear more like fresh bloody cuts rather than 'faint scars'
- − Some lighting artifacts on the torch flame itself
Wan 2.7
- + Excellent skin texture with realistic dirt and faint scarring
- + Beads in the hair feel historically appropriate and well-integrated
- + Realistic cloth and leather textures underneath the armor
- − The engravings on the armor are less sharp and detailed than Model A
- − The torch in the background is slightly less defined
Verdict: Wan 2.1 is the winner for its superior realism and material accuracy, particularly in the hair beads and the subtle rendering of scars and dirt on the skin. While Qwen Image provides sharper armor engravings and dramatic lighting, its choice of colorful beads and fresh wounds detracts from the 'battle-worn paladin' aesthetic requested in the prompt.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
Qwen Image
- + Strong minimalist aesthetic with clean white space
- + Follows the bold sans-serif font requirement perfectly
- + Vibrant background colors in the photo grid create a modern look
- − Nonsense filler text under the headings
- − Less variety in food items featured in the photos
Wan 2.7
- + Excellent photo-realistic food images with high variety
- + High informational density including prices, descriptions, and social media handles
- + More complex layout including a QR code and realistic staging elements
- − Font choice is thinner and less 'bold' than requested
- − Layout feels slightly cluttered for a 'minimalist' prompt
Verdict: Qwen Image delivers a superior minimalist design that perfectly matches the requested bold typography and clean layout, though its text is mostly gibberish. Wan 2.7 provides a much more functional and realistic menu with coherent food items and prices, but it leans more toward a traditional casual dining style rather than the requested 'minimalist' aesthetic. Qwen Image is preferred for pure design adherence, while Wan 2.7 is better for a realistic use-case.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
Qwen Image
- + Excellent photorealistic texture on the meat patty and bun.
- + Clean, modern glowing typography for the 'MAGIC BURGER' title.
- + Effective use of depth of field with the fiery background.
- − The burger is not truly 'exploded' as most components are still stacked in the center.
- − The starburst for the price is very simplistic and lacks the fiery glow requested.
Wan 2.7
- + Perfect interpretation of the 'exploded' burger concept with clearly separated components.
- + Highly creative fiery font effect that blends well with the theme.
- + Dynamic splattering of sauces and seeds that enhances the sense of motion.
- − The meat patty looks a bit more like an illustration than a high-resolution photo.
- − The price starburst is a bit cluttered with rays and borders.
Verdict: Wan 2.7 followed the prompt much more accurately by providing a truly 'exploded' view of the burger where all layers are separated, whereas Qwen Image kept the burger mostly assembled. While Qwen Image has slightly superior material realism on the meat and bun, Wan 2.7 captures the dynamic motion and fiery typography requirements far more effectively.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
Qwen Image
- + The handwriting style feels very authentic and organic with natural slanting.
- + The chalk smudges and texture on the board add to the realism of a busy café.
- − There is a significant typo in the year, rendering it as '20026'.
- − The title is not in 'elegant cursive' as requested, but rather a blocky print.
- − It missed the word 'Risotto' on the first line, breaking it awkwardly.
Wan 2.7
- + Perfect text accuracy, including the specific date and all menu items.
- + The composition is clean and well-balanced within the frame.
- + The handwriting looks consistent and aesthetic, maintaining the same style throughout.
- − The text looks slightly like a digital font overlay rather than genuine chalk marks on a surface.
- − The title lacks the requested 'elegant cursive' style, opting for stylized print instead.
Verdict: Wan 2.7 is the clear winner for its superior text rendering, correctly spelling every item and the date exactly as requested, whereas Qwen Image failed significantly by writing the year as '20026'. While Qwen Image had a more realistic 'hand-drawn' feel with better chalk textures, Wan 2.7's reliability and complete adherence to the specific menu items make it the better overall output.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
Qwen Image
- + Excellent scale and lighting on the Earth background.
- + Anatomically coherent horse posture for a space environment.
- + Clean, cinematic aesthetic with smooth transitions.
- − The astronaut's left leg appears slightly stiff and poorly angled relative to the saddle.
- − The horse's hind legs have some minor clipping issues at the hooves.
Wan 2.7
- + Higher level of background detail with multiple galaxies and planets.
- + Better rendering of the astronaut's face visible through the visor.
- + Stronger texture rendering on the horse's coat and mane.
- − The horse's front left leg has an anatomically awkward joint and hoof angle.
- − The lighting on the astronaut is slightly inconsistent with the background light source.
Verdict: Both models successfully follow the prompt, but Qwen Image offers a more cohesive and cinematic composition with superior lighting integration. While Wan 2.7 provides more intricate background details and better facial rendering within the helmet, it suffers from minor anatomical issues in the horse's legs.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
Qwen Image
- + Excellent photorealistic lighting and depth of field.
- + Accurately places the passenger in the back seat as requested.
- + Very clean textures on the capybara's fur and the jacket.
- − The capybara's hands/paws look more like primate hands than capybara paws.
- − The taxi sign on top is misspelled as 'YOXI'.
Wan 2.7
- + Features a highly detailed capybara with realistic claws on the steering wheel.
- + The city background through the window feels very authentic to Manhattan.
- + Good adherence to the requested clothing and accessories.
- − The passenger is sitting in the front passenger seat instead of the back seat.
- − The passenger's hands and the phone have significant anatomy artifacts.
- − The seatbelt passes through the capybara's neck.
Verdict: Qwen Image followed the spatial instructions much better by placing the businesswoman in the back seat, whereas Wan 2.1 placed her in the front passenger seat. While Wan 2.1 had more realistic capybara anatomy (claws), Qwen Image produced a much cleaner overall composition with fewer anatomical artifacts on the human and better logical placement.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
Qwen Image
- + Strong cinematic lighting with atmospheric depth
- + Accurately represents the 'dark parchment' and gothic aesthetic
- + Unique thorn-style border correlates well with the prompt
- − Significant spelling error in the main title ('Halle Party')
Wan 2.7
- + Perfect text rendering for all requested fields
- + Ornate and highly detailed border including skulls and lanterns
- + Excellent layout and composition that feels like a completed graphic design
- − Lighting is more illustrative/flat rather than 'cinematic'
- − Border thorns look more like leafy vines
Verdict: While Qwen Image captures a more atmospheric and moody gothic vibe, it fails on the basic requirement of spelling the primary title correctly. Wan 2.7 provides a highly polished, professional invitation layout with perfect text accuracy and creative additions like the cauldron and ravens. Wan 2.7 is the preferred choice for a usable invitation design.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
Qwen Image
- + Excellent 3D miniature 'diorama' aesthetic with a thick base
- + Very soft, cohesive clay-like textures
- + Strong implementation of the requested isometric 45-degree angle
- − Text is slightly cut off at the top
- − The second flag icon is partially glitched into the text line
Wan 2.7
- + Perfectly rendered and centered text
- + Greater variety of sushi types and garnishes
- + Excellent material realism with subsurface scattering on the fish
- − The 'diorama base' feels more like a thick plate than a miniature scene
- − Missing the physical flag prop requested in favor of just a digital icon
Verdict: Qwen Image better captures the 'miniature diorama' feel requested in the prompt, using a thick layered base and a physical flag prop. However, Wan 2.7 provides superior text rendering, better material textures for the sushi, and a cleaner overall composition. Wan 2.7 is the likely winner for its professional polish and perfect adherence to the text layout requirements.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
Qwen Image
- + Strong composition with a clear central focus on the golden retriever.
- + Atmospheric lighting with soft, glistening bokeh and dewy effects.
- + Adorable character expressions that match the 'joyful vibe' well.
- − The fox's front legs look slightly unnatural in their pose.
- − The animals appear somewhat static rather than 'playfully chasing' or 'tumbling'.
Wan 2.7
- + Better dynamic movement, with the retriever and kitten actually appearing to run.
- + More varied butterfly designs and colors.
- + Highly detailed individual blades of grass and wildflowers.
- − The kitten has an anatomical error with a third paw resting on the bunny.
- − The lighting on the animals feels a bit flat compared to the bright background god rays.
- − The fox's face looks a bit mature and less like a 'kit' compared to the others.
Verdict: Both models successfully included all four requested animals and hit the atmospheric requirements. Qwen Image produces a more cohesive and aesthetically pleasing 'masterpiece' look with softer lighting, while Wan 2.7 captures the action of 'chasing' much better despite a significant anatomical glitch on the kitten.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
Qwen Image
- + Strong minimalist vector aesthetic
- + Correct spelling of 'Caffè'
- + Applies the requested subtle texture effectively
- − Confusing typography with overlapping, illegible letters
- − Composition is a bit bottom-heavy with the large banner
Wan 2.7
- + Excellent balanced emblem composition
- + Clean, high-quality vector line work
- + Clear and professional classic typography
- − Spelling error in the main name ('Florion' instead of 'Florian')
- − Steam icon is slightly disconnected from the cloche knob
Verdict: Wan 2.7 provides a much more professional and aesthetically pleasing emblem layout that perfectly captures the vintage restaurant vibe, though it suffers from a minor spelling error. Qwen Image adheres to the spelling correctly but fails on typography execution, creating a messy overlap of text that makes the brand name difficult to read.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
Qwen Image
- + Stronger vector art style with bold colors
- + Clear and charming icons for the rocket and lunar module
- − Poor text rendering with numerous typos and garbled words
- − Mistakenly included instructional text like '(Stop at landing)' inside the image
Wan 2.7
- + Excellent adherence to all 6 requested steps in sequence
- + Superior text rendering and information density
- + Professional infographic layout with a consistent NASA-inspired palette
- − Minor spelling errors in smaller text like 'Descript' and 'Tranquiliry'
- − The icons are much smaller and less impactful than the other model
Verdict: Wan 2.7 is the clear winner as it successfully followed the complex multi-step instructions and maintained a professional infographic layout. Qwen Image failed to include all steps and incorporated the prompt's negative constraints as actual text headers, resulting in a confusing and garbled final product.
Explore each model
Alibaba's Wan 2.7 image generation and editing model for text-to-image, reference-guided generation, and instruction-based image edits