Alibaba's Qwen image model
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
Qwen Image
#38 of 62 in Text-to-Image
Wan 2.5 (Preview)
#27 of 62 in Text-to-Image
Where the votes landed
Qwen Image
0.0%
win rate
Ties
0.0%
Wan 2.5 (Preview)
100.0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
Qwen Image
- + Excellent adherence to the glass cube geometry.
- + Clean, minimalist composition that feels very photographic.
- + Accurate reflection of the blue sphere on the cube's base.
- − The plant is positioned more to the side than directly behind the cube as viewed through the glass.
Wan 2.5 (Preview)
- + Higher texture detail on the red book and wooden table.
- + Successfully places the plant directly behind the cube so it is visible through the glass.
- + Includes atmospheric dust particles in the soft window light.
- − The blue sphere appears to be floating slightly above the floor of the cube.
- − The perspective of the cube's top front edge is slightly warped under the book.
Verdict: Both models followed the complex spatial instructions well. Qwen Image produced a cleaner, more architecturally sound glass cube with realistic reflections, while Wan 2.5 (Preview) excelled at the 'partially visible through the glass' requirement and provided much richer surface textures and lighting details. Qwen Image is slightly preferred for its superior geometric coherence.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
Qwen Image
- + Excellent handling of motion blur on the background vehicles
- + Natural composition that feels like a genuine candid street photo
- + Accurate depiction of wet pavement and soft reflections
- − Anatomical issues with the man's hands merging into the bicycle seat
- − The bicycle framing is a bit awkward with the wheels appearing misaligned
Wan 2.5 (Preview)
- + Highly realistic skin textures and facial details on the man
- + Very intricate environment details including tools on the ground and rain droplets
- + Beautiful use of shallow depth of field and bokeh
- − Weak logic on the bicycle mechanics, such as the kickstand and floating chain elements
- − Missed the 'motion blur' request for the passing cars, which appear mostly static
Verdict: Both models followed the prompt well, but Wan 2.5 (Preview) produced a significantly more high-quality and realistic image with superior textures and lighting. While Qwen Image handled the 'motion blur' request better, it suffered from noticeable AI artifacts in the man's hands and the bike's structure, making Wan 2.5 the preferred choice for a professional cinematic look.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
Qwen Image
- + Excellent execution of engraving details on the plate armor.
- + Complex and colorful bead patterns in the hair braids.
- + Very dramatic and high-contrast lighting from the torch.
- − The facial scars look a bit like paint or superficial scratches rather than skin-level trauma.
- − Spark effects appear somewhat artificial and overly sharp.
Wan 2.5 (Preview)
- + Extremely realistic skin texture with lifelike eyes and subtle dirt/scar integration.
- + Superior rendering of frayed cloth and stitching on the underlayer.
- + The lighting on the face feels more natural and integrated into the scene.
- − The beads in the hair are monotonous in color compared to the prompt's potential.
- − A slightly younger-looking character that looks less 'battle-worn' than Model A's subject.
Verdict: Wan 2.5 (Preview) produces a more photographically realistic image with superior textures on skin and cloth, making the character feel more lifelike. Qwen Image has more intricate armor engravings and creative hair styling but suffers from slightly more digital-looking effects and less realistic skin rendering.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
Qwen Image
- + Strong layout that balances a high-quality vertical grid of images with a dedicated text column.
- + Clean, professional typography that uses bold sans-serif fonts effectively for headings.
- + Modern aesthetic that creates a distinct separation between menu items and prices.
- − The text content is mostly gibberish with minor spelling artifacts in the main header.
- − The food images are somewhat repetitive, featuring many similar salad bowls.
Wan 2.5 (Preview)
- + Excellent adherence to the grid layout with distinct sections for Appetizers, Pizza, and Mains.
- + High visual quality of food items with vibrant colors and appetizing presentation.
- + Included functional design elements like colored dividers to separate sections.
- − The font choice for some headers is a bit playful/stylized rather than the requested bold sans-serif.
- − The text is largely illegible across the entire menu.
Verdict: Qwen Image delivers a superior professional layout that feels like a real-world minimalist menu template, with great use of white space and a clean type-driven design. While Wan 2.5 (Preview) excels in the individual quality and variety of the food photography, its overall composition is slightly less 'modern minimalist' and more cluttered than Qwen Image.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
Qwen Image
- + Excellent text legibility and font layout.
- + Highly detailed food textures, particularly the grilled patty and sesame seed bun.
- + Effective use of a starburst element as requested by the prompt.
- − The burger is not fully 'exploded'; many components are still stacked together in the center.
- − The background fire feels a bit soft and lacks the sharp 'motion' requested.
Wan 2.5 (Preview)
- + Perfect interpretation of the 'exploded' request with every component clearly suspended in mid-air.
- + Dynamic sense of motion with flying droplets of sauce and varying angles for ingredients.
- + Creative fiery 'dripping' effect on the typography and a very energetic starburst.
- − The bottom bun texture looks slightly less realistic compared to the top bun.
- − Overall lighting is a bit more illustrative/digital than strictly photorealistic.
Verdict: Wan 2.5 (Preview) produced a much more dynamic and accurate interpretation of an 'exploded' burger, whereas Qwen Image kept the central ingredients mostly stacked. Wan 2.5 also followed the stylistic requests for fiery effects and starburst placement with more creativity, making it the better advertisement overall.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
Qwen Image
- + Follows the text request accurately including menu items and prices.
- + Clean, legible layout with consistent spacing.
- + Good chalkboard texture with realistic smudging in the background.
- − Text rendering is too clean and feels like a digital font rather than natural chalk handwriting.
- − Hallucinated the year as '20026' instead of '2026'.
- − Spelling error in 'Risotto' (spelled as 'Risoto').
Wan 2.5 (Preview)
- + Excellent chalk texture on the lettering with realistic pressure and grain.
- + Superior handwriting style that feels authentic and manually written.
- + Correctly rendered the year '2026' and all menu item names.
- − The second menu item is cut off ('& $28' instead of '& Herbs').
- − Repeat pricing on the third item creates a slightly cluttered look.
- − The background cafe scene is slightly more blurred/less defined than Image A.
Verdict: Wan 2.5 (Preview) produced a much more realistic chalk effect and handwriting style, which was a core requirement of the prompt, and correctly rendered the date. While Qwen Image followed the specific item list more closely, its text looks like a digital font overlay and it failed significantly on the requested year.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
Qwen Image
- + High resolution and clean rendering of the space suit and saddle details.
- + Great use of cinematic depth with the planet in the background.
- − Failed the specific prompt instruction to have the horse on top of the astronaut.
Wan 2.5 (Preview)
- + Dynamic composition with a sense of motion and vibrant nebula colors.
- + Excellent texture on the horse's coat and mane.
- − Failed the specific prompt instruction to have the horse on top of the astronaut.
- − Anatomical issues where the horse's front leg blends awkwardly with the astronaut's leg.
Verdict: Both models failed the specific spatial logic test in the prompt requesting the 'horse on top' of the astronaut, instead defaulting to the common trope of an astronaut riding a horse. Qwen Image produced a cleaner, more coherent image, whereas Wan 2.5 (Preview) had significant anatomy merging issues between the rider and the mount's front legs, despite its better color palette.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
Qwen Image
- + Natural composition and lighting.
- + The capybara's fur texture and expression are highly realistic.
- − The hands on the steering wheel look more like primate hands than capybara paws.
- − The taxi sign on top is stylized incorrectly as 'YOXI'.
Wan 2.5 (Preview)
- + Excellent depiction of a rainy Times Square lighting environment.
- + Better paw representation for the capybara.
- + The businesswoman's bored expression perfectly matches the prompt.
- − Perspective of the steering wheel and dashboard feels slightly flat.
- − The glass physics of the windshield and the taxi sign text are slightly garbled.
Verdict: Both models followed the prompt well, but Wan 2.5 (Preview) captured the specific atmosphere of a rainy New York night and the requested bored expression of the passenger more effectively. Qwen Image produced a cleaner, more photographic look for the capybara itself, but the strange primate-like hands were a significant anatomical error.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
Qwen Image
- + Excellent adherence to the 'dark parchment' and cinematic lighting prompt
- + Very clean layout with clear hierarchy of information
- + Successful rendering of the thorn and web border within a square format
- − Spelling error in the main title ('Halle Party Invitation')
- − The 'Halloween' text at the very top is garbled
Wan 2.5 (Preview)
- + Perfect text rendering for both the title and the secondary banner
- + Highly detailed jack-o-lantern with creative internal fire effects
- + Rich colors and intricate border details combining thorns and webs
- − Title text is slightly overlapping the border at the top
- − The layout feels a bit more cluttered compared to the atmospheric original prompt
Verdict: Both models followed the prompt closely, but Wan 2.5 (Preview) produced a superior image primarily due to its flawless text rendering, whereas Qwen Image had a significant spelling error in the word 'Halloween'. While Qwen Image captured the 'dark parchment' aesthetic more accurately, Wan 2.5's technical execution of the details and legibility makes it the more usable invitation.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
Qwen Image
- + Excellent adherence to the isometric perspective and miniature diorama base instructions.
- + Includes a variety of sushi types and props like chopsticks and ginger for a complete scene.
- + Text rendering is clean and bold as requested.
- − The flag icon in the text area is slightly distorted compared to a standard Japanese flag.
Wan 2.5 (Preview)
- + Features very realistic PBR materials with high-quality highlights on the fish texture.
- + Text rendering and flag icon are crisp and accurately formatted.
- + Clean, minimalist composition that feels high-end.
- − Missed the specific 'isometric' perspective, opting for a standard eye-level 3D view.
- − The diorama base is very simple and lacks the 'miniature scene' feel requested in the prompt.
Verdict: Qwen Image followed the complex layout instructions much better, delivering a true isometric miniature diorama with varied elements. While Wan 2.5 (Preview) had superior PBR material shaders and lighting, it failed to capture the requested isometric perspective and miniature scene depth.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
Qwen Image
- + Excellent soft lighting and realistic god rays.
- + Fur textures look natural and well-blended with the environment.
- + Composition feels harmonious with a nice depth of field.
- − The characters are mostly static rather than 'tumbling together' as requested.
- − The butterfly on the right is merging slightly with the fox's ear.
Wan 2.5 (Preview)
- + Successfully captures the 'tumbling' and 'chasing' motion described in the prompt.
- + Each of the four requested animals is distinct and clearly visible.
- + Vibrant colors and highly expressive facial features on the animals.
- − The fox's eyes have unnatural, glowing blue artifacts.
- − The floating dew drops look like glass orbs rather than natural moisture.
- − The kitten's tail is abnormally long and stiff.
Verdict: Qwen Image delivers a more aesthetically pleasing and realistic photographic style with superior lighting, though it lacks the dynamic motion requested. Wan 2.5 (Preview) better captures the energy of the prompt's 'tumbling' action but suffers from anatomical issues and unnatural 'sparkle' artifacts in the air. Qwen Image is the preferred choice for its higher visual coherence and professional finish.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
Qwen Image
- + Strong minimalist vector aesthetic
- + High contrast and bold design
- + Correct color palette implementation
- − Text layout for 'Florian' is crowded and messy
- − Banner design feels a bit heavy and flat
Wan 2.5 (Preview)
- + Superior typography that is clear and elegant
- + Refined illustrative details on the cloche and steam
- + Excellent textures and vintage paper border
- − Text on the cloche has slight lighting artifacts
- − Wait for generation could be higher if using the full model
Verdict: Wan 2.5 (Preview) is the clear winner as it produces a professional, legible logo with elegant typography and a beautiful vintage paper texture. Qwen Image adheres well to the minimalist style but fails significantly on the text rendering, resulting in an illegible 'Florian' and a disjointed layout.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
Qwen Image
- + Excellent adherence to the requested NASA-inspired color palette
- + Clean vector shapes with a friendly, modern design aesthetic
- − Poor text rendering with several misspellings like 'Sarth' and 'Apoll'
- − Included 'Stop at landing' from the prompt as actual title text
- − Failed to depict all 6 numbered steps in order
Wan 2.5 (Preview)
- + Superior text legibility and mostly accurate spelling for stage names
- + Better logical flow follow-through on the specific stages requested
- + More realistic icons for the Saturn V and Lunar Module
- − Included a strange rocket icon for Descent that looks like a shuttle
- − Mismatched portrait styles for the crew members
- − Missed the specified 'muted red' by using a brighter red
Verdict: Wan 2.5 (Preview) is the clear winner as it successfully organized the requested steps into a logical flow with legible, correctly spelled text. While Qwen Image captured the flat-vector style and specific color palette more accurately, its failure to provide readable text and its inclusion of prompt meta-instructions ('Stop at landing') as text make it less functional as an infographic.
Explore each model
Alibaba's text-to-image and image-to-image generation model from the Wan AI suite, offering high-quality visual generation capabilities