Improved version of Alibaba's Qwen image model with better text rendering, finer natural textures, and more realistic human generation.
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
Qwen Image 2512
#30 of 62 in Text-to-Image
Wan 2.6
#28 of 62 in Text-to-Image
Where the votes landed
Qwen Image 2512
33.3%
win rate
Ties
0.0%
Wan 2.6
66.7%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
Qwen Image 2512
- + Excellent realization of the glass cube with a realistic blue sphere and reflections.
- + Very clean, modern photographic aesthetic with soft, high-quality lighting.
- + Subtle and accurate rendering of the plant through the glass panes.
- − The plant is very close to the cube, making the 'behind' instruction a bit cramped.
- − The cube has a slight teal tint rather than being purely clear glass.
Wan 2.6
- + Perfectly follows the spatial instruction of the plant being behind the cube.
- + High level of detail on the book's texture and the wooden table's grain.
- + Wonderful interaction of light, specifically the caustic reflections on the table.
- − The book appears slightly small relative to the cube.
- − The perspective on the sphere's reflection at the base of the cube looks a bit distorted.
Verdict: Both models followed every instruction in the prompt perfectly. Qwen Image 2512 produces a cleaner, more professional-looking photograph with more realistic glass properties, while Wan 2.1 offers a more charming, rustic composition with superior spatial distancing between objects. Qwen Image 2512 is the likely winner for its superior visual clarity and the more natural integration of the sphere within the glass cube.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
Qwen Image 2512
- + Excellent high-resolution skin texture and facial details.
- + Accurate bike construction and seat design.
- + Beautiful bokeh and authentic automotive lighting.
- − The subject is posing for the camera rather than 'repairing' the bike.
- − Lacks the requested motion blur on the passing cars.
Wan 2.6
- + Perfect adherence to the 'repairing' action with tools visible.
- + Captures the 'candid' feel with the subject looking at his work.
- + Very realistic depiction of raindrops on the jacket and pavement reflections.
- − The facial anatomy is slightly less sharp than Model A.
- − Some minor structural confusion where the hands meet the bike chain.
Verdict: While Qwen Image 2512 produces a stunningly clear portrait, Wan 2.6 better interprets the specific narrative requirements of the prompt, showing the man actively repairing the bicycle with tools on the ground. Wan 2.6 feels more like a spontaneous street photo, whereas Qwen Image 2512 feels like a professional posed photograph.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
Qwen Image 2512
- + Exquisite detailing on the intricate engravings of the plate armor.
- + Superior skin texture and realistic facial scarring that feels integrated into the skin.
- + Excellent use of shallow depth of field and bokeh sparks to create atmosphere.
- − The leather straps look a bit flat compared to the realism of the armor.
Wan 2.6
- + Strong depiction of the 'battle-worn' aesthetic with heavy dirt and grime.
- + Great texture on the frayed cloth underlayer as requested in the prompt.
- + Dynamic lighting with two distinct light sources reflecting off the character.
- − The facial features look slightly smoothed/plasticky under the dirt.
- − The braid's beads look a bit like floating assets rather than being woven into the hair.
Verdict: Qwen Image 2512 produces a much more lifelike portrait with superior facial details and incredibly intricate armor engravings. While Wan 2.6 captures the 'dirt' and 'cloth underlayer' aspects of the prompt very well, its overall composition and facial rendering are less realistic than Qwen's output.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
Qwen Image 2512
- + Excellent photographic quality and consistency in the food grid
- + Very intentional use of bold, condensed sans-serif fonts
- + Clean, highly structured layout with clear vertical divisions
- − Text rendering is mostly gibberish despite looking stylistically correct
- − Lacks the 'vibrant accents' requested, sticking to a strict black and white theme
Wan 2.6
- + Successfully incorporates vibrant color accents as requested
- + Better text legibility overall
- + Clearly defined sections for Appetizers, Pizza, and Mains as per the prompt
- − The grid layout is slightly messy with uneven borders and clipping
- − Food photography is a bit less consistent in lighting and angle compared to Model A
Verdict: Wan 2.6 adhered better to the prompt's specific request for sections (Appetizers/Pizza/Mains) and vibrant accents, whereas Qwen Image 2512 combined sections into a 'Pizza/Means' header and used a purely monochrome UI. While Qwen produced more consistent food photography, Wan 2.6's design feels more like a complete, colorful casual dining menu.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
Qwen Image 2512
- + Excellent typography and graphic design layout.
- + Crisp, photorealistic textures on the burger ingredients.
- + Effective use of a starburst element for the price as requested.
- − The 'exploded' effect is more of a stack than a dynamic dispersal.
- − The price starburst looks a bit like a flat sticker rather than being integrated into the 3D scene.
Wan 2.6
- + Dynamic sense of motion with sauce drips and splashing embers.
- + True 'exploded' view with ingredients floating at varied angles.
- + Consistent fiery glow across all text elements.
- − The meat patty texture is slightly less realistic than Model A's.
- − Lacks the specific 'starburst' style requested for the price, looking more like a standard comic bubble.
Verdict: Qwen Image 2512 produces a cleaner, more professional advertisement with superior texture detail and graphic layout. However, Wan 2.6 better captures the 'exploded' and 'dynamic' motion aspects of the prompt, creating a more energetic composition even if the fine details are slightly softer.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
Qwen Image 2512
- + Excellent text legibility and spelling consistency
- + Accurately captures the elegant cursive request for the title while maintaining a consistent style throughout
- + Superior chalk texture and smudging realism on the board surface
- − Slightly mispelled 'Risotto' as 'Risitto'
- − The handwriting looks somewhat digitally clean compared to the rougher chalk request
Wan 2.6
- + Features a very authentic, crumbly chalk texture with visible dust
- + Successfully rendered all three menu items clearly
- + The 'handwritten' feel is more organic with natural grit
- − Redundant price rendering for the first two items ($24 and $28 are written twice)
- − The 'TODAY's Specials' title uses a mix of cases that wasn't specifically requested
- − The perspective of the board is slightly more distorted than Model A
Verdict: Qwen Image 2512 provides a much cleaner and more professional-looking menu with superior spelling accuracy, despite a minor typo in 'Risotto'. Wan 2.1 captures the gritty, dusty texture of chalk more realistically but suffers from distracting logic errors, such as repeating the prices twice for the same line item.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
Qwen Image 2512
- + Excellent photorealism in the textures of the horse's coat and the spacesuit material
- + High anatomical accuracy for both the person and the animal.
- + Captures a cinematic perspective with the Earth's curvature in the background.
- − Fails the specific prompt instruction to have the horse on top of the astronaut.
- − The lighting on the astronaut's face inside the helmet looks slightly flat compared to the suit.
Wan 2.6
- + Beautiful surreal atmosphere with vibrant nebulae and cosmic lighting.
- + Highly detailed rendering of the horse's mane and tail interacting with the environment.
- + Strong cinematic composition with a dynamic sense of motion.
- − Fails the specific prompt instruction to have the horse on top of the astronaut.
- − Minor anatomical issues with the horse's rear right leg.
Verdict: Both models failed the specific logical constraint of the prompt to place the horse on top of the astronaut, choosing instead to generate a standard astronaut-riding-horse image. Qwen Image 2512 provides a more grounded, photorealistic look with better human anatomy, while Wan 2.6 offers a much more vibrant and surreal artistic style that fits the 'cinematic' and 'surreal' keywords better.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
Qwen Image 2512
- + Excellent photorealism in the fur texture and clothing materials.
- + Captures the perfectly bored and mundane expression of the passenger.
- + Superior lighting and depth of field within the car interior.
- − The capybara's paws look somewhat human-like and uncanny.
- − The perspective from the dashboard makes it harder to see the 'taxi' exterior context.
Wan 2.6
- + Classic side-profile composition clearly shows the yellow taxi exterior and city lights.
- + The capybara's anatomy, specifically the paws and fur, looks more natural.
- + Great atmosphere with reflections on the taxi and rainy windows.
- − The passenger's scale and positioning in the back seat look slightly off.
- − The interior of the car looks significantly more weathered/deteriorated than a standard taxi.
Verdict: Qwen Image 2512 produces a more polished and high-fidelity image with a focus on textures and character expressions, particularly the bored passenger. However, Wan 2.6 provides a better composition for the prompt, showing the iconic yellow taxi exterior and a more believable capybara physiology, even if the overall image quality is slightly less crisp than the competitor.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
Qwen Image 2512
- + Strong cinematic lighting and glow effects.
- + Adheres closely to the requested border elements of thorns and webs.
- + Layout is well-balanced and reads clearly as an invitation.
- − Includes a spelling error in the main title ('Hallowern' instead of 'Halloween').
Wan 2.6
- + Perfect spelling in the main title.
- + High-quality texture on the trees and parchment background.
- + Beautiful gold-foiled effect on the gothic lettering.
- − The 'You are invited' banner is quite small and placed awkwardly high in the composition.
- − The pumpkin carving looks slightly more generic/less polished than the other model.
Verdict: Both models followed the prompt's structural requirements very well. Qwen Image 2512 has more impactful cinematic lighting and a better overall composition, but it fails on the basic spelling of the header. Wan 2.1 produces a higher quality texture and accurate text, making it the more functional invitation despite the slightly crowded upper half.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
Qwen Image 2512
- + Excellent text styling with a 3D-effect border
- + Rich detail in the sushi textures and vegetable garnishes
- + Better adherence to the 'small diorama' request with complex base layering
- − The flag icon is slightly off-center relative to the text
Wan 2.6
- + Very clean, minimalist aesthetic
- + Perfectly aligned typography
- + High-quality realistic PBR textures on the fish and wood
- − The sushi models are slightly floating or poorly integrated with the rice textures
- − The flag icon is placed to the left rather than 'below' the primary text as implied by the hierarchy
Verdict: Qwen Image 2512 better captures the 'miniature 3D cartoon scene' via more elaborate diorama details and stylized typography. Wan 2.6 is technically very clean with great light behavior, but it feels slightly more like a generic product render than a curated isometric diorama.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
Qwen Image 2512
- + Excellent fur texture and fine detail on the rabbit and fox facial features.
- + Strong adherence to the 'big expressive eyes' and 'god rays' portion of the prompt.
- + Very clear and vivid colors with sharp focus on all four animals.
- − Static composition where animals are posing rather than 'playfully chasing' or 'tumbling'.
- − The kitten is larger than the fox kit, creating an unnatural scale issue.
Wan 2.6
- + Dynamic composition that perfectly captures the 'playfully chasing and tumbling' action.
- + Beautiful atmospheric lighting with realistic dew sparkles and backlighting.
- + More natural scale and interaction between the species in a meadow setting.
- − The kitten has some minor anatomical blurring in its back leg/tail area.
- − The rabbit's eyes are less expressive/detailed compared to the others.
Verdict: While Qwen Image 2512 produces a very high-quality 'portrait' of the animals with incredible fur detail, Wan 2.6 captures the actual spirit of the prompt by showing the animals in motion, chasing butterflies and tumbling. Wan 2.6 feels much more like a consistent, lived-in scene with superior atmospheric lighting and dew effects.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
Qwen Image 2512
- + Perfect text rendering for both the name and the banner
- + High-quality illustrative detail with a classic emblem feel
- + Excellent use of the requested color palette and textures
- − Less 'minimalist' than requested, leaning more into detailed illustration
- − The steam is a bit oversized compared to the dome
Wan 2.6
- + Strictly follows the 'minimalist' and 'vector' keywords
- + Clean, simple composition suitable for a modern logo
- + Accurate representation of a cloche dome
- − The banner is very small and awkwardly placed
- − The background texture is limited to the edges
- − Lacks the 'vintage' richness requested
Verdict: Qwen Image 2512 produces a much more professional and aesthetically pleasing result that captures the 'vintage' and 'classic' requirements perfectly, even if it is less minimalist than requested. Wan 2.6 captures the minimalist vector style well, but the banner is poorly integrated and the overall design feels a bit generic.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
Qwen Image 2512
- + Successfully included all six requested infographic steps and their associated icons.
- + Followed the color palette and flat-vector infographic style perfectly.
- + Text rendering is reasonably legible and provides logical numbering for the mission phases.
- − Contains minor spelling errors in labels like 'Desceeint' and 'Translaurtoit'.
- − Includes redundant numbering for step 2 and 3.
Wan 2.6
- + Clean white circles with corectly spelled names of the crew members.
- + High contrast text that is easy to read.
- − Failed to include any of the six requested infographic steps or icons.
- − The composition is mostly empty space and does not function as an infographic poster.
- − The texture appears more like a fabric/towel than a vector poster.
Verdict: Qwen Image 2512 followed the detailed prompt instructions almost perfectly, creating a comprehensive infographic with all requested steps, whereas Wan 2.6 failed to generate most of the content. Qwen Image 2512's minor spelling issues are far outweighed by its superior layout and adherence to the vector art style requested.
Explore each model
Alibaba's multimodal generation model from the Wan AI suite, supporting text-to-video, image-to-video, reference-to-video with audio, and text-to-image, in both Chinese and English