Improved version of Alibaba's Qwen image model with better text rendering, finer natural textures, and more realistic human generation.
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
Qwen Image 2512
#30 of 62 in Text-to-Image
Wan 2.7
#39 of 62 in Text-to-Image
Where the votes landed
Qwen Image 2512
0%
win rate
Ties
0%
Wan 2.7
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
Qwen Image 2512
- + Excellent depiction of depth and glass physics
- + Highly realistic lighting and soft shadows
- + Very clean composition with cinematic focus
- − The glass box has some internal reflection inconsistencies
Wan 2.7
- + Perfect adherence to all prompt elements
- + Realistic texture on the wooden table and book cover
- + Clear visibility of the plant through the glass
- − The glass box construction shows some geometric errors particularly at the vertical corners
- − The reflection of the sphere on the left side of the glass appears solid rather than a reflection
Verdict: Both models followed the prompt instructions perfectly, including the specific spatial arrangement of the blue sphere, glass cube, red book, and green plant. Qwen Image 2512 is the winner because it handles the physics of light and glass much more convincingly, whereas Wan 2.7 has some confusing structural artifacts in the corners of the glass cube and less realistic reflections.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
Qwen Image 2512
- + Excellent shallow depth of field and bokeh effect
- + Remarkable skin texture and facial realism on the man
- + Clearly depicts the light rain and atmospheric reflections
- − The man is posing/looking at the camera, which is less 'candid' than requested
- − Lacks the requested motion blur on the passing cars
Wan 2.7
- + Natural, candid posture that follows the 'repairing' action more accurately
- + Good implementation of the requested camera angle and environment
- + Excellent wet pavement reflections
- − The bicycle appears to have three wheels or a distorted double frame
- − Does not capture the shallow depth of field or 50mm look as effectively as Model A
- − Cars in the background are too static and sharp, missing the motion blur requested
Verdict: Qwen Image 2512 produces a much more visually stunning and high-quality image with superior skin textures and cinematic lighting, but it misses the candid nature of the prompt by having the subject look at the camera. Wan 2.7 captures the 'candid' and 'repairing' aspects better, but it suffers from significant anatomical errors in the bicycle structure and lacks the 50mm lens depth requested.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
Qwen Image 2512
- + Superior texture on the engraved armor plates
- + Intense, lifelike gaze with realistic iris detail
- + Excellent implementation of warm torchlight and reflected glow on skin and metal
- − The leather straps look slightly flatter compared to the armor detail
Wan 2.7
- + Strong adherence to the braided hair and bead prompt
- + Good use of bokeh sparks in the background
- + Excellent leather strap and buckle texture
- − Skin texture appears slightly more digital/smooth than Model A
- − Lighting on the face is a bit flatter
- − The armor engraving is less intricate
Verdict: Qwen Image 2512 produces a superior cinematic portrait with more realistic skin textures and a highly sophisticated lighting setup that pulls the viewer in. While Wan 2.7 does a great job with the leather and hair details, it lacks the fine-grained 'battle-worn' realism and depth of field captured by Qwen Image 2512.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
Qwen Image 2512
- + Features a clear grid layout for images
- + Includes bold sans-serif header fonts as requested
- + Vibrant color-coded accents for different menu categories
- − Text is largely illegible/gibberish
- − The 'grid' format results in some cropped food items
- − Lower overall aesthetic polish compared to a real menu
Wan 2.7
- + Highly legible and realistic text rendering
- + Exceptional composition that mimics a professional modern menu
- + Includes realistic branding elements like a QR code and social media handles
- − The grid includes desserts and burgers not explicitly requested in the section list
- − Uses a serif font for some descriptions despite the 'bold sans-serif' prompt
Verdict: Wan 2.7 delivers a significantly more professional and usable result with nearly perfect text rendering and a sophisticated layout. While Qwen Image 2512 followed the layout requirements for a grid of photos, it failed to produce legible text and lacks the professional finish and secondary details provided by Wan 2.7.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
Qwen Image 2512
- + Excellent photorealistic texture on the meat patty and fresh produce
- + Strong sense of depth and atmospheric lighting
- + Clean and readable text integration with consistent fiery effects
- − The burger feels a bit too 'stacked' rather than fully 'exploded' compared to the other model
Wan 2.7
- + Dynamic composition with wide dispersion of ingredients matching the 'exploded' prompt
- + Creative use of smoke and sauce splashes to enhance the motion
- + Includes additional details like flying sesame seeds and pickles
- − The 'MAGIC BURGER' text has a stylized, slightly more illustrative look rather than high-end photorealism
- − The burger bun appears somewhat flat and less realistic than in Image A
Verdict: Qwen Image 2512 produces a highly polished, professional-looking advertisement with superior photorealistic textures and lighting. While Wan 2.7 captures the 'exploded' motion better through its composition and flying ingredients, Qwen's overall image quality and more realistic food rendering make it the more effective ad.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
Qwen Image 2512
- + Excellent chalk texture with realistic smudges and dust on the board.
- + Authentic handwritten variation in cursive style that looks truly manual.
- + Correct interpretation of a framed tabletop chalkboard in a cafe setting.
- − Spelling error in 'Risotto' (spelled 'Risitto').
Wan 2.7
- + Perfect spelling for all requested menu items.
- + Very clear and legible text alignment.
- + Good lighting and background environment detail.
- − The text looks more like a digital font or vector asset with a drop shadow rather than genuine chalk on a board.
- − The 'handwriting' is too uniform, lacking the natural variation requested in the prompt.
Verdict: Qwen Image 2512 produces a much more authentic 'handwritten' chalkboard look with realistic textures and manual flourishes, despite a minor spelling error. Wan 2.7 has perfect spelling but fails the stylistic requirement, as the text appears as a clean digital overlay rather than realistic chalk.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
Qwen Image 2512
- + High photographic realism in the lighting and textures of the spacesuit and horse.
- + Dynamic and cinematic composition with a clear view of the Earth's horizon.
- − Failed the negative constraint to have the horse on top of the astronaut.
- − Anatomical artifact where the horse's back leg appears to be missing or merged into the tail area.
Wan 2.7
- + Beautiful celestial background with detailed galaxies and planets.
- + Clean rendering of the astronaut's gear and horse anatomy.
- − Failed the specific instruction to place the horse on top of the astronaut.
- − Composition is a bit more generic and less cinematic than the competitor.
Verdict: Both Qwen Image 2512 and Wan 2.7 failed the primary challenge of the prompt, which was the 'horse on top' spatial arrangement, instead opting for the common 'astronaut riding a horse' trope. Qwen Image 2512 produces a more photorealistic and cinematic image, though it suffers from a significant anatomical glitch on the horse's hindquarters, whereas Wan 2.7 is cleaner but feels more like a standard digital illustration.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
Qwen Image 2512
- + Excellent adherence to the 'bored' facial expression on the human passenger.
- + Correct positioning of the human in the back seat as per the prompt.
- + Highly detailed fur texture and realistic reflection on the windshield.
- − The paws on the steering wheel look slightly anthropomorphized and claw-like.
- − The camera perspective through the windshield is slightly crowded.
Wan 2.7
- + Natural profile lighting on the capybara's face.
- + Good anatomy on the capybara's paws holding the wheel.
- + Clear view of the Manhattan street environment.
- − Failed to place the human in the back seat; she is in the front passenger seat.
- − The capybara's fur has a slightly more artificial, repetitive texture.
- − The passenger is looking at the phone but is sitting next to the driver, changing the requested dynamic.
Verdict: Qwen Image 2512 followed the prompt's spatial instructions much better, correctly placing the human passenger in the back seat with a distinct 'bored' expression. Wan 2.7 generated a high-quality image but failed on the layout by placing the passenger in the front seat, which misses the intended 'taxi ride' narrative.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
Qwen Image 2512
- + Atmospheric cinematic lighting that captures a moody night sky.
- + Highly accurate rendering of complex scroll banner text.
- + Coherent border design featuring thorns and webs exactly as requested.
- − Spelling error in the main title ('Hallowern' instead of 'Halloween').
Wan 2.7
- + Flawless spelling in the main title text.
- + Excellent composition with vintage aesthetic, including extra creative elements like skulls and books.
- + Clean and legible event details at the bottom of the poster.
- − The parchment is a bit bright and illustrative rather than 'dark parchment'.
- − Smaller text on the banner is slightly less crisp than the larger header.
Verdict: Both models followed the prompt instructions very well, capturing the gothic vintage aesthetic effectively. Wan 2.1 is the winner because it correctly spelled 'Halloween', whereas Qwen Image 2512 included a significant typo in the primary title, despite having slightly more atmospheric lighting.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
Qwen Image 2512
- + Excellent text rendering with clean outlines
- + Very charming 3D diorama base with organic details
- + Sophisticated lighting and color grading
- − The sushi variety is a bit limited compared to Model B
Wan 2.7
- + Features a wider variety of realistic-looking sushi pieces
- + Very clean, minimalist aesthetic aligned perfectly with the prompt
- + Accurate text and flag placement
- − The sushi variety includes a shrimp with a slightly odd, eye-like dot
- − The composition feels a bit more sparse compared to Model A
Verdict: Both models followed the prompt exceptionally well, producing clean 3D isometric designs with accurate text. Qwen Image 2512 has a slightly more cohesive 'toy' aesthetic with its layered diorama base, while Wan 2.1 offers a more varied and minimalist arrangement of sushi. Qwen Image 2512 wins slightly on art direction and the quality of the 'JAPAN' text styling.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
Qwen Image 2512
- + Excellent fur texture and facial detail on all animals
- + Strong atmospheric lighting with effective god rays
- + Coherent composition with animal interactions
- − The animals are static/composed rather than 'chasing and tumbling' as requested
- − The size scale of the rabbit and fox kit versus the kitten is slightly off
Wan 2.7
- + Successfully captures the action of 'chasing and tumbling' requested in the prompt
- + More realistic animal proportions relative to one another
- + Beautiful distribution of dew sparkles and wildflowers
- − The kitten has a strange facial expression and somewhat messy fur rendering
- − The fox's anatomy looks a bit stiff compared to the other animals
Verdict: Qwen Image 2512 produces a more polished, high-detail portrait with superior fur texture, but it fails to capture the 'chasing' action requested. Wan 2.7 better represents the playful movement and dynamic energy of the prompt, despite having slightly lower facial definition on the kitten. Wan 2.7 is the preferred choice for accurately translating the active scene description.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
Qwen Image 2512
- + Excellent typography with perfect spelling and an accent on the 'è'.
- + Beautiful vintage illustration style with high-quality cross-hatching and shading.
- + Strong composition where the cloche, steam, and banner feel integrated and professional.
- − The steam is a bit more illustrative and heavy than a typical minimalist logo might use.
Wan 2.7
- + Successfully creates a balanced circular emblem style.
- + Adheres well to the warm brown and cream color palette.
- + Includes all requested elements like the banner and cloche.
- − Significant typo in the main text, spelling it 'Florion' instead of 'Florian'.
- − The cloche dome is transparent, which looks more like a cake stand than a traditional metal cloche.
- − The steam is overly simplified and lacks the 'retro' feel of the rest of the piece.
Verdict: Qwen Image 2512 is the clear winner as it delivered flawless typography and a cohesive, professional vintage aesthetic. In contrast, Wan 2.7 failed at a basic level by misspelling the brand name as 'Florion' and produced a cloche that lacked the requested retro charm.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
Qwen Image 2512
- + Features more detailed and visually appealing illustrations of the Saturn V and Lunar Module.
- + Colors are well-balanced and strictly adhere to the NASA-inspired palette.
- − Poor text rendering with several misspellings and redundant labels.
- − The flow of information is disorganized with numbers jumping around and confusing placement.
Wan 2.7
- + Excellent typography with clear, mostly correct labels for each step.
- + Strong infographic composition with a logical vertical flow and consistent flat-vector iconography.
- + Excellent adherence to the 'flat-vector' style requested in the prompt.
- − The Saturn V icon is a bit generic compared to the detailed render in the other model.
- − The background stars are a bit cluttered and uniform.
Verdict: Wan 2.1 is the clear winner for its superior infographic layout and legible text, following the logical sequence of the Apollo mission perfectly. While Qwen Image 2512 has more attractive individual illustrations, its failure to organize the steps logically and its severe spelling errors make it a poor infographic.
Explore each model
Alibaba's Wan 2.7 image generation and editing model for text-to-image, reference-guided generation, and instruction-based image edits