Alibaba's Qwen image model
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
Qwen Image
#40 of 62 in Text-to-Image
Z-Image Turbo
#12 of 62 in Text-to-Image
Where the votes landed
Qwen Image
0%
win rate
Ties
0%
Z-Image Turbo
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
Qwen Image
- + Excellent adherence to the 'plant behind the cube visible through the glass' requirement
- + High realism in light refraction and table reflections
- + Balanced composition with good depth of field
- − Internal reflections in the glass cube create vertical lines that look like extra panes
Z-Image Turbo
- + Natural look to the red book's texture and wear
- + Clear sphere placement with clean reflections
- − Fails to show the green plant through the glass cube (pane is opaque/empty)
- − The book appears to slightly float above the glass surface
Verdict: Qwen Image followed the complex spatial instructions much better by accurately showing the green plant visible through the glass cube. Z-Image Turbo failed this specific prompt element, showing an empty/silver interior for the cube while the plant exists only in the background. Qwen Image also captured the soft window lighting more effectively across the wood grain.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
Qwen Image
- + Excellent cinematic lighting and atmosphere matches the prompt perfectly.
- + Realistic reflections on the wet pavement.
- + Accurate shallow depth of field and soft background blur.
- − Physical logic errors with the bicycle frame and kickstand area.
- − The subject is not explicitly 'repairing' the bike, more leaning over it.
Z-Image Turbo
- + Captures more realistic skin texture and detail on the man's arms and face.
- + Includes subtle rain streaks visible against the background.
- − Failed the motion blur request for passing cars.
- − Composition feels a bit flatter and less 'cinematic' than requested.
- − The car behind the bicycle is fused with the handles/man's arm.
Verdict: Qwen Image delivers a much more cinematic and atmospheric result that adheres better to the photographic style requested (50mm lens, motion blur, and wet pavement reflections). While Z-Image Turbo has slightly more detail in the skin texture, its failure to incorporate motion blur and the clipping/fusion artifacts with the background car make it a less successful interpretation of the prompt.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
Qwen Image
- + Excellent detail on the engraved metal scrollwork
- + Strong contrast and dramatic lighting from the torch
- + Clear, high-quality rendering of various textures like leather and chainmail
- − The beads in the hair look a bit like modern plastic jewelry
- − The fire and sparks have a slightly synthetic, illustrative feel
Z-Image Turbo
- + More realistic, subtle scarring and dirt application
- + Highly cinematic atmosphere and lighting
- + Realistic hair texture and beads that integrate naturally into the braids
- − The detail on the engravings is slightly softer than in the other image
- − The composition is a slightly wider than a 'close portrait'
Verdict: Both models followed the prompt exceptionally well. Qwen Image provides sharper, more defined textures on the armor and leather, but Z-Image Turbo achieves a more realistic and cinematic look, with the skin, scars, and hair beads appearing more natural and less like a digital illustration. Z-Image Turbo is the likely winner for its more cohesive and lifelike atmosphere.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
Qwen Image
- + Features a very clean, professional layout that matches the 'modern minimalist' prompt perfectly.
- + The grid system is well-organized with nice color-blocked accents.
- + Text rendering for titles is bold and stylishly integrated.
- − The body text for menu items is mostly garbled and illegible.
- − Some of the food photos in the grid lack variety, showing several very similar salads.
Z-Image Turbo
- + Includes a wider variety of realistic-looking food imagery including pasta, meat, and pizza.
- + Better legibility of the list pricing and menu item names compared to the other model.
- + Layout effectively uses orange accents to create a vibrant casual dining feel.
- − Includes typos in major headers like 'PIZZA MANS' and 'SE TIIION'.
- − The composition feels slightly more cluttered and less 'minimalist' than requested.
Verdict: Qwen Image delivers a superior graphic design aesthetic that perfectly captures the modern minimalist requirement, though the text content is less legible. Z-Image Turbo provides better food variety and clearer pricing structures, but is marred by significant spelling errors in the main headers. Qwen Image is the winner for its professional layout and adherence to the specific design style requested.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
Qwen Image
- + Excellent adherence to the 'exploded' concept with components clearly suspended in mid-air
- + Modern neon-style typography for the main title
- + High clarity and photorealistic textures on the floating ingredients like the tomato and lettuce
- − The starburst element for the price is a bit simple compared to the other graphics
- − The main burger core is still mostly assembled rather than fully exploded
Z-Image Turbo
- + Strong commercial aesthetic with vibrant glowing text effects
- + Very high quality rendering of the burger meat and melted cheese
- + Text integration feels professional and well-aligned
- − Fails to follow the 'exploded' instruction, showing a mostly assembled burger
- − Less sense of motion compared to the first image
Verdict: Qwen Image followed the core creative prompt much better by actually 'exploding' the burger components into the air, whereas Z-Image Turbo rendered a standard stacked burger. While Z-Image Turbo had slightly more professional-looking text integration, Qwen Image captured the requested dynamic motion and unique layout more effectively.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
Qwen Image
- + Excellent chalk texture on the board surface including realistic smudges
- + Very legible text with high contrast
- − Failed the date significantly by rendering '20026' instead of '2026'
- − Spelling error in 'Risotto' (spelled 'Risoto')
- − Handwriting looks slightly more digital/font-like rather than natural chalk strokes
Z-Image Turbo
- + Correctly rendered the date '2026'
- + Handwriting has a more authentic chalk-on-blackboard variation and texture
- + Completed the incomplete sentence from the prompt in a logical way
- − Spelling error in 'Mushroom' (spelled 'Mustroom')
- − Composition is a bit tight at the top and bottom margins
Verdict: Z-Image Turbo is the clear winner for following the specific date requested and capturing a more authentic handwritten chalk aesthetic. While both models had minor spelling errors, Qwen Image failed on the primary date requirement and used a typeface that felt less like real handwriting.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
Qwen Image
- + High cinematic quality with beautiful planetary backgrounds.
- + Clean rendering of the astronaut suit and horse's muscles.
- − Failed the spatial reasoning instruction; the astronaut is on the horse.
- − The horse's back right leg has an anatomical issue where the hoof is detached.
Z-Image Turbo
- + Realistic lighting on the horse's coat.
- + Good rendering of the astronaut's visor and equipment.
- − Failed the spatial reasoning instruction; the astronaut is on top.
- − The background is quite empty compared to the cinematic request.
Verdict: Both Qwen Image and Z-Image Turbo failed the negative constraint/spatial challenge to place the horse on top of the astronaut, instead providing the common 'astronaut riding a horse' trope. Qwen Image is slightly better due to its much more cinematic and detailed background, despite an anatomical error on one of the horse's legs.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
Qwen Image
- + Excellent photorealistic texture on the capybara fur and jacket.
- + Clear composition that shows both the driver and the passenger as requested.
- + The lighting and bokeh background perfectly capture the New York at night aesthetic.
- − The hands on the steering wheel look like monkey hands rather than capybara paws.
- − The taxi sign on top says "YOXI", which is a minor text hallucination.
Z-Image Turbo
- + Accurately depicts both paws on the steering wheel.
- + The capybara's expression is very calm and professional as requested.
- − The human passenger's hands and phone are blurred and poorly defined.
- − The composition feels tighter and slightly more claustrophobic compared to the first image.
- − The capybara's hands look more like human/primate hands than paws.
Verdict: Qwen Image is the superior model because it delivers a higher level of photorealistic detail and better lighting, creating a more convincing movie-still aesthetic. While both models struggled with rendering realistic capybara paws (giving the animal primate-like fingers instead), Qwen Image managed the human passenger and the background environmental details much more effectively than Z-Image Turbo.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
Qwen Image
- + Successfully included all elements including trees, bats, and jack-o-lantern.
- + The text rendering for the event details is crisp and accurate.
- + Creative framing with the thorn and webbing border matching the parchment style.
- − The main title font has a spelling error 'Halle Party'.
- − The upper calligraphy text is illegible and messy.
Z-Image Turbo
- + Excellent texture on the parchment and atmospheric lighting.
- + Perfectly spelled and high-quality gothic typography for the title.
- + Good integration of the scroll banners for subtext and details.
- − Spelling error in the location text ('Archves' instead of 'Arches').
- − The composition feels a bit cramped compared to the other model.
Verdict: Both models followed the prompt well, but Z-Image Turbo produced a more visually striking aesthetic with superior parchment textures and a much cleaner main title font. While Qwen Image handled the smaller event detail text slightly better, its misspelling of the main title 'Halle Party' makes it less effective as a functional invitation.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
Qwen Image
- + Perfect text rendering of 'JAPAN' and 'SUSHI'.
- + Accurate inclusion of the Japanese flag icon and a physical flag prop.
- + Superior diorama composition with multiple elements like chopsticks and maki.
- − The lighting is a bit flat across the text area.
Z-Image Turbo
- + Clean isometric diorama base.
- + High-quality specular highlights on the sushi fish.
- − Incorrect flag icon (displays the flag of China instead of Japan).
- − The text 'SUSHI' is slightly off-center and the font is less bold than requested.
- − Lacks the variety of sushi and accessories seen in the other model.
Verdict: Qwen Image significantly outperforms Z-Image Turbo by accurately adhering to the cultural context of the prompt, including the correct Japanese flag. Qwen Image also provides a more detailed miniature scene with chopsticks and varied sushi types, whereas Z-Image Turbo mistakenly uses a Chinese flag and a much simpler composition.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
Qwen Image
- + Expertly rendered god rays and soft morning lighting.
- + Superior texture on the animals' fur and the delicate dew sparkles.
- + Balanced composition with clear interaction between all subjects.
- − The fox's anatomy looks slightly stylized rather than purely photorealistic.
- − The puppy's expression is very calm, whereas the prompt requested 'playfully chasing'.
Z-Image Turbo
- + Captures the 'playful chasing and tumbling' action more dynamically.
- + Excellent expressions on the kitten and puppy that convey joy.
- + Accurately includes all four requested species with clear distinctions.
- − Noticeable anatomical errors, such as the puppy's paw having too many claws/toes.
- − Visible artifacts where the animals overlap, particularly the kitten's placement behind the rabbit.
- − The butterflies are less integrated into the lighting of the scene.
Verdict: Qwen Image delivers a more polished, high-resolution aesthetic with professional lighting and beautiful dew details, though it is a bit static. Z-Image Turbo captures the joyful energy and interaction of the prompt much better, but suffers from significant anatomical glitched and less refined fur textures. Qwen Image is preferred for its overall visual quality and lack of distracting artifacts.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
Qwen Image
- + Strong minimalist vector illustration style.
- + Includes the banner element requested in the prompt.
- + Atmospheric steam effect on the cloche.
- − Severely garbled typography with overlapping and illegible letters.
- − Poor use of space in the text area.
Z-Image Turbo
- + Perfectly legible and classic typography.
- + Clean, professional logo composition with high visual appeal.
- + Effective minimalist interpretation of the cloche icon.
- − Missing the 'banner' element specified for the 'Est. 1720' text.
- − Slightly less 'vintage texture' than requested.
Verdict: While Qwen Image follows the composition instructions regarding the banner, its text is completely illegible and visually messy. Z-Image Turbo produces a professional, clean logo with perfect spelling and high-quality typography, making it the superior choice for a usable logo design despite missing the banner element.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
Qwen Image
- + Followed the specific 6-step sequence requested in the prompt.
- + Consistent color palette that reflects the NASA-inspired navy and muted red requirements.
- + Cleaner layout with a logical visual flow connecting the mission phases.
- − Several spelling errors in the labels like 'Tranar' and 'Aldin'.
- − Included a strange seventh 'stop at landing' text next to the title which is redundant.
Z-Image Turbo
- + Stronger vector aesthetic with crisp lines and balanced negative space.
- + Iconography is professional and maintains a consistent flat-vector style.
- + Text is highly legible despite spelling errors.
- − Failed to include all 6 specific steps requested, skipping several icons.
- − Significant spelling errors in the main title ('Apolio E 11') and labels ('Translurian', 'Descenty').
- − The rocket design is less faithful to the Saturn V silhouette than Model A.
Verdict: Qwen Image followed the complex instructions much better, successfully visualizing all six requested mission steps in a logical sequence, whereas Z-Image Turbo missed several steps entirely. While both models struggled with text accuracy, Qwen Image's layout and adherence to the content requirements make it the more useful infographic.
Explore each model
Tongyi-MAI's 6-billion parameter distilled text-to-image model optimized for speed, achieving high-quality generation in 8 steps or fewer with support for bilingual text rendering