Alibaba's Qwen Image 2.0 model with enhanced text rendering, supporting both Chinese and English prompts with up to 6 images per request
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
Qwen Image 2.0
#34 of 62 in Text-to-Image
Z-Image Turbo
#12 of 62 in Text-to-Image
Where the votes landed
Qwen Image 2.0
0.0%
win rate
Ties
0.0%
Z-Image Turbo
100.0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent photographic texture on the book and wooden table
- + High level of detail in the glass reflections and refractivity
- + Realistic lighting interaction with the surrounding environment
- − The sphere is floating unnaturally in the center of the cube
- − The sphere appears duplicated or reflected incorrectly on the side panels of the glass
Z-Image Turbo
- + Natural placement of the sphere resting on the bottom of the cube
- + Follows all spatial instructions accurately
- + Clean composition with a realistic depth of field
- − Slightly lower textural detail on the red book compared to its competitor
Verdict: Both models followed the prompt requirements perfectly, including the complex request for a plant to be visible through the glass cube. Qwen Image 2.0 has superior texture and lighting, but Z-Image Turbo is preferred for its more logical physical arrangement, as the sphere sits on the base of the cube rather than floating in mid-air.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent skin texture and hyper-realistic facial details
- + Accurate shallow depth of field and beautiful reflections on the wet pavement
- + Strong storytelling with the man actively repairing the chain mechanism
- − The bike chain and pedal geometry is slightly nonsensical and tangled
- − The hand interacting with the pedal looks somewhat distorted
Z-Image Turbo
- + Successfully captures the light rain and wet road texture
- + Good overall composition and clear depiction of the red bicycle
- − The man's hands are very poorly rendered and fused with the handlebars
- − Lacks the 'candid' feel and looks more like he is just standing with the bike
- − Skin textures look smoother and less natural compared to the other model
Verdict: Qwen Image 2.0 captures the 'candid' and 'cinematic' requirements significantly better than Z-Image Turbo, providing a photo-real quality with convincing skin textures and complex reflections. While both models struggle slightly with the mechanical precision of the bicycle and hand-to-object interaction, Qwen Image 2.0 maintains a much higher level of visual fidelity and mood.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent depiction of battle-worn skin texture and grit
- + High-contrast lighting with vibrant bokeh sparks
- + Ornate engraving details on the chest plate are very clear
- − The eyes appear slightly unnatural and glowing
- − The hands and fingers have anatomical inaccuracies
Z-Image Turbo
- + Natural and lifelike eye rendering
- + Sophisticated jewelry and beadwork in the hair braids
- + Superior rendering of cloth underlayers and chainmail
- − The torch flame looks a bit like a digital overlay
- − The 'worn' aspect of the plating is more subtle than Model A
Verdict: Qwen Image 2.0 excels at capturing the grit and intense lighting of a battle-worn warrior, but struggles with the anatomy of the hands. Z-Image Turbo provides a more cohesive and technically cleaner image, with realistic eyes and a better representation of the layered textures like chainmail and fine leather. Z-Image Turbo is preferred for its superior fidelity and natural composition.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent visual quality of food photography with high lighting and texture detail.
- + Perfectly balanced grid layout that aligns with the requested minimalist design.
- + Clear headlining for all three requested sections: Appetizers, Pizza, and Mains.
- − Text rendering below images is gibberish.
- − The grid proportions are slightly inconsistent across the three rows.
Z-Image Turbo
- + Includes realistic pricing lists which enhance the 'menu' feel.
- + Adds vibrant orange accents as requested in the prompt.
- + Professional shadow effect making it look like a physical menu card.
- − Failed to create a section for 'Mains', merging it with 'Pizza' instead.
- − Food images are less distinct and higher in artifacts compared to Model A.
- − Typography layout is cramped and messy in the bottom right corner.
Verdict: Qwen Image 2.0 is the clear winner for its superior image quality and perfect adherence to the structural requirements of the prompt, including the specific sections requested. While Z-Image Turbo attempted to include prices and color accents more effectively, it failed on the logical organization of the menu categories and produced lower-quality food photography.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent adherence to the 'exploded' burger requirement with clearly suspended individual layers.
- + Highly effective use of fire and ember particle effects that blend with the product.
- + Text rendering is clean and follows the fiery effect prompt perfectly.
- − The 'LIMITED TIME ONLY' text is slightly small and less integrated than the main header.
Z-Image Turbo
- + Clean, readable typography for the primary and secondary messages.
- + Strong glowing aura around the price starburst.
- + Solid photorealistic textures on the beef patties.
- − Failed to create an 'exploded' burger, showing a mostly assembled stack instead.
- − The background is more of a warm blur rather than a dark, fiery scene with distinct embers.
- − The starburst is placed in a way that feels a bit generic compared to the first image.
Verdict: Qwen Image 2.0 is the clear winner as it fully captured the 'exploded' nature of the prompt, whereas Z-Image Turbo presented a standard stacked burger. Qwen Image 2.0 also demonstrated superior creative integration of the fiery theme, using embers and smoke to bridge the gap between the food and the background text.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent chalk texture with realistic smudges and strokes.
- + Accurate spelling on all menu items.
- + Beautifully rendered cafe environment in the background adds to the 'cozy' atmosphere.
- − The cursive title requested is partially printed/block-style rather than 'elegant cursive'.
- − The slant in the handwriting is a bit inconsistent across words.
Z-Image Turbo
- + Text is extremely legible and centered well within the frame.
- + Captures a very consistent handwriting style throughout the entire board.
- − Includes a spelling error ('Mustroom' instead of 'Mushroom').
- − The chalk texture is very clean, almost looking like a digital font rather than natural chalk.
- − Completely ignores the 'elegant cursive' requirement for the title.
Verdict: Qwen Image 2.0 is the superior result because it follows the spelling requirements perfectly and provides a much more convincing chalk texture with realistic environmental details. While Z-Image Turbo has very neat alignment, the spelling error and lack of requested cursive style make it less successful for this specific prompt.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent adherence to the 'surreal' aesthetic with shimmering, scale-like horse hair and floating water droplets.
- + Dynamic composition with a beautiful cinematic background featuring Earth and a rich starfield.
- + Highly detailed textures on both the spacesuit and the horse's musculature.
- − The horse's front legs have a slightly unnatural anatomical structure even for a surreal concept.
Z-Image Turbo
- + Crisp, clean rendering of the astronaut and the saddle equipment.
- + Clear interpretation of the 'horse on top' (astronaut riding horse) instruction.
- − The background is very plain and lacks the cinematic, space-like quality requested.
- − The lighting on the horse is flat and feels like a standard composite rather than a surreal space scene.
- − The astronaut's glove and reins have some minor clipping issues.
Verdict: Qwen Image 2.0 followed the 'surreal' and 'cinematic' prompts much more effectively, creating an ethereal scene with scales, floating droplets, and a vibrant galactic background. Z-Image Turbo produced a much more literal and plain image that lacks the requested level of detail and atmosphere. While both followed the instruction of putting the horse on the bottom (astronaut on top), Qwen Image 2.0 is the superior visual work.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent handling of complex anatomy with the capybara's paws on the steering wheel.
- + High level of texture detail in the capybara's fur and the passenger's coat.
- + Strong night-time lighting and city reflections on the window glass.
- − The passenger appears to be in the front passenger seat rather than the back seat as requested.
Z-Image Turbo
- + Successfully placed the passenger in the back seat as specified in the prompt.
- + Clean, professional-looking taxi driver cap with a badge.
- − Human passenger's hand holding the phone looks distorted.
- − The capybara's arm appears elongated and anatomically awkward compared to the body.
- − Lighting is a bit flat for a night-time New York street scene.
Verdict: Qwen Image 2.0 produced a much more realistic and detailed image with superior texture and lighting, though it failed the spatial positioning of the passenger. Z-Image Turbo followed the layout instructions better by placing the woman in the back, but suffered from significant anatomical defects in the hands and arms. Qwen Image 2.0 is the preferred choice for its photographic quality and convincing execution of the primary subject.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent typography with a cohesive gothic font style across all text.
- + High-quality, cinematic rendering of the jack-o-lantern and misty forest.
- + Perfect spelling of all requested event details.
Z-Image Turbo
- + Creative use of layering with the torn parchment on a dark background.
- + Good use of multiple scrolls for different text sections.
- + Includes graveyards and additional spooky elements for atmosphere.
- − Typos in the text: 'The Archves' instead of 'The Arches' and missing letters in the top banner.
- − The text font is inconsistent between the title and the details.
- − Composition is a bit cluttered with overlapping thorn and web elements.
Verdict: Qwen Image 2.0 followed all instructions perfectly, delivering an elegant and legible design with no spelling errors. While Z-Image Turbo had a creative parchment layered effect, it failed on fine details like spelling ('Archves') and font consistency, making Qwen Image 2.0 the superior choice for a usable invitation.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent text rendering and placement.
- + Authentic and correct Japanese flag icon.
- + Higher complexity in the sushi variety while maintaining a clean aesthetic.
- − Lean more toward a realistic photograph than the requested '3D cartoon' style.
- − The wooden board is a standard tray rather than a stylized diorama base.
Z-Image Turbo
- + Perfectly captures the '3D cartoon' miniature aesthetic with soft textures.
- + Strict adherence to the 45-degree isometric projection and diorama base.
- + Clean, rounded shapes that match the 'miniature' prompt.
- − Includes the flag of China instead of Japan.
- − Text is slightly off-center and the flag placement is awkward.
- − The sushi piece has a strange cross-section/internal element.
Verdict: Qwen Image 2.0 produced a much more accurate image regarding the specific subject matter, correctly depicting the Japanese flag and high-quality text. However, Z-Image Turbo better followed the stylistic instruction for a '3D cartoon' isometric miniature, though it failed the critical detail of the national flag. Qwen Image 2.0 is the preferred winner because it successfully integrated the text components and correct iconography without errors.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent fur textures that meet the 'ultra-detailed' request
- + Complex interaction with animals actually tumbling together
- + Superior background detail with diverse wildflowers and convincing god rays
- − The fox kit has a slightly distorted facial expression while lying on its back
Z-Image Turbo
- + Cute, expressive facial features on all four animals
- + Clear and clean composition
- − The kitten and fox look more like static figurines than playing animals
- − The 'tumbling together' part of the prompt is not represented
- − Butterflies and lighting effects look more artificial and less integrated into the scene
Verdict: Qwen Image 2.0 followed the complex action requirements of the prompt much better, showing the animals actually tumbling and interacting in a dense, detailed meadow. Z-Image Turbo produced a cute static group shot, but it lacked the realistic fur textures and dynamic energy requested in the prompt. Qwen Image 2.0's superior lighting and background detail make it a much more successful 8K masterpiece.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
Qwen Image 2.0
- + Accurate rendering of the accent mark in 'Caffè'
- + Creative integration of the steam inside the cloche
- + Successfully includes the requested banner element
- − The ribbon/banner geometry is slightly inconsistent on the right side
- − The '1720' text is slightly misaligned within the banner
Z-Image Turbo
- + Stronger minimalist vector aesthetic
- + Very clean typography and professional layout
- + Better adherence to the 'minimalist' instruction
- − Failed to include the requested banner element, using simple lines instead
- − Steam detail is very tiny and lacks impact
Verdict: Both models followed the prompt well, but Qwen Image 2.0 captures more specific details like the banner and the specific 'Caffè' accent mark. While Z-Image Turbo is more 'minimalist', it failed to include the requested banner element, making Qwen Image 2.0 the more accurate interpretation of the prompt's specific requirements.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent adherence to the infographic structure, following all six steps in order.
- + Highly legible text with relatively accurate spelling (one minor typo in 'Translunjar').
- + Sophisticated vertical composition that tells a clear narrative story.
- − Includes a minor spelling error ('Translunjar' instead of Translunar).
- − The icons for descent and landing are a bit more detailed rather than 'flat-vector' style.
Z-Image Turbo
- + Strong flat-vector aesthetic with bold, clean icons.
- + High-quality color palette that matches the NASA-inspired theme.
- − Fails to follow the requested six-step sequence, grouping text haphazardly.
- − Contains significant spelling errors ('APOLIO E 11', 'Translurian', 'Descenty').
- − The Saturn V icon has incorrect proportions and rocket stage layout.
Verdict: Qwen Image 2.0 is the clear winner as it successfully understands the complex instruction to create a sequential, six-step infographic with corresponding icons. While it has a minor spelling typo, it is vastly more functional and accurate than Z-Image Turbo, which fails to structure the steps correctly and contains multiple egregious spelling mistakes.
Explore each model
Tongyi-MAI's 6-billion parameter distilled text-to-image model optimized for speed, achieving high-quality generation in 8 steps or fewer with support for bilingual text rendering