The Max series of Tongyi Qwen’s image generation model excels across a wide range of generation tasks. Compared with the Plus series, it significantly reduces the “AI-like” feel in generated images, enhancing their realism. It delivers more lifelike material textures for human subjects, finer and more detailed natural textures, and more visually appealing text rendering.
Settled by community votes across 8 shared challenges, with an AI judge weighing in on each.
Qwen Image Max
#35 of 62 in Text-to-Image
Z-Image Turbo
#12 of 62 in Text-to-Image
Where the votes landed
Qwen Image Max
50.0%
win rate
Ties
0.0%
Z-Image Turbo
50.0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
Qwen Image Max
- + Excellent handling of complex light refractions and reflections within the glass.
- + High level of photographic realism and sharpness.
- + Creative floating effect for the sphere adds visual interest.
- − The sphere appears significantly larger than the 'small' descriptor in the prompt.
- − The glass cube has some internal structural inconsistencies where the back panels meet.
Z-Image Turbo
- + Perfectly follows the size descriptor for a 'small' blue sphere.
- + Very clean and simple composition that feels grounded.
- + Accurately places the plant behind the cube as requested.
- − The plant is quite blurry and lacks detail compared to Image A.
- − The reflection of the sphere on the bottom of the cube is slightly offset from the sphere's actual position.
Verdict: Both models followed the prompt instructions perfectly, including the specific spatial relationships of the objects. Qwen Image Max produced a more visually striking image with superior lighting and texture, although Z-Image Turbo followed the 'small' scale of the sphere more accurately.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
Qwen Image Max
- + Exceptional detail on the engraved armor patterns
- + Stronger adherence to the request for detailed texture on leather and cloth
- + Lifelike skin texture with realistic dirt and weathering
- − The torch in the background is slightly blurry/indistinct compared to Model B
Z-Image Turbo
- + Excellent highlight reflections on the metal plate
- + Clearer inclusion of the torch within the frame
- + Good shallow depth of field effect
- − Armor engravings are less intricate than Model A
- − Leather and cloth textures are softer and lack the requested high-detail grit
- − Braid beads are very small and less prominent
Verdict: Qwen Image Max captures the 'battle-worn' aesthetic much more effectively through highly detailed skin textures, realistic dirt, and frayed leather straps. While Z-Image Turbo produces a high-quality cinematic image with great lighting, Qwen Image Max follows the specific texture and engraving requirements of the prompt more closely.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
Qwen Image Max
- + Excellent sense of motion and 'exploded' effect as requested.
- + Highly realistic food textures on the patty and bun.
- + Vibrant and high-contrast fiery background with effective smoke.
- − The main title text is slightly thin and less 'professional' in its font choice.
- − The burger is slightly tilted in a way that feels less balanced than a standard ad.
Z-Image Turbo
- + Perfect text rendering with high-quality glowing font treatments.
- + Great implementation of the starburst for the price tag.
- + Excellent lighting on the burger patties and cheese.
- − Failed the 'exploded burger' requirement; the burger is mostly stacked rather than having suspended components.
- − The composition is a bit more static compared to the dynamic motion requested.
Verdict: Qwen Image Max successfully captured the 'exploded' motion requested in the prompt, creating a much more dynamic and energetic visual. However, Z-Image Turbo produced superior typography and a cleaner overall advertising layout, despite failing to actually separate the burger's components in mid-air.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
Qwen Image Max
- + Excellent interior detail with realistic textures and lighting.
- + The woman's expression perfectly matches the 'bored' prompt requirement.
- + High-quality fur rendering and natural integration of the capybara's head.
- − The capybara's hands look more like monkey hands than capybara paws.
- − The taxi ceiling has some strange cracking artifacts.
Z-Image Turbo
- + The taxi driver cap has a more authentic 'professional' look with the emblem.
- + Clean, sharp composition that clearly shows both the driver and passenger.
- + The capybara's expression is very calm and stony.
- − The lighting feels slightly flatter and less night-like than Image A.
- − The passenger is looking away but doesn't quite convey the same level of 'boredom' as requested.
- − The capybara's paws are somewhat distorted on the steering wheel.
Verdict: Qwen Image Max is the winner for its superior atmospheric lighting and better character acting; the woman's 'bored' expression perfectly captures the humor of the prompt. While Z-Image Turbo creates a cleaner image with a better taxi cap, it lacks the cinematic realism and fine interior detail found in Qwen Image Max.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
Qwen Image Max
- + Perfect text rendering for all information including title, banner, and date.
- + Highly cohesive composition with a professional border and moody lighting.
- + Accurate adherence to all requested elements including the scroll banner and twisted trees.
- − The parchment texture is slightly less distressed than Model B.
Z-Image Turbo
- + Authentic parchment texture with torn edges and layered depth.
- + Maintains high contrast and a spooky aesthetic.
- + Strong visual quality of the jack-o-lantern and surrounding atmosphere.
- − Spelling error in the location text ('The Archves' instead of 'The Arches').
- − The banner for the sub-text is missing, with text placed directly on the parchment instead.
- − The banner scrolls at the center are empty and purposeless.
Verdict: Qwen Image Max is the clear winner as it followed all textual instructions perfectly and integrated the scroll banner as requested. While Z-Image Turbo has a nice layered parchment effect, it failed on the location spelling and placed the invite text outside of the banner elements it created.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
Qwen Image Max
- + Perfectly renders the Requested text and the correct Japanese flag
- + Excellent rendering of various sushi types with high-quality PBR textures
- + Creative use of a glass diorama base that adds a premium feel
- − The sushi variety exceeds the 'minimal garnish and plate' request slightly
Z-Image Turbo
- + Simple, clean isometric composition adhering to the miniature style
- + Soft, clay-like cartoon textures that match the stylistic request
- − Factual error: displays the flag of China instead of Japan
- − Sushi construction is nonsensical with a green slice inside the rice
Verdict: Qwen Image Max is the clear winner as it correctly follows all instructions, including the text and the specific flag of the country mentioned. Z-Image Turbo unfortunately included the flag of China despite the text explicitly saying 'JAPAN', and the quality of the 3D model is much lower than Qwen Image Max.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
Qwen Image Max
- + Excellent depiction of god rays and dynamic lighting
- + Beautiful, lush environment with high flower density
- + Captures a complex 'tumblingtogether' action between characters
- − Missed the rabbit character entirely
- − Replaced the rabbit with a second golden retriever puppy
- − Fur textures feel slightly more digital and less natural than the competitor
Z-Image Turbo
- + Successfully includes all four requested animals: puppy, kitten, bunny, and fox
- + Incredible fur detail and soft, photorealistic texture
- + Playful, expressive faces that perfectly match the 'wholesome joyful vibe'
- − Lighting is a bit more washed out with less distinct god rays than requested
- − Composition is a bit more static, with animals standing rather than 'tumbling'
Verdict: Z-Image Turbo is the clear winner as it successfully rendered all four specific animals requested in the prompt, whereas Qwen Image Max failed by omitting the rabbit and doubling up on the puppy. While Qwen Image Max had more dramatic lighting and environmental detail, Z-Image Turbo's superior prompt adherence and beautiful character rendering make it the better image.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
Qwen Image Max
- + Excellent typography with correct accents and distinct styling
- + Beautifully detailed engraving style that fits the 'vintage' prompt
- + Perfectly followed the instruction for an 'Est. 1720' banner
- − Texture on the background is a bit heavy/distracting compared to the logo
- − Slightly less 'minimalist' than a standard vector logo
Z-Image Turbo
- + True minimalist vector aesthetic
- + Clean and simple composition for a logo
- + Accurate spelling and accent mark
- − Missed the 'banner' element for the date
- − Typography feels a bit generic compared to the vintage request
- − The steam effect is very basic
Verdict: Qwen Image Max better captured the requested 'vintage' and 'banner' elements with superior typography and a sophisticated engraving style. Z-Image Turbo followed the 'minimalist' instruction more closely but failed to include the banner for the established date and lacked the character of a classic cafe logo.
Explore each model
Tongyi-MAI's 6-billion parameter distilled text-to-image model optimized for speed, achieving high-quality generation in 8 steps or fewer with support for bilingual text rendering