Alibaba's Qwen Image 2.0 model with enhanced text rendering, supporting both Chinese and English prompts with up to 6 images per request
Settled by community votes across 8 shared challenges, with an AI judge weighing in on each.
Qwen Image 2.0
#34 of 62 in Text-to-Image
Qwen Image Max
#35 of 62 in Text-to-Image
Where the votes landed
Qwen Image 2.0
0%
win rate
Ties
0%
Qwen Image Max
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent photorealistic texture on the wooden table and red book.
- + Highly accurate glass reflections and light refraction.
- + Perfect directional lighting consistent with window placement.
- − The sphere appears to be floating mid-air without physical support, which looks slightly unnatural.
Qwen Image Max
- + Superior composition with the plant potted and integrated into the scene.
- + Clean and sharp rendering of the glass cube and its edges.
- + Good use of shadows and light play across the table surface.
- − The book has black geometric shapes not requested in the prompt.
- − The sphere suffers from similar 'floating' physics as Model A.
Verdict: Both models followed the complex spatial instructions perfectly. Qwen Image 2.0 is the winner due to its superior photorealism, particularly in the textures of the book cover and the wood grain, whereas Qwen Image Max added unnecessary graphics to the book and had slightly softer details on the plant.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent character pose and narrative feel with the hand resting on a sword hilt.
- + Vibrant color palette with realistic light interaction on the polished armor.
- + Highly detailed facial texture including realistic wrinkles and scars.
- − The eyes have a slightly unnatural yellowish/red tint that makes the character look less human/paladin-like.
- − The armor engraving is a bit simpler compared to the competitor.
Qwen Image Max
- + Incredible detail in the ornate engravings on the plate armor.
- + Superior texture work on the weather-beaten leather straps and torn cloth.
- + Stronger adherence to the 'warm torchlight' prompt with the torch being visible in the background.
- − The sparks look slightly more like static glowing dots rather than dynamic embers.
- − The transition between the hair braids and the armor collar is a bit crowded.
Verdict: Both models performed exceptionally well on this complex prompt. Qwen Image Max is the winner because its attention to the 'ornate engraved plate' and 'leather straps and cloth underlayer' instructions resulted in a much more intricate and realistic texture, whereas Qwen Image 2.0 has more basic engraving. While Qwen Image 2.0 has a great composition with the character's hand, Qwen Image Max captures the gritty, battle-worn atmosphere more effectively through its superior detailing and lighting.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent text integration with clean, legible fonts
- + Better representation of an 'exploded' view with multiple tiers of ingredients
- + Very high photorealistic detail on the meat texture and sauce drips
- − The composition feels slightly more static compared to the diagonal motion of the competitor
Qwen Image Max
- + Stronger sense of motion and 'dynamic' energy with diagonal tilt and motion streaks
- + Highly vibrant lighting and impactful starburst effect for the price
- + Great charred texture on the burger patty
- − The starburst overlaps the bun slightly awkwardly
- − The ingredients are less 'exploded' or separated than requested
Verdict: Both models followed the complex prompt exceptionally well, particularly with the text rendering. Qwen Image 2.0 provides a better 'exploded' view that clearly shows every ingredient layer, while Qwen Image Max captures the 'dynamic motion' and fiery atmosphere with more intensity. Qwen Image 2.0 is slightly preferred for its cleaner layout and superior ingredient separation.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent close-up detail on the capybara's fur and expression.
- + Captures the professional taxi driver cap style requested.
- + Good dynamic lighting and window reflections.
- − The passenger is sitting in the front seat instead of the back seat as requested.
- − The capybara's paws look somewhat mangled and unnatural on the steering wheel.
Qwen Image Max
- + Correctly places the passenger in the back seat as specified.
- + Superior interior composition with a clear view of both subjects and the street.
- + Includes realistic taxi details like the cracked ceiling and dashboard gauges.
- − The hat is a standard baseball cap rather than a traditional taxi driver cap.
- − The capybara's hands look more like human/primate hands than capybara paws.
Verdict: While Qwen Image 2.0 has vibrant lighting and a great character design, it failed the spatial requirement of having the passenger in the back seat. Qwen Image Max followed the prompt more accurately, capturing the correct seating arrangement and a very realistic 'bored' expression on the passenger, making it the more effective storytelling image.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent typography with perfect spelling in all text elements.
- + Clean, professional layout that feels like a real invitation.
- + High clarity and balanced composition.
- − The jack-o'-lantern lighting is a bit flat compared to the other model.
Qwen Image Max
- + Superior cinematic lighting with a more dramatic 'glow' effect.
- + Intricate, three-dimensional border design that feels tactile.
- + Detailed and well-rendered spooky bats.
- − The banner text contains a slight typo ('frights' is missing the 's' or it is poorly rendered).
- − The scroll banner has some structural clipping issues on the right side.
Verdict: Both models followed the prompt exceptionally well, but Qwen Image 2.0 is the winner due to its perfect text rendering across all sections of the invitation. While Qwen Image Max had more impressive lighting and a richer border, it struggled with the fine details of the banner text and scroll structure.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent photographic realism in textures
- + Clean typography and clear flag icon
- − Fails the isometric requirement, opting for a standard perspective
- − Does not follow the 3D cartoon/miniature aesthetic requested
- − Flag icon is placed to the side instead of top-center
Qwen Image Max
- + Perfectly follows the 45-degree isometric 3D cartoon style
- + Accurately represents the diorama base using a glass/plastic material
- + Correct top-center typography and layout placement
- − The white text on a light blue background has slightly less contrast than the black text in Model A
- − One sushi roll has a slightly distorted internal texture
Verdict: Qwen Image Max is the clear winner as it perfectly adheres to the 'isometric' and '3D cartoon miniature' stylistic requirements of the prompt, whereas Qwen Image 2.0 produced a traditional high-realism photograph. Qwen Image Max also correctly followed the 'top-center' layout for all text and icons.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
Qwen Image 2.0
- + Successfully included all four requested animals (dog, cat, rabbit, fox).
- + Naturalistic interaction with animals physically touching in a playful way.
- + Excellent lighting effects with subtle god rays and dew sparkles on the flowers.
- − The fox kit's head is slightly distorted where it meets the retriever's paw.
- − The kitten's anatomy is a bit stiff in its upright pose.
Qwen Image Max
- + Very high detail in the fur texture and colorful butterfly wings.
- + Vibrant color palette that enhances the 'joyful vibe' requested.
- + Composition feels balanced and fills the frame well with wildflowers.
- − Failed to include the requested baby bunny.
- − Included a second golden retriever puppy which was not requested.
- − The fox's front leg appears to grow directly out of the puppy's head.
Verdict: Qwen Image 2.0 is the clear winner for prompt adherence, as it correctly included all four specific animals while Qwen Image Max missed the rabbit entirely and added a duplicating puppy. While Qwen Image Max has slightly more vibrant colors, it suffers from significant anatomical merging issues where the animals overlap and fails the basic instruction count.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent typography rendering with the correct grave accent on 'Caffè'.
- + Clean vector-style execution with smooth gradients.
- + Good use of the requested banner element for the date.
- − The steam effect is inside the dome, which looks a bit like a flame rather than rising steam.
- − The banner is slightly asymmetrical and disconnected from the main emblem.
Qwen Image Max
- + Excellent 'vintage' feel with the textured paper background.
- + Rising steam from the top of the cloche is more logical and artistic.
- + Sophisticated woodcut/engraving texture on the cloche and banner.
- − Used an acute accent instead of the correct grave accent on 'CAFFÉ'.
- − The text is slightly more cramped than in the other version.
Verdict: Both models followed the prompt well, but Qwen Image Max produced a more cohesive 'vintage' brand mark with superior texture and logical steam placement. Qwen Image 2.0 has cleaner vector lines and more accurate character rendering, but the design feels less mature as a logo compared to the secondary output.
Explore each model
The Max series of Tongyi Qwen’s image generation model excels across a wide range of generation tasks. Compared with the Plus series, it significantly reduces the “AI-like” feel in generated images, enhancing their realism. It delivers more lifelike material textures for human subjects, finer and more detailed natural textures, and more visually appealing text rendering.