Alibaba's Qwen image model
Settled by community votes across 8 shared challenges, with an AI judge weighing in on each.
Qwen Image
#40 of 62 in Text-to-Image
Qwen Image Max
#35 of 62 in Text-to-Image
Where the votes landed
Qwen Image
0%
win rate
Ties
0%
Qwen Image Max
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
Qwen Image
- + Excellent realization of soft light reflections on the wooden table.
- + Clean and simple object rendering with realistic glass thickness.
- + The plant is positioned exactly behind the cube as requested.
- − The sphere appears to be sitting on a mirrored bottom pane rather than floating/centered.
- − The plant detail is slightly blurry compared to the foreground objects.
Qwen Image Max
- + High level of texture detail on the book cover and wooden grain.
- + Sophisticated lighting with strong directional shadows and light rays.
- + The sphere has a nice satin finish and appears more centrally placed in the volume.
- − The plant is moved to the side rather than being strictly 'behind' the cube.
- − The sphere reflection on the left side of the glass looks slightly unnatural.
Verdict: Both models followed the prompt instructions very well, capturing the complex lighting and transparency. Qwen Image adhered more closely to the spatial request of placing the plant directly behind the cube, while Qwen Image Max produced a more visually striking image with superior textures and more dramatic lighting.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
Qwen Image
- + Excellent engraving detail on the pauldrons and gorget.
- + Dynamic lighting with a strong torchlight source visible in the composition.
- + Accurate beads and braid styles as requested.
- − Scars look more like surface-level scratches or face paint than weathered battle wounds.
- − The sparks have a slightly synthetic, star-shaped filter look.
Qwen Image Max
- + Superior skin texture and 'battle-worn' appearance with believable scars and dirt.
- + Highly detailed leather texture and aged, realistic cloth underlayer.
- + Excellent adherence to the 'faint scars' and 'shallow depth of field' requirements.
- − The torch itself is partially cut off at the edge of the frame.
- − The bokeh sparks are a bit uniform across the image.
Verdict: Qwen Image Max is the winner for its superior rendering of 'battle-worn' textures, particularly the skin, scars, and aged leather, which feel much more grounded and lifelike than Qwen Image. While Qwen Image handles the ornate metal engravings beautifully, the character's facial features and weathering in Qwen Image Max better capture the gritty essence of the prompt.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
Qwen Image
- + Excellent text legibility and clean graphic design elements
- + Captures an 'exploded' view better with many individual floating ingredients
- + Vibrant colors and high-quality rendering of food textures
- − The main burger assembly isn't fully 'exploded' as much as surrounded by debris
- − The fiery background is a bit generic and blurred
Qwen Image Max
- + Applied a superior fiery, glowing effect to all text as requested
- + Strong sense of motion and impact with speed lines and flying embers
- + Photorealistic texture on the grilled patty and fresh lettuce
- − Failed to provide an actually 'exploded' burger, showing a mostly assembled one instead
- − The starburst is a bit chaotic compared to the cleaner design in the other image
Verdict: Both models followed the text instructions perfectly. Qwen Image provides a cleaner, more traditional commercial layout with a better 'exploded' feel, whereas Qwen Image Max delivers a much more cinematic atmosphere with superior fiery text effects and dynamic lighting, despite keeping the burger mostly intact.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
Qwen Image
- + Excellent adherence to the 'bored expression' prompt for the passenger.
- + The cap is a more formal taxi driver style, matching the professional description.
- + Good use of depth of field to separate the subjects from the background.
- − The hands on the steering wheel look more like primate hands than capybara paws.
- − The taxi sign on the roof appears somewhat distorted and says 'YOYI' instead of 'TAXI'.
Qwen Image Max
- + Highly realistic interior textures and lighting, including the worn ceiling of the taxi.
- + Superior fur texture and facial lighting on the capybara.
- + Better camera angle that captures the street environment more effectively.
- − The passenger looks sad or tired rather than 'bored' and appearing to ignore the capybara.
- − The capybara's hands are also inaccurately rendered, resembling human/primate hands.
Verdict: Qwen Image Max is the superior model in this comparison due to its significantly higher level of detail in the textures and lighting, creating a more convincing 'film-like' quality. While Qwen Image followed the passenger's expression prompt slightly better, Qwen Image Max provided a much more immersive and detailed environment that captured the essence of a gritty New York taxi at night.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
Qwen Image
- + Features a unique 'parchment within a parchment' aesthetic that fits the dark moody theme.
- + Includes the specific thorn and web border as requested.
- + The glowing pumpkin provides excellent central lighting.
- − The main title text 'Halloween Party Invitation' contains spelling errors ('Halle Party').
- − The scroll banner text is split awkwardly above and on the banner.
Qwen Image Max
- + Perfect text rendering for all requested strings, including the title and scroll.
- + Superior visual quality with more detailed twisted trees and a more menacing jack-o-lantern.
- + The composition feels more cohesive as a single invitation poster.
- − The thorns in the border are slightly less distinct than in Model A.
- − The background lighting is a bit brighter, sacrificing some of the 'dark parchment' feel.
Verdict: Qwen Image Max is the clear winner because it successfully rendered all requested text accurately, whereas Qwen Image failed on the primary title. Qwen Image Max also provided a more polished artistic style with better integration of the bats, trees, and jack-o-lantern within the gothic theme.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
Qwen Image
- + Excellent adherence to the 'cartoon' and 'miniature' aesthetic.
- + Clean, playful layout with a charming diorama base.
- + Accurate text rendering and includes all requested elements like the flag icon.
- − The salmon nigiri texture is a bit simplistic/plastic compared to Model B.
- − Chopsticks are slightly thick and toy-like.
Qwen Image Max
- + High-quality PBR materials with realistic textures on the salmon and eel.
- + Clean, modern text layout with subtle drop shadows.
- + Elegant glass diorama base adds a premium feel.
- − The sushi models are less 'cartoon' than requested, leaning toward photo-realism.
- − Lacks the 3D-molded flag requested in the scene, opting for a 2D icon next to text.
Verdict: Qwen Image (Model A) followed the prompt's 'cartoon' and 'diorama' instructions more faithfully, creating a cohesive miniature scene that includes a physical flag. Qwen Image Max (Model B) produced significantly better textures and more realistic materials, but it missed the 'cartoon' style request and left the flag out of the physical 3D scene. Model A is the likely winner for better stylistic consistency with the prompt.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
Qwen Image
- + Excellent adherence to the list of specific animals (pupil, kitten, bunny, and fox).
- + Clear, high-quality lighting with well-defined god rays.
- + Clean character silhouettes without significant anatomical merging.
- − The posing is somewhat static and posed rather than 'tumbling'.
- − The kitten has an anatomically odd front paw that looks like a hand.
Qwen Image Max
- + Successfully captures the 'tumbling together' aspect of the prompt.
- + Lush environment with a high density of flowers and colorful butterflies.
- + Dynamic movement and interaction between the animals.
- − Missing the baby bunny requested in the prompt.
- − The golden retriever is duplicated, with two puppies appearing instead of one.
- − The fox's front paw is touching the dog's head in a slightly distorted way.
Verdict: Qwen Image accurately included all four animals requested in the prompt, whereas Qwen Image Max missed the baby bunny and duplicated the puppy. While Qwen Image Max better captured the energy of 'tumbling,' Qwen Image is the superior output due to its strict adherence to the subject list and cleaner anatomical rendering.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
Qwen Image
- + Strong minimalist vector aesthetic
- + Clean and simple color palette
- − Confused and overlapping typography in the main brand name
- − Banner design feels a bit heavy and flat
Qwen Image Max
- + Clear and legible typography for the main title
- + Excellent use of texture and shading to create a vintage feel
- + Sophisticated rendering of the cloche and steam
- − Less minimalist than requested due to high detail density
Verdict: Qwen Image Max is the superior choice because it correctly renders the text 'Caffè Florian', whereas Qwen Image produced a jumbled, overlapping mess of letters. Qwen Image Max also excelled at the 'vintage' aspect of the prompt, providing a beautiful paper texture and elegant vector shading that gives the logo more depth and professional appeal.
Explore each model
The Max series of Tongyi Qwen’s image generation model excels across a wide range of generation tasks. Compared with the Plus series, it significantly reduces the “AI-like” feel in generated images, enhancing their realism. It delivers more lifelike material textures for human subjects, finer and more detailed natural textures, and more visually appealing text rendering.