Head to head
Esc

Models · slot A

to navigate to pick

Qwen Image Alibaba Qwen Image 2.0 Alibaba

Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.

Qwen Image

20.3 arena score

#39 of 62 in Text-to-Image

Skill signature · Text-to-Image

Qwen Image 2.0

21.7 arena score

#34 of 62 in Text-to-Image

Vote tally

Where the votes landed

Qwen Image

0%

win rate

Ties

0%

Qwen Image 2.0

0%

win rate

Shared challenges 13

Challenge by challenge

The strongest take from each model on every shared challenge, with the AI judge's read.

Geometric Composition

Text-to-Image

“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”

Qwen Image
Qwen Image 2.0

AI Judge Analysis

Qwen Image

  • + Excellent photographic quality with realistic depth of field.
  • + Accurate representation of glass reflections and material properties.
  • + Clean and balanced composition following all prompt instructions.
  • The blue sphere is sitting on the bottom rather than being suspended, though the prompt was ambiguous on this.

Qwen Image 2.0

  • + Creative interpretation with the sphere suspended in the center.
  • + Rich texture on the book cover and wooden table.
  • + Clear visibility of the plant through the glass panels.
  • The reflections in the glass are physically confusing and cluttered.
  • The sphere appears to be levitating without a clear medium like water to support it.
  • The glass cube edges look more like a frame than a solid glass object.

Verdict: Both models followed the prompt's spatial instructions perfectly. Qwen Image is preferred because it achieves a more realistic and cohesive photographic look, whereas Qwen Image 2.0 introduces confusing reflections and a floating sphere that feels less grounded in reality.

Candid Street Photography

Text-to-Image

“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”

Qwen Image
Qwen Image 2.0

AI Judge Analysis

Qwen Image

  • + Excellent full-body composition and lighting
  • + Captures the atmosphere of light rain and street reflections beautifully
  • + Realistic 50mm-style depth of field
  • Physical logic errors with the bicycle frame and pedal placement
  • The car in the background lacks sufficient motion blur requested in the prompt

Qwen Image 2.0

  • + Highly realistic skin textures and facial details
  • + Authentic 'imperfect' framing that feels like a candid snapshot
  • + Superior mechanical detail on the bicycle
  • The subject's hands have significant anatomical errors (merged fingers/extra joints)
  • Motion blur on the background car is still minimal

Verdict: Both models followed the prompt well, but Qwen Image 2.0 produced a much more realistic texture and 'candid' feel, capturing the natural skin and 50mm lens look more effectively. Qwen Image produced a more aesthetically pleasing full-body shot, but it failed significantly on the structural logic of the bicycle. However, Qwen Image 2.0's failure on the hand anatomy is a significant detractor compared to the cleaner execution in Qwen Image.

Fantasy Warrior

Text-to-Image

“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”

Qwen Image
Qwen Image 2.0

AI Judge Analysis

Qwen Image

  • + Excellent depiction of ornate engraved armor with detailed filigree
  • + Accurate representation of warmth from a torchlight source
  • + High-quality rendering of leather straps and buckles
  • The sparks look like flat star-shaped graphics rather than photographic bokeh
  • The facial scars appear more like fresh paint than healed battle wounds

Qwen Image 2.0

  • + Extremely realistic skin texture, dirt, and lifelike eyes
  • + Natural-looking hair braiding and beads integrated into the character design
  • + Superior bokeh effect and photographic depth of field
  • The lighting is bright and outdoor-like, missing the specific 'warm torchlight' atmosphere
  • Portions of the hand and sword hilt are slightly blurry or less defined

Verdict: Qwen Image 2.0 produces a significantly more realistic and 'lifelike' image with professional-grade skin textures and lighting, whereas Qwen Image has a more illustrative, digital-art feel. While Qwen Image followed the specific 'torchlight' lighting prompt better, Qwen Image 2.0 is the superior image due to its incredible detail in the face, hair, and believable battle-worn weathering.

Modern Clean Menu

Text-to-Image

“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”

Qwen Image
Qwen Image 2.0

AI Judge Analysis

Qwen Image

  • + Excellent typography layout for a modern menu
  • + Clean white space and Professional design aesthetic
  • + Correct sectioning for Appetizers and Pizza/Mains
  • Nonsense 'Pizzaurant' header text
  • The food images are repetitive and less realistic
  • Text is somewhat garbled in the smaller sections

Qwen Image 2.0

  • + High-quality, appetizing food photography
  • + Clear, bold sans-serif headers for all three requested sections
  • + More diverse food representation aligned with categories
  • Text descriptions are gibberish
  • Missing the price/description lines typically found in menus
  • Layout feels more like a grid gallery than a functional menu

Verdict: Qwen Image (Model A) delivers a much more realistic menu layout that feels like a professional design piece, despite some text errors. Qwen Image 2.0 (Model B) has significantly better food photography but fails to capture the 'menu' structure, presenting more as a digital food gallery. Model A is preferred for capturing the specific layout and design intent of the prompt.

Magic Burger Explosion: Fiery Photorealism Challenge

Text-to-Image

“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”

Qwen Image
Qwen Image 2.0

AI Judge Analysis

Qwen Image

  • + Excellent text readability and layout.
  • + Vibrant color palette with a clear 'fiery' aesthetic.
  • + Included all requested price and promotional text correctly.
  • The burger is less 'exploded' and more of an assembled burger with flying garnish.
  • Graphic elements like the starburst look a bit generic/clip-art style.

Qwen Image 2.0

  • + Successfully applied the fiery, glowing effect to the text as requested.
  • + High level of photorealistic detail in the food textures and sauce drips.
  • + Good sense of vertical motion and depth with the embers and smoke.
  • The 'starburst' for the price is more of an explosion effect, making the text slightly harder to read.
  • Composition feels a bit more cramped at the top with the title and subtitle overlapping.

Verdict: Qwen Image 2.0 provides a more photorealistic and stylistically coherent interpretation of the 'fiery' prompt, especially with the burning text effect. While Qwen Image has cleaner graphic design for the price tag, Qwen Image 2.0 captures the professional food photography look and atmospheric motion more effectively.

Chalkboard Menu

Text-to-Image

“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”

Qwen Image
Qwen Image 2.0

AI Judge Analysis

Qwen Image

  • + Excellent legibility and adherence to the specific text requested.
  • + Good framing of the chalkboard within the cafe environment.
  • Year is incorrectly rendered as '20026' instead of '2026'.
  • The text looks more like a digital marker font than authentic chalk texture.
  • Spelling error in 'Risoto' instead of 'Risotto'.

Qwen Image 2.0

  • + Perfect date rendering '2026' as requested.
  • + Highly realistic chalk texture with smudges and varying pressure that feels authentic.
  • + Excellent handwriting style that looks truly hand-drawn rather than font-based.
  • The word 'Chip' is slightly separated and messy in the third item.
  • Missing the '$' sign on the bottom price in the original prompt (though it included '9').

Verdict: Qwen Image 2.0 significantly outperforms the previous version by capturing the authentic texture of a chalkboard, including realistic smudges and hand-lettered variation. While Qwen Image followed the text closely, it failed on the year '2026' and the typography appeared too digital; Qwen Image 2.0 felt more like a real photograph of a cafe menu.

The Reversed Rodeo

Text-to-Image

“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”

Qwen Image
Qwen Image 2.0

AI Judge Analysis

Qwen Image

  • + Clean cinematic lighting with good depth
  • + Realistic horse anatomy and texture
  • + Clear focus on the subject against a defined planetary background
  • The astronaut's hands and the reins have some merging issues
  • Less creative interpretation of the 'surreal' aspect compared to the scale patterns

Qwen Image 2.0

  • + Enhanced surrealism with scale-like patterns on the horse
  • + Dynamic composition with floating water-like droplets
  • + Better integration of the astronaut's hands with the reins
  • The white-on-white color palette makes the astronaut blend into the horse slightly
  • Back leg anatomy is a bit spindly and distorted

Verdict: Both models successfully followed the specific spatial prompt instructions. Qwen Image provides a cleaner, more grounded cinematic look, while Qwen Image 2.0 leans harder into the 'surreal' prompt requirement with scale textures and floating elements, making it a more interesting artistic interpretation.

The Capybara Taxi Driver

Text-to-Image

“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”

Qwen Image
Qwen Image 2.0

AI Judge Analysis

Qwen Image

  • + Excellent adherence to the 'professional expression' and 'both paws on steering wheel' prompt.
  • + Very clean, high-resolution rendering with a polished cinematic feel.
  • + Great character consistency for the capybara's face and uniform.
  • The hands/paws on the steering wheel look more like primate hands than capybara paws.
  • The taxi sign on top displays the nonsensical text 'YOXI'.

Qwen Image 2.0

  • + Naturalistic lighting and reflections that enhance the nighttime Manhattan atmosphere.
  • + The woman's 'bored' expression perfectly matches the prompt's request for apathy.
  • + Paws look more like actual capybara webbing/claws compared to Model A.
  • The woman appears to be sitting in the front passenger seat rather than the back seat as requested.
  • The capybara's left paw is positioned awkwardly and blending into the wheel.
  • Lower overall sharpness compared to Model A.

Verdict: Qwen Image (Model A) is the clear winner for its superior technical quality and adherence to the layout requested in the prompt. While Qwen Image 2.0 (Model B) has great lighting, it failed to place the passenger in the back seat and possesses more internal logic errors in its composition. Qwen Image captured the specific detail of both paws on the wheel and a more professional chauffeur-like appearance for the capybara.

The Halloween Invitation

Text-to-Image

“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”

Qwen Image
Qwen Image 2.0

AI Judge Analysis

Qwen Image

  • + Excellent cinematic lighting and contrast that captures the moody night sky perfectly.
  • + Very clean, modern-gothic font choice for the event details at the bottom.
  • + Higher adherence to the 'dark parchment' request by creating a layered effect with the internal illustration.
  • Significant spelling error in the main title, which reads 'Halle Party Invitation'.
  • Small text at the top of the illustration is garbled and illegible.

Qwen Image 2.0

  • + Perfect text rendering for the main title with no spelling errors.
  • + Better adherence to the scroll banner request, which looks more integrated and vintage.
  • + The jack-o-lantern carving is more expressive and detailed.
  • The parchment texture looks a bit more like a generic filter compared to the cut-out style of Model A.
  • Event details at the bottom are slightly less legible against the background texture.

Verdict: Qwen Image 2.0 is the superior choice because it successfully rendered the main title without the spelling errors found in Qwen Image. While Qwen Image had slightly better cinematic lighting, Qwen Image 2.0 followed all textual instructions accurately and provided a more polished gothic aesthetic for a printable invitation.

Isometric Miniature Diorama Scenes

Text-to-Image

“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”

Qwen Image
Qwen Image 2.0

AI Judge Analysis

Qwen Image

  • + Excellent adherence to the '3D cartoon' and 'miniature' aesthetic.
  • + Creative use of a 3D diorama base that fits the isometric theme.
  • + Legible white text that blends well with the stylized scene.
  • The flag icon in the text area is slightly distorted.
  • The scale of the chopsticks is a bit thick compared to the sushi.

Qwen Image 2.0

  • + High realism in the textures of the fish and rice.
  • + Very clean typography and flag icon representation.
  • + Accurate interpretation of 'PBR materials' for a realistic look.
  • Fails the 'cartoon scene' requirement by being too photo-realistic.
  • The wooden base is less of a 'diorama base' and more of a standard cutting board.

Verdict: Qwen Image is the preferred choice as it perfectly captures the '3D cartoon scene' and 'miniature' request, whereas Qwen Image 2.0 ignored the cartoon styling in favor of high realism. Qwen Image feel more like a cohesive isometric graphic, while Qwen Image 2.0 looks like a standard product photograph with text overlay.

Adorable Baby Animals in Sunny Meadow

Text-to-Image

“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”

Qwen Image
Qwen Image 2.0

AI Judge Analysis

Qwen Image

  • + Excellent depiction of warm golden light and prominent god rays.
  • + Whimsical, clean composition with all animals clearly visible and facing forward.
  • + Strong adherence to the 'big expressive eyes' and 'dew sparkles' descriptors.
  • The lighting feels slightly more filtered/digital rather than purely photorealistic.
  • The poses are a bit static, looking more like a staged photoshoot than 'tumbling together'.

Qwen Image 2.0

  • + Successfully captures the 'tumbling together' prompt with dynamic, playful poses.
  • + Highly realistic fur textures and natural lighting integration.
  • + Superior wildflower meadow detail with a more varied and organic plant life.
  • The rabbit feels slightly disconnected from the main action of the other three animals.
  • The fox's face is somewhat obscured by the angle of the tumble.

Verdict: While both models followed the prompt exceptionally well, Qwen Image 2.0 is the winner for its superior interpretation of movement and 'tumbling' described in the prompt. Qwen Image produced a beautiful, clean image, but Qwen Image 2.0 felt more authentic to the request for photorealism and dynamic playfulness.

Vintage Cafe Logo

Text-to-Image

“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”

Qwen Image
Qwen Image 2.0

AI Judge Analysis

Qwen Image

  • + Successfully includes all text elements requested
  • + Captures a flat vector emblem style well
  • + Uses a consistent and appropriate warm brown and cream palette
  • Confusing typography layout with overlapping and redundant text
  • Steam element looks somewhat disconnected from the cloche knob

Qwen Image 2.0

  • + Excellent typography rendering with clean, professional fonts
  • + Superior illustration quality with better shading and lighting on the cloche
  • + The banner element is more sophisticated and well-integrated
  • Steam is placed inside or in front of the cloche rather than rising from it
  • Slightly less 'minimalist' than the first option

Verdict: Qwen Image 2.0 significantly outperforms the original by providing legible, elegant typography and a much more polished vector illustration. While Qwen Image followed the layout instructions, the chaotic text overlap makes it unusable as a logo, whereas Qwen Image 2.0 produces a professional-grade brand mark.

Apollo 11: Journey to Tranquility

Text-to-Image

“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”

Qwen Image
Qwen Image 2.0

AI Judge Analysis

Qwen Image

  • + Excellent adherence to the 'flat-vector' style with crisp, clean shapes.
  • + Includes secondary icons like the astronauts at the bottom with a cohesive color palette.
  • + Good use of the specific NASA-inspired color scheme.
  • Confusing flow and numbering, skipping steps or mislabeling them (e.g., 'Tranar Orbit').
  • The rocket design looks more like a modern missile than the historic Saturn V requested.
  • Incorrect text rendering for names and headers.

Qwen Image 2.0

  • + Perfect adherence to the requested 6-step sequence in chronological order.
  • + Superior text rendering with almost perfect spelling for complex terms like 'Translunjar' (small typo) and 'Tranquility'.
  • + Icons accurately represent the specific hardware of the Apollo mission, such as the Lunar Module.
  • The layout is a bit cramped vertically, leaving little negative space.
  • Some icons overlap with text (like 'Lunar Orbit'), making it slightly harder to read.

Verdict: Qwen Image 2.0 is the clear winner as it followed the complex 6-step instruction set perfectly and maintained a logical chronological flow. While Qwen Image had a slightly cleaner 'flat' aesthetic, its logic and text were disjointed and failed to capture the specific sequence requested by the prompt.

Next steps

Explore each model