Head to head
Esc

Models · slot A

to navigate to pick

Qwen Image 2.0 Alibaba Stable Diffusion 3.5 Medium Stability AI

Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.

Qwen Image 2.0

21.7 arena score

#34 of 62 in Text-to-Image

Skill signature · Text-to-Image

Stable Diffusion 3.5 Medium

16.9 arena score

#57 of 62 in Text-to-Image

Vote tally

Where the votes landed

Qwen Image 2.0

33.3%

win rate

Ties

33.3%

Stable Diffusion 3.5 Medium

33.3%

win rate

33.3% 33.3% ties 33.3%
Shared challenges 13

Challenge by challenge

The strongest take from each model on every shared challenge, with the AI judge's read.

Geometric Composition

Text-to-Image

“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”

Qwen Image 2.0
Stable Diffusion 3.5 Medium

AI Judge Analysis

Qwen Image 2.0

  • + Excellent photorealistic texture on the wooden table and red book.
  • + Highly accurate glass reflections and refractions.
  • + Superior composition with more natural soft lighting from the window.
  • Internal reflections of the sphere appear a bit cluttered.

Stable Diffusion 3.5 Medium

  • + Successfully follows all prompt instructions including object placement.
  • + Bright, vibrant colors.
  • The book appears merged into the glass cube rather than sitting on top of it.
  • Lower image resolution and visible noise in the blurry background.
  • The blue sphere looks somewhat digital and flat compared to the surrounding environment.

Verdict: Qwen Image 2.0 is the clear winner due to its significantly higher visual quality and realistic rendering of materials. While both models followed the spatial instructions perfectly, Stable Diffusion 3.5 Medium suffered from poor integration of the book and cube, resulting in a less convincing physical interaction.

Candid Street Photography

Text-to-Image

“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”

Qwen Image 2.0
Stable Diffusion 3.5 Medium

AI Judge Analysis

Qwen Image 2.0

  • + Excellent skin texture and realistic weathering on the face
  • + Captures the 'imperfect framing' and 'candid' feel much more effectively
  • + Correct mechanical depiction of circular bike chain and pedaling assembly
  • The motion blur on the car in the background is subtle, appearing more as a static bokeh
  • Hand anatomy on the bicycle pedal is slightly messy upon close inspection

Stable Diffusion 3.5 Medium

  • + Stronger atmospheric effect with visible rain droplets and wet street reflections
  • + Good use of color and cinematic lighting
  • + Effective shallow depth of field
  • The man is hunched over and leaning on the bike rather than actively repairing it
  • Severe anatomy issues with the man's hands merging into the handlebars
  • The bicycle geometry is distorted with the frame and tires appearing warped

Verdict: Qwen Image 2.0 much better fulfilled the request for a realistic, candid photo with natural skin textures and believable subject matter. While Stable Diffusion 3.5 Medium captured the moody atmosphere and rain effects well, it suffered from significant anatomical and structural failures in the man's hands and the bike's frame.

Fantasy Warrior

Text-to-Image

“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”

Qwen Image 2.0
Stable Diffusion 3.5 Medium

AI Judge Analysis

Qwen Image 2.0

  • + Excellent high-frequency detail on the skin texture and scars
  • + Highly detailed ornate engraving on the breastplate and shoulder armor
  • + Perfect adherence to hair braids with distinct small beads
  • The hand resting on the hilt has some anatomical awkwardness with the finger lengths
  • The background fire is slightly overwhelming and lacks subtle bokeh depth

Stable Diffusion 3.5 Medium

  • + Beautiful cinematic lighting on the face with effective highlight/shadow interplay
  • + Captures an intense and lifelike expression in the eyes
  • + Good use of bokeh sparks and shallow depth of field
  • Missed the request for beads in the hair braids
  • The engraving on the armor is less sharp and intricate compared to the competitor
  • Armor texture looks a bit flat and less authentic than the other version

Verdict: Qwen Image 2.0 provided much better prompt adherence by including the specific 'small beads' in the hair and showing much more intricate detail in the armor engravings. While Stable Diffusion 3.5 Medium has very strong cinematic lighting and character expression, it failed on the specific bead detail and had less texture on the primary gear elements.

Modern Clean Menu

Text-to-Image

“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”

Qwen Image 2.0
Stable Diffusion 3.5 Medium

AI Judge Analysis

Qwen Image 2.0

  • + Clean, professional grid layout that perfectly matches the 'modern minimalist' prompt.
  • + High-quality food photography with consistent lighting and vibrant colors.
  • + Excellent text rendering for section headers like 'APPETIZERS', 'PIZZA', and 'MAINS'.
  • Item names are largely gibberish despite the clear headers.
  • Placeholder prices are repetitive, using '100' for almost every item.

Stable Diffusion 3.5 Medium

  • + Attempts a multi-page or bifold layout more common in physical menus.
  • + Includes small text details like dotted leaders for prices.
  • Layout is cluttered and lacks the 'modern minimalist' aesthetic requested.
  • Images are lower resolution with significant compression artifacts and 'bloody' color bleeding.
  • Text rendering is very poor, with distorted characters and unreadable headers.

Verdict: Qwen Image 2.0 followed the prompt much more effectively, delivering a clean, professional grid layout with high-quality food photography and crisp main headers. Stable Diffusion 3.5 Medium produced a cluttered design with significant text distortion and lower-quality imagery that did not capture the 'minimalist' or 'modern' feel requested.

Magic Burger Explosion: Fiery Photorealism Challenge

Text-to-Image

“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”

Qwen Image 2.0
Stable Diffusion 3.5 Medium

AI Judge Analysis

Qwen Image 2.0

  • + Excellent typography with a cohesive fiery glowing effect that matches the prompt perfectly.
  • + Outstanding photorealistic detail in the food textures, especially the melting cheese and fresh lettuce.
  • + Dynamic composition with a sense of motion and 'exploded' levitation.
  • The 'LIMITED TIME ONLY' text is slightly clipped/crowded on the right edge.
  • The background is slightly more generic than the foreground quality.

Stable Diffusion 3.5 Medium

  • + Successfully renders all requested text with correct spelling.
  • + Good use of vertical space with various font sizes.
  • + Vibrant fire background provides a strong sense of heat.
  • The burger is not truly 'exploded' as requested, appearing mostly assembled and static.
  • The starburst graphic for the price is very simplistic and lacks the 'fiery, glowing effect' requested.
  • The 'MAGIC BURGER' title lacks the requested fiery rendering effect.

Verdict: Qwen Image 2.0 followed the prompt details much more effectively, particularly the 'exploded' motion of the burger components and the fiery styling of the text. Stable Diffusion 3.5 Medium produced a largely static burger and failed to apply the requested glowing fire effect to the typography, resulting in a less professional advertisement look.

Chalkboard Menu

Text-to-Image

“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”

Qwen Image 2.0
Stable Diffusion 3.5 Medium

AI Judge Analysis

Qwen Image 2.0

  • + Excellent text rendering with no spelling errors.
  • + Responds perfectly to the requested date and pricing details.
  • + Highly realistic chalk texture and natural handwriting variations.
  • The title is in a clean sans-serif/print hybrid rather than the requested 'elegant cursive'.

Stable Diffusion 3.5 Medium

  • + Stylized chalk borders and artistic layout.
  • + Hand-drawn decorative elements add to the cafe aesthetic.
  • Significant spelling errors and illegible text throughout the board.
  • Failed to include the requested date 'April 30, 2026' accurately.
  • Text looks heavily garbled compared to the natural handwriting request.

Verdict: Qwen Image 2.0 followed the complex text instructions almost perfectly, rendering the full menu accurately with realistic chalk smears and natural handwriting. Stable Diffusion 3.5 Medium struggled with the text generation, resulting in numerous spelling errors and a cluttered, illegible layout that failed to meet the prompt's detail requirements.

The Reversed Rodeo

Text-to-Image

“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”

Qwen Image 2.0
Stable Diffusion 3.5 Medium
33% wins 33% ties 33% wins

AI Judge Analysis

Qwen Image 2.0

  • + Excellent high-frequency details on the spacesuit and horse's mane
  • + Cinematic lighting with a clear sense of depth and scale against Earth
  • + Creative surreal touches like the scales on the horse and floating water droplets
  • Failed the specific spatial instruction for the horse to be 'on top' of the astronaut

Stable Diffusion 3.5 Medium

  • + Naturalistic posing of the rider on the horse
  • + Good wide-angle cinematic composition
  • + Clear rendering of the Earth's atmosphere and surface below
  • Failed the specific spatial instruction for the horse to be 'on top' of the astronaut
  • Noticeable anatomical errors in the horse's legs
  • Lower texture detail compared to the competitor

Verdict: Both Qwen Image 2.0 and Stable Diffusion 3.5 Medium failed the negative/spatial constraint to have the 'horse on top' of the astronaut, instead providing the standard rider configuration. Qwen Image 2.0 is the superior image due to its much higher level of detail, beautiful rendering of the spacesuit, and creative surreal elements like the horse's scales and floating droplets.

The Capybara Taxi Driver

Text-to-Image

“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”

Qwen Image 2.0
Stable Diffusion 3.5 Medium

AI Judge Analysis

Qwen Image 2.0

  • + Excellent adherence to the 'paws on steering wheel' instruction.
  • + The passenger is clearly looking at a phone with a bored expression as requested.
  • + Realistic lighting and texture on the capybara's fur and the taxi interior.
  • The passenger's hands look slightly distorted around the phone.

Stable Diffusion 3.5 Medium

  • + High cinematic quality with a nicely detailed chauffeur hat.
  • + Good background blur and lighting bokeh for a Manhattan night feel.
  • The capybara's paws are not on the steering wheel, failing a key part of the prompt.
  • The passenger is not looking at a phone as requested.
  • The perspective makes it look like the capybara is sitting in the center or on the dashboard rather than in the driver's seat.

Verdict: Qwen Image 2.0 followed every specific detail of the prompt, including the complex interaction of the capybara's paws on the wheel and the passenger's expression and activity. Stable Diffusion 3.5 Medium produced an aesthetically pleasing image but failed to include the phone or the correct driving posture.

The Halloween Invitation

Text-to-Image

“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”

Qwen Image 2.0
Stable Diffusion 3.5 Medium

AI Judge Analysis

Qwen Image 2.0

  • + Excellent typography with perfect spelling of all requested text.
  • + High-quality, cinematic lighting on the central jack-o-lantern.
  • + Strong adherence to the border of webs and thorns requested in the prompt.
  • The parchment texture is a bit safe and flat compared to the complex background.

Stable Diffusion 3.5 Medium

  • + Dynamic composition with multiple pumpkins and a torn edges effect on the parchment.
  • + Atmospheric color palette with a good contrast between the blues and oranges.
  • Numerous spelling errors including 'Halloweeen Inviloween' and 'The Aches'.
  • Failed to include the central jack-o-lantern, placing them in corners instead.
  • Did not include the requested 'scroll banner' for the secondary text.

Verdict: Qwen Image 2.0 followed the prompt instructions perfectly, rendering all text accurately and placing all compositional elements exactly where requested. Stable Diffusion 3.5 Medium struggled significantly with the text rendering and failed to include specific layout elements like the scroll banner and central jack-o-lantern.

Isometric Miniature Diorama Scenes

Text-to-Image

“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”

Qwen Image 2.0
Stable Diffusion 3.5 Medium

AI Judge Analysis

Qwen Image 2.0

  • + Excellent typography rendering with crisp, bold characters.
  • + Includes all requested elements including the flag icon.
  • + Higher level of detail in the food textures and variety of sushi.
  • The style leans more toward a realistic photograph than a 3D cartoon diorama.
  • The wooden base is a bit large, making it feel less like a 'miniature' scale.

Stable Diffusion 3.5 Medium

  • + Successfully captures the 3D cartoon/miniature aesthetic requested.
  • + Good use of isometric perspective and lighting.
  • + The simplified shapes align well with a stylized UI or game asset look.
  • Missing the requested flag icon.
  • Text rendering is slightly inconsistent with artifacts around the letters 'S' and 'U'.
  • The 'JAPAN' text is small and lacks the 'large bold' emphasis requested.

Verdict: Qwen Image 2.0 followed the prompt instructions more comprehensively, including all text and icons with perfect clarity. While Stable Diffusion 3.5 Medium captured the 'cartoon' style better, it failed to include the flag and struggled with the aesthetic cleanliness of the typography.

Adorable Baby Animals in Sunny Meadow

Text-to-Image

“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”

Qwen Image 2.0
Stable Diffusion 3.5 Medium

AI Judge Analysis

Qwen Image 2.0

  • + Successfully includes all four requested animals (puppy, kitten, bunny, and fox).
  • + Excellent rendering of lighting and 'god rays' through the background trees.
  • + Realistic fur texture and interaction between the animals.
  • The fox kit on the ground has slightly distorted anatomy/face due to the pose.

Stable Diffusion 3.5 Medium

  • + Vibrant colors and a very clean, high-saturation aesthetic.
  • + Well-defined facial features and expressive eyes on the foreground animals.
  • Failed to include the bunny, which was a specific part of the prompt.
  • The 'kitten' looks more like a hybrid creature or a fox-eared cat.
  • Missing the 'tumbling together' interaction, feeling more like a posed lineup.

Verdict: Qwen Image 2.0 is the clear winner as it followed the complex prompt instructions to include four different animals and showed them actually interacting. Stable Diffusion 3.5 Medium missed the bunny entirely and produced a more generic, static composition with less realistic fur textures.

Vintage Cafe Logo

Text-to-Image

“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”

Qwen Image 2.0
Stable Diffusion 3.5 Medium

AI Judge Analysis

Qwen Image 2.0

  • + Excellent typography including the grave accent in 'Caffè'.
  • + Perfect adherence to all prompt elements like the steam, cloche, and banner.
  • + Clean, vector-style execution with accurate date rendering.
  • The banner lines are slightly disconnected from the main emblem.
  • The 'steam' looks a bit more like a flame icon.

Stable Diffusion 3.5 Medium

  • + Beautiful hand-drawn etching style that fits the vintage theme.
  • + Nice use of color and texture for a 'classic' look.
  • Significant spelling errors: 'Florrian' and 'Est 170' instead of 1720.
  • The cloche dome shape is distorted and looks more like a rounded bulb or hat.
  • Fails to include the steam element clearly.

Verdict: Qwen Image 2.0 followed the prompt perfectly, producing a clean, professional logo with accurate text and symbols. Stable Diffusion 3.5 Medium had a more interesting artistic texture but failed significantly on the text rendering and the specific shape of the requested cloche dome.

Apollo 11: Journey to Tranquility

Text-to-Image

“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”

Qwen Image 2.0
Stable Diffusion 3.5 Medium

AI Judge Analysis

Qwen Image 2.0

  • + Excellent adherence to the requested sequence of steps
  • + Most text is correctly spelled and logically placed
  • + Clean, professional vector illustration style with a clear vertical flow
  • Typo in 'Translunjar' (contains an extra 'j')
  • The 'Launch' icon is very small compared to other elements

Stable Diffusion 3.5 Medium

  • + Consistent color palette following the NASA-inspired prompt
  • + Artistic composition that feels space-themed
  • Severe text rendering issues with nonsense words
  • Incorrect iconography for the steps (e.g., a person inside the 'Descant' step)
  • Fails to follow the logical sequence of the mission steps

Verdict: Qwen Image 2.0 followed the prompt's logical instructions perfectly, creating a clear and readable infographic that accurately depicts the mission stages. Stable Diffusion 3.5 Medium failed significantly on the text rendering and the specific iconography required for each step, resulting in a disorganized and nonsensical layout. Qwen Image 2.0 is the clear winner for its functional design and high prompt adherence.

Next steps

Explore each model