Head to head
Esc

Models · slot A

to navigate to pick

Qwen Image Max Alibaba Stable Diffusion 3.5 Medium Stability AI

Settled by community votes across 8 shared challenges, with an AI judge weighing in on each.

Qwen Image Max

21.3 arena score

#35 of 62 in Text-to-Image

Skill signature · Text-to-Image

Stable Diffusion 3.5 Medium

16.8 arena score

#58 of 62 in Text-to-Image

Vote tally

Where the votes landed

Qwen Image Max

0%

win rate

Ties

0%

Stable Diffusion 3.5 Medium

0%

win rate

Shared challenges 8

Challenge by challenge

The strongest take from each model on every shared challenge, with the AI judge's read.

Geometric Composition

Text-to-Image

“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”

Qwen Image Max
Stable Diffusion 3.5 Medium

AI Judge Analysis

Qwen Image Max

  • + Excellent photorealistic lighting and textures, especially on the wood grain.
  • + Superior handling of reflections and refractions within the glass cube.
  • + High resolution and clear detail on all objects.
  • The plant is more beside the cube than behind it, making the 'visible through glass' effect less central.

Stable Diffusion 3.5 Medium

  • + Accurately depicts the plant behind the cube, visible through the glass.
  • + Follows all prompt elements including color and positioning.
  • The cube geometry is slightly warped and the glass quality looks less realistic.
  • The blue sphere has an odd internal artifact or texture that looks like a smudge.
  • Lower overall visual fidelity compared to the competitor.

Verdict: Qwen Image Max produces a significantly more polished and aesthetically pleasing image with professional-grade lighting and texture work. While Stable Diffusion 3.5 Medium technically follows the spatial positioning of the plant better, its poor glass rendering and distorted cube geometry make it less successful overall.

Fantasy Warrior

Text-to-Image

“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”

Qwen Image Max
Stable Diffusion 3.5 Medium

AI Judge Analysis

Qwen Image Max

  • + Exquisite detail on the engraved armor and weathered leather straps.
  • + Perfect adherence to all prompt elements including hair beads and lifelike eyes.
  • + Excellent depth of field with a clearly defined torch in the background.
  • The sparks have a slightly synthetic, floating appearance in some areas.

Stable Diffusion 3.5 Medium

  • + Intense, high-contrast lighting creates a strong dramatic effect.
  • + Good rendering of the engraved patterns on the pauldrons.
  • + Strong facial expression that conveys a battle-worn character well.
  • Failed to include the requested beads in the braided hair.
  • The dirt on the face looks more like skin texture or freckles than actual grime.
  • The leather and cloth textures are less defined compared to the other model.

Verdict: Qwen Image Max followed the prompt much more accurately, specifically including the requested hair beads and creating a more convincing representation of battle-worn equipment. Stable Diffusion 3.5 Medium produced a striking image with great lighting, but missed key details and had less realistic material transitions between the armor and underlayers.

Magic Burger Explosion: Fiery Photorealism Challenge

Text-to-Image

“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”

Qwen Image Max
Stable Diffusion 3.5 Medium

AI Judge Analysis

Qwen Image Max

  • + Excellent typography with a glowing fire effect that matches the theme.
  • + High degree of photorealism and dynamic 'exploded' composition as requested.
  • + Effective use of lighting and embers to create a cinematic atmosphere.
  • The burger is slightly more 'assembled' than 'exploded', though the ingredients are clearly floating.

Stable Diffusion 3.5 Medium

  • + Clean text rendering for all three requested messages.
  • + Good integration of the background fire and embers.
  • Failed the 'exploded burger' requirement; the burger is fully assembled and merely hovering.
  • The text lacks the requested 'fiery, glowing effect' and appears flat.
  • Composition is static and lacks the sense of motion requested in the prompt.

Verdict: Qwen Image Max is the clear winner as it successfully captured the high-energy, 'exploded' motion requested in the prompt and applied a beautiful fiery glow effect to the text. Stable Diffusion 3.5 Medium produced a static, fully-assembled burger and failed to apply the requested stylistic effects to the typography, resulting in a less professional advertisement.

The Capybara Taxi Driver

Text-to-Image

“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”

Qwen Image Max
Stable Diffusion 3.5 Medium

AI Judge Analysis

Qwen Image Max

  • + Excellent photorealistic texture on the capybara fur and clothing.
  • + Accurately depicts the businesswoman looking at her phone.
  • + Great composition showing the scale and depth of the taxi interior.
  • The hands on the steering wheel look more like primate hands than capybara paws.
  • The taxi ceiling has some strange cracking artifacts.

Stable Diffusion 3.5 Medium

  • + Stronger capybara-specific facial features and front paws.
  • + Good adherence to the 'taxi driver cap' style with the badge detail.
  • + Vibrant bokeh lighting in the background creates a nice night atmosphere.
  • Failed the prompt requirement for the businesswoman to be 'looking at her phone'.
  • The businesswoman's face is quite blurry and lacks detail.
  • Lower overall photorealism compared to Model A.

Verdict: Qwen Image Max is the clear winner as it followed all aspects of the prompt, specifically the interaction of the passenger with her phone. While Stable Diffusion 3.5 Medium captured the capybara's anatomy more accurately, it failed to render the passenger's action and lacked the sharp photorealistic quality found in Qwen Image Max.

The Halloween Invitation

Text-to-Image

“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”

Qwen Image Max
Stable Diffusion 3.5 Medium

AI Judge Analysis

Qwen Image Max

  • + Excellent typography with perfect spelling in all requested areas
  • + Superior composition with a high-quality gothic border and scrolls
  • + Professional aesthetic that looks like a finished graphic design product
  • None notable for the requested prompt

Stable Diffusion 3.5 Medium

  • + Successfully included all elements including trees, bats, and pumpkins
  • + Nice contrast between the parchment and the background elements
  • Numerous spelling errors including Halloweeen and Inviloween
  • Missed the scroll banner element as a distinct graphic
  • The central focus is distorted by text overlays rather than being integrated

Verdict: Qwen Image Max perfectly followed every instruction, including specific text strings and graphic design elements like the scroll banner and thorn border. Stable Diffusion 3.5 Medium failed significantly on text rendering, resulting in multiple typos and a cluttered layout that lacked the requested 'polished' feel.

Isometric Miniature Diorama Scenes

Text-to-Image

“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”

Qwen Image Max
Stable Diffusion 3.5 Medium

AI Judge Analysis

Qwen Image Max

  • + Perfectly followed text instructions including word placement, font weight, and the flag icon.
  • + Excellent miniature 3D aesthetic with a clear glass diorama base as requested.
  • + High-quality textures on the fish, rice, and tamago that feel both realistic and stylized.
  • The perspective is slightly flatter than a true 45-degree angle.

Stable Diffusion 3.5 Medium

  • + Clean, vibrant colors that fit a cartoonish theme.
  • + Good adherence to the isometric perspective and square format.
  • Failed to include the flag icon and the bold text 'JAPAN' is too small.
  • The 3D rendering lacks the 'miniature diorama' base requested.
  • Visual artifacts in the 'SUSHI' text where the letters overlap or have weird transparency.

Verdict: Qwen Image Max followed every specific detail of the prompt, including the complex text layout, flag iconography, and the glass diorama base. Stable Diffusion 3.5 Medium struggled with the text hierarchy, omitted the flag, and the rendering quality of the sushi itself was significantly less detailed than Qwen's output.

Adorable Baby Animals in Sunny Meadow

Text-to-Image

“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”

Qwen Image Max
Stable Diffusion 3.5 Medium

AI Judge Analysis

Qwen Image Max

  • + Successfully included multiple interactions and a complex 'tumbling' dynamic
  • + Features beautiful god rays and a lush, varied floral environment
  • + Maintains high anatomical detail across multiple distinct species
  • Failed to include the baby bunny
  • Includes a fifth animal (a second puppy) not requested in the prompt

Stable Diffusion 3.5 Medium

  • + Excellent vibrant colors and soft, fluffy fur textures
  • + Warm golden hour lighting is very effective
  • + Clean, high-quality rendering of facial features
  • Failed to include both the tabby kitten and the baby bunny
  • Static composition lacks the 'tumbling' and 'chasing' actions requested
  • The central animal is a hybrid-looking creature rather than a clear species from the prompt

Verdict: Qwen Image Max captures the 'joyful wholesome vibe' and the energetic interaction of animals much better, despite missing the bunny and adding an extra dog. Stable Diffusion 3.5 Medium produced a more static, portrait-like image and failed to include half of the requested animal types. Qwen's lighting, complex composition, and adherence to the 'tumbling' action make it the superior choice.

Vintage Cafe Logo

Text-to-Image

“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”

Qwen Image Max
Stable Diffusion 3.5 Medium

AI Judge Analysis

Qwen Image Max

  • + Excellent typography including the requested accent mark on 'Caffè'.
  • + Clean vector emblem style with clear cross-hatching and shading.
  • + High prompt adherence for both the cloche dome and the banner.
  • The steam is a bit literal and somewhat thick for a minimalist logo.

Stable Diffusion 3.5 Medium

  • + Elegant warm brown and cream color palette.
  • + Intricate hand-drawn aesthetic that captures a vintage feel.
  • Several spelling errors including 'Florrian' and 'Est 170'.
  • The cloche dome is oddly shaped and lacks the requested steam.
  • The composition is cluttered and fails the 'minimalist' requirement.

Verdict: Qwen Image Max followed every specific instruction in the prompt, resulting in a clean, professional logo with perfect spelling and a clear vector style. Stable Diffusion 3.5 Medium failed on text accuracy and provided a cluttered illustration rather than a minimalist logo.

Next steps

Explore each model

The Max series of Tongyi Qwen’s image generation model excels across a wide range of generation tasks. Compared with the Plus series, it significantly reduces the “AI-like” feel in generated images, enhancing their realism. It delivers more lifelike material textures for human subjects, finer and more detailed natural textures, and more visually appealing text rendering.