Head to head
Esc

Models · slot A

to navigate to pick

Qwen Image Max Alibaba Stable Diffusion 3.5 Large Stability AI

Settled by community votes across 8 shared challenges, with an AI judge weighing in on each.

Qwen Image Max

21.3 arena score

#35 of 62 in Text-to-Image

Skill signature · Text-to-Image

Stable Diffusion 3.5 Large

22.9 arena score

#29 of 62 in Text-to-Image

Vote tally

Where the votes landed

Qwen Image Max

0%

win rate

Ties

0%

Stable Diffusion 3.5 Large

0%

win rate

Shared challenges 8

Challenge by challenge

The strongest take from each model on every shared challenge, with the AI judge's read.

Geometric Composition

Text-to-Image

“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”

Qwen Image Max
Stable Diffusion 3.5 Large

AI Judge Analysis

Qwen Image Max

  • + Perfectly adheres to the requested spatial arrangement with the book on top.
  • + Excellent rendering of light and reflections within the glass.
  • + High photographic realism and clean composition.
  • The sphere appears to be floating without visible support, though this is a minor artistic choice.

Stable Diffusion 3.5 Large

  • + Strong handling of light and shadow on the table surface.
  • + Good rendering of the glass material's thickness and edges.
  • Failed the spatial positioning request by putting the book inside/under the cube instead of on top.
  • Composition is slightly cluttered with the background furniture.

Verdict: Qwen Image Max followed every spatial instruction perfectly, placing the red book on top of the cube and the blue sphere inside. Stable Diffusion 3.5 Large failed the primary layout challenge by placing the red book inside the cube at the bottom, and the blue sphere resting on the book. Qwen Image Max also delivered a more aesthetically pleasing, professional photographic quality.

Fantasy Warrior

Text-to-Image

“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”

Qwen Image Max
Stable Diffusion 3.5 Large

AI Judge Analysis

Qwen Image Max

  • + Excellent adherence to all prompt details including beads in hair and leather straps
  • + Deeply atmospheric lighting with realistic reflections on the metal
  • + Highly detailed skin textures showing battle-worn features and age
  • The bokeh sparks appear a bit uniform and overlay-like in some areas

Stable Diffusion 3.5 Large

  • + Very intricate engraving on the plate armor
  • + Good shallow depth of field with a cinematic feel
  • + Strong character expression and clear facial features
  • Failed to include the beads in the hair as requested
  • Armor feels slightly flat compared to Model A's lighting and material texture

Verdict: Qwen Image Max followed the prompt much more closely, including the specific request for beads in the hair and providing superior texture on the leather and cloth underlayers. While Stable Diffusion 3.5 Large produced a high-quality cinematic portrait, it missed some of the finer descriptive details and the lighting felt less integrated with the environment than in Qwen's output.

Magic Burger Explosion: Fiery Photorealism Challenge

Text-to-Image

“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”

Qwen Image Max
Stable Diffusion 3.5 Large

AI Judge Analysis

Qwen Image Max

  • + Perfect adherence to text requirements with clean, stylized, and readable fonts.
  • + Excellent sense of motion with the 'exploded' concept where ingredients are flying apart.
  • + Superior photographic quality and lighting on the food items.
  • The burger is tilted at a slightly aggressive angle that some might find less appetizing.

Stable Diffusion 3.5 Large

  • + Impressive fire effects and embers interacting with the burger components.
  • + Good photorealism on the textures of the meat and bun.
  • Completely failed to include any of the requested text.
  • The burger is 'floating' but not 'exploded'; the layers are still mostly stacked.
  • The composition is more of a generic food shot rather than a dynamic advertisement.

Verdict: Qwen Image Max followed every instruction including complex text rendering and the 'exploded' burger layout, resulting in a professional-looking advertisement. Stable Diffusion 3.5 Large produced a high-quality image of a burger over fire, but it failed to include any text and did not adhere to the dynamic 'exploded' motion requested in the prompt.

The Capybara Taxi Driver

Text-to-Image

“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”

Qwen Image Max
Stable Diffusion 3.5 Large

AI Judge Analysis

Qwen Image Max

  • + Excellent adherence to the full prompt including the businesswoman in the back seat.
  • + Highly realistic taxi interior with accurate lighting and detailed car textures.
  • + Natural integration of the capybara into a human-like pose with hands correctly placed on the wheel.
  • The hands on the steering wheel appear more like primate hands than capybara paws.
  • Minor roof lining artifacts in the upper center of the image.

Stable Diffusion 3.5 Large

  • + Strong vibrant colors and high contrast aesthetic.
  • + Clear capybara facial features with nice whisker detail.
  • Fails to include the requested passenger in the back seat.
  • The capybara's paws are not on the steering wheel as requested, appearing to hover or float instead.
  • Anatomical issues with the capybara's legs and the way they merge into the seat.

Verdict: Qwen Image Max followed every instruction in the prompt, including the complex addition of the bored businesswoman in the background, whereas Stable Diffusion 3.5 Large missed the passenger entirely. Qwen Image Max also delivered a more coherent and realistic composition with the capybara actually interacting with the vehicle controls.

The Halloween Invitation

Text-to-Image

“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”

Qwen Image Max
Stable Diffusion 3.5 Large

AI Judge Analysis

Qwen Image Max

  • + Perfect text rendering for all lines of requested information.
  • + Excellent composition with a central glowing jack-o-lantern as requested.
  • + Highly detailed gothic border featuring thorns and webs exactly as described.
  • The parchment texture is a bit flat compared to the jagged edge of the other model.

Stable Diffusion 3.5 Large

  • + Atmospheric lighting with a large moon and twisted tree silhouettes.
  • + Great aged parchment paper effect with burned/torn edges.
  • Failed to include the specific event details (Date, Time, Location) at the bottom.
  • The jack-o-lantern is not the central focus as requested by the prompt.
  • Graphic artifacts and messy text/symbols appear on the scroll banner.

Verdict: Qwen Image Max followed every specific instruction, including the complex request for multiple lines of exact text and specific border elements. Stable Diffusion 3.5 Large produced a moody image but failed to include the event details and suffered from muddy text rendering in the scroll area.

Isometric Miniature Diorama Scenes

Text-to-Image

“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”

Qwen Image Max
Stable Diffusion 3.5 Large

AI Judge Analysis

Qwen Image Max

  • + Excellent typography rendering with clean, clear text placement
  • + Highly accurate 3D diorama/miniature aesthetic with a professional glassy base
  • + Perfect adherence to the 'solid light blue background' and 'square format' constraints
  • The sushi models are slightly more simplified/generic compared to realistic textures

Stable Diffusion 3.5 Large

  • + High textural detail in the sushi rice and fish
  • + Complex isometric composition with many individual elements
  • Failed to place text at 'top-center', instead putting it on a small card
  • Background has a textured pattern instead of being 'solid light blue'
  • Composition is cluttered with extra elements like chopsticks and side bowls, ignoring the 'minimal garnish' request

Verdict: Qwen Image Max followed every specific detail of the prompt, particularly the challenging typography and diorama-base requirements, resulting in a cleaner and more professional graphic. Stable Diffusion 3.5 Large produced high-quality textures but failed significantly on the layout instructions, misplacing the text and adding excessive background noise.

Adorable Baby Animals in Sunny Meadow

Text-to-Image

“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”

Qwen Image Max
Stable Diffusion 3.5 Large

AI Judge Analysis

Qwen Image Max

  • + Excellent fur texture and sharpness across all animals.
  • + Vibrant color palette with a high density of flowers.
  • + Complex interaction between animals (tumbling/playing) is well depicted.
  • Completely failed to include the baby bunny requested in the prompt.
  • Included two golden retriever puppies instead of one.
  • Butterflies feel a bit static and superimposed.

Stable Diffusion 3.5 Large

  • + Correctly included all four requested animals: puppy, kitten, bunny, and fox.
  • + Captured the 'chasing' aspect of the prompt with dynamic movement.
  • + Beautiful use of bokeh and soft sunrise lighting for a dreamy atmosphere.
  • Anatomical issues with the animals' paws and legs during the running motion.
  • The fox kit has slightly distorted facial features.
  • The kitten's fur looks a bit less 'ultra-detailed' compared to Model A.

Verdict: Stable Diffusion 3.5 Large is the winner because it successfully followed the complex prompt requirements by including all four specific animals, whereas Qwen Image Max missed the bunny entirely and duplicated the puppy. While Qwen Image Max had superior sharpness and fur detail, Stable Diffusion 3.5 Large better captured the joyful action and specific character list requested.

Vintage Cafe Logo

Text-to-Image

“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”

Qwen Image Max
Stable Diffusion 3.5 Large

AI Judge Analysis

Qwen Image Max

  • + Excellent typography with perfect spelling and accent mark.
  • + Professional engraving style with consistent shading.
  • + Strong execution of the 'Est. 1720' banner as requested.
  • The cloche is closed, meaning the steam appears to be coming through the metal instead of from food underneath.

Stable Diffusion 3.5 Large

  • + Includes decorative corner flourishes that enhance the vintage aesthetic.
  • + Creative interpretation of the cloche showing steam from the contents below.
  • Spelling error in the main text ('Cafféé' instead of 'Caffè').
  • The cloche handle and steam shapes look slightly lopsided and amateur.
  • The 'Est. 1720' is not on the banner as requested; the banner contains the name instead.

Verdict: Qwen Image Max produced a much more professional and polished logo with flawless typography and better adherence to the specific 'Est. 1720' banner placement. Stable Diffusion 3.5 Large struggled with spelling and failed to follow the specific layout instructions for the banner.

Next steps

Explore each model

The Max series of Tongyi Qwen’s image generation model excels across a wide range of generation tasks. Compared with the Plus series, it significantly reduces the “AI-like” feel in generated images, enhancing their realism. It delivers more lifelike material textures for human subjects, finer and more detailed natural textures, and more visually appealing text rendering.