The Max series of Tongyi Qwen’s image generation model excels across a wide range of generation tasks. Compared with the Plus series, it significantly reduces the “AI-like” feel in generated images, enhancing their realism. It delivers more lifelike material textures for human subjects, finer and more detailed natural textures, and more visually appealing text rendering.
Settled by community votes across 8 shared challenges, with an AI judge weighing in on each.
Qwen Image Max
#35 of 62 in Text-to-Image
Stable Diffusion 3.5 Large Turbo
#61 of 62 in Text-to-Image
Where the votes landed
Qwen Image Max
100.0%
win rate
Ties
0.0%
Stable Diffusion 3.5 Large Turbo
0.0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
Qwen Image Max
- + Excellent photo-realism and natural lighting from the left
- + Perfect spatial adherence with the book on top and sphere inside
- + Realistic reflections and refractions within the glass cube
- − The sphere appears to be floating without a visible support
Stable Diffusion 3.5 Large Turbo
- + Sharp, clean rendering with high contrast
- + Distinct textures on the wooden table and plant leaves
- − Failed the spatial positioning of the red book, placing it inside the cube rather than on top
- − The plant is to the left of the cube rather than behind it as requested
- − Minor artifacting on the edge of the glass frame
Verdict: Qwen Image Max followed the spatial instructions perfectly, accurately placing the red book on top of the cube and the plant behind it with realistic lighting. Stable Diffusion 3.5 Large Turbo struggled with the layout, placing the book inside the cube and the plant to the side, resulting in a less accurate interpretation of the prompt.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
Qwen Image Max
- + Excellent adherence to the 'battle-worn' aesthetic with realistic scars, dirt, and facial expressions.
- + Highly detailed engraving on the armor and rich texture on the leather straps and frayed cloth.
- + Superior lighting and atmosphere, with warm torchlight correctly interacting with the metallic surfaces and background sparks.
- − Some of the braided hair strands merge slightly into the armor shoulder pieces.
Stable Diffusion 3.5 Large Turbo
- + Clear, thick braids and recognizable ornate metal engraving on the armor.
- + Good contrast in colors between the green cloth and metallic armor.
- − The character appears too clean and 'model-like' for a battle-worn description, with bloodstains looking more like paint.
- − The lighting is flat and lacks the specified 'warm torchlight' atmosphere.
- − Overall texture is noticeably smoother and less lifelike than requested.
Verdict: Qwen Image Max followed the prompt across every metric, producing a character that feels authentically battle-worn with exceptional textures on the armor, skin, and leather. Stable Diffusion 3.5 Large Turbo produced a much cleaner, more stylized image that missed the lighting atmosphere and the grit requested in the prompt. Qwen's interpretation of 'lifelike eyes' and 'ornate engraved plate' is significantly more detailed and cinematic.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
Qwen Image Max
- + Excellent photorealistic texture on the meat and bun
- + Flawless text rendering for all three requested elements
- + Dynamic composition with a strong sense of motion and embers
- − The burger components are partially assembled rather than fully exploded/suspended individually
Stable Diffusion 3.5 Large Turbo
- + Strong vertical alignment and levitation effect
- + Vibrant colors and glowing fire effect
- − Completely failed to include any of the requested text
- − Rendering style is more illustrational/3D-render than photorealistic
- − Poor anatomy of the burger with messy textures at the bottom
Verdict: Qwen Image Max is the clear winner as it successfully integrated all complex text requirements and followed the stylistic request for photorealism. Stable Diffusion 3.5 Large Turbo failed to include 'MAGIC BURGER', the price, or the secondary message, and produced a less realistic image.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
Qwen Image Max
- + Excellent photorealistic texture on the fur and leather seating
- + Accurate depiction of a businesswoman looking at a phone in the background
- + Captures the gritty, detailed atmosphere of a New York night street through the windows
- − The hands on the steering wheel look more like primate hands than capybara paws
- − The cap is missing a 'Taxi' logo which was implied by the role
Stable Diffusion 3.5 Large Turbo
- + Successfully includes a 'Taxi' logo on the cap
- + Paws on the steering wheel look more morphologically accurate for a capybara
- + Bold, clean lighting and color palette
- − Fails to include the businesswoman in the back seat (she appears to be in the front passenger seat)
- − The businesswoman is not looking at a phone as requested
- − The interior is much less detailed and looks more like a 3D render than a photorealistic image
Verdict: Qwen Image Max is the clear winner as it successfully captures all elements of the prompt, including the specific bored expression of the businesswoman on her phone. Stable Diffusion 3.5 Large Turbo fails on several prompt instructions, positioning the passenger incorrectly and omitting her interaction with a phone.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
Qwen Image Max
- + Excellent typography with perfect adherence to the requested text and date.
- + Highly detailed and cohesive gothic aesthetic with intricate thorn and web borders.
- + Superior cinematic lighting and atmosphere that feels like a polished invitation.
- − None notable; it successfully captured every element of the prompt.
Stable Diffusion 3.5 Large Turbo
- + Bold, clean graphics that fit a simple poster style.
- + Good use of negative space in the center.
- − Failed to include most of the requested text, missing the invitation details entirely.
- − The title is truncated to 'Halloween Party' instead of the full requested title.
- − Lacks the 'vintage parchment' and 'moody night sky' complexity requested.
Verdict: Qwen Image Max successfully followed every instruction, including the specific event details and the scroll banner, while maintaining a high-quality gothic aesthetic. In contrast, Stable Diffusion 3.5 Large Turbo failed to include the event details and used a much simpler, less atmospheric visual style. Qwen's attention to text rendering and composition makes it the clear winner for this specific invitation request.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
Qwen Image Max
- + Excellent text rendering with clean, professional typography and a correct flag icon.
- + Higher quality textures and realistic PBR materials, especially on the salmon and glass base.
- + Perfectly centered composition adhering to the square format requirement.
- − The diorama base is a simple glass plinth rather than a more traditional wooden/miniature set design.
Stable Diffusion 3.5 Large Turbo
- + Good interpretation of the 'diorama' request with the signage and wooden base.
- + Clear 45-degree isometric perspective.
- − Failed the text requirement with a misspelling ('SIIHI') and missing the 'SUSHI' text below 'JAPAN'.
- − The flag icon is incorrect, appearing more like a red and white bicolour flag than the Japanese flag.
- − Overall lower rendering quality with less refined textures compared to the competitor.
Verdict: Qwen Image Max is the clear winner as it followed all prompt instructions, including complex text rendering and specific material requirements. Stable Diffusion 3.5 Large Turbo failed significantly on the text elements and the flag icon, while also producing less realistic textures.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
Qwen Image Max
- + Successfully includes all four requested animals (dog, cat, fox, and bunny).
- + High level of detail in fur textures and environmental elements like god rays and wildflowers.
- + Dynamic and interaction-heavy composition that matches the 'tumbling' prompt.
- − The fox's anatomy is slightly distorted where it interacts with the dog.
- − Includes an extra puppy (two golden retrievers) not explicitly requested.
Stable Diffusion 3.5 Large Turbo
- + Features big expressive eyes as requested in the style description.
- + Clean lighting with backlight highlights on the fur.
- − Fails to include the bunny and the fox kit, missing half the requested animals.
- − The image has a more digital, illustrative feel rather than 'hyper-photorealistic'.
- − The composition is a static portrait rather than animals 'playfully chasing butterflies'.
Verdict: Qwen Image Max followed the complex prompt much better, including all the requested animal species and capturing the energetic 'tumbling' action in a detailed meadow. Stable Diffusion 3.5 Large Turbo failed to include several requested elements—most notably the bunny and fox—and produced a more generic, simplified composition.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
Qwen Image Max
- + Perfect text rendering with the correct 'é' accent
- + Exemplary adherence to the 'cloche dome' description
- + Clean, balanced vector-style composition
- − The steam effect is slightly more whimsical than minimalist
Stable Diffusion 3.5 Large Turbo
- + Strong vintage aesthetic with high contrast
- + Good use of cross-hatching for texture
- − Spelling error in the name ('Caffeé')
- − Garbled text/artifacts on the end of the name banner
- − The illustration is a confusing hybrid of a cloche and a coffee mug
Verdict: Qwen Image Max follows all prompt instructions perfectly, including accurate spelling and a clear rendition of a cloche dome. Stable Diffusion 3.5 Large Turbo struggles with the spelling of the requested entity and provides a confusing visual metaphor that combines a cloche with a mug, resulting in a less professional logo.
Explore each model
Distilled version of SD 3.5 Large that generates high-quality images in just 4 steps, offering faster inference and reduced costs