Alibaba's Qwen Image 2.0 model with enhanced text rendering, supporting both Chinese and English prompts with up to 6 images per request
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
Qwen Image 2.0
#34 of 62 in Text-to-Image
Stable Diffusion 3.5 Large
#28 of 62 in Text-to-Image
Where the votes landed
Qwen Image 2.0
0%
win rate
Ties
0%
Stable Diffusion 3.5 Large
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
Qwen Image 2.0
- + Perfect adherence to spatial instructions with the book on top and sphere inside
- + Realistic glass reflections and light refraction
- + Excellent soft window lighting from the left as requested
- − The sphere appears to be floating without a visible support structure
- − The glass cube has internal partitions that make it look like multiple boxes merged
Stable Diffusion 3.5 Large
- + Very high clarity and realistic materials
- + Accurately depicts a green plant partially visible through the glass
- − Failed the spatial prompt: the red book is inside the cube rather than on top of it
- − The blue sphere is on top of the book, which was not the requested arrangement
Verdict: Qwen Image 2.0 followed all spatial instructions perfectly, placing the sphere inside the cube and the book on top. Stable Diffusion 3.5 Large produced a high-quality image but failed the core prompt logic by putting the book inside and the sphere on top of the book.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent skin texture and facial details that feel authentic and non-AI.
- + The shallow depth of field and color grading perfectly match the 'cinematic' and 'no stylization' requirements.
- + The composition feels like a genuine candid street photo with a very natural crop.
- − The car in the background lacks the requested motion blur.
- − Minor anatomical confusion where the hands interact with the bicycle chain.
Stable Diffusion 3.5 Large
- + Successfully captures the light rain effect and wet pavement reflections.
- + Accurately represents the request for a red bicycle and an elderly man.
- + Good use of vertical space in the composition.
- − The image has an obvious 'AI-generated' smooth aesthetic that ignores the 'no stylization' and 'natural skin' request.
- − The man's hands are mangled and fused into the bicycle handlebars.
- − The background car lacks motion blur and looks Static.
Verdict: Qwen Image 2.0 is the clear winner as it adheres to the technical photographic requirements, particularly the natural skin texture and cinematic-yet-realistic lighting. Stable Diffusion 3.5 Large suffers from significant anatomical distortions in the hands and a plastic-like texture that fails the 'no stylization' prompt.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent depiction of beads within the braided hair.
- + Highly detailed textures on the leather and pitted metal armor.
- + Superior lighting effects with warm reflections and bokeh sparks.
- − The eyes appear slightly glowing or unnatural rather than just lifelike.
- − Composition is a bit crowded with the hand placement.
Stable Diffusion 3.5 Large
- + Incredible level of fine detail in the ornate armor engravings.
- + Very lifelike eyes and realistic facial anatomy.
- + Excellent shallow depth of field with a distinct background.
- − Failed to include the requested beads in the braids.
- − The lighting feels a bit cooler and less 'torchlit' than requested.
Verdict: Qwen Image 2.0 followed the specific prompt details much better, successfully incorporating the beads in the hair and the warm torchlight atmosphere. While Stable Diffusion 3.5 Large produced stunningly intricate armor engravings and a very realistic face, it missed the bead requirement and lacked the warm, spark-filled ambiance of the other model.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent photographic quality and realism in the food imagery.
- + Clean, logical grid layout that perfectly matches the 'minimalist' request.
- + Accurate categorization into Appetizers, Pizza, and Mains.
- − Nonsense text for dish names and prices.
- − Formatting of some text characters is glitchy and inconsistent.
Stable Diffusion 3.5 Large
- + Interesting use of a side-scrolling floral/food border for a creative layout.
- + Large, bold typography for the main header.
- − Extremely poor text rendering with significant gibberish and artifacts.
- − Layout is cluttered and does not feel 'minimalist' or professional for a menu.
- − Categorization labels are misspelled and confusing.
Verdict: Qwen Image 2.0 provides a much more professional and usable design that adheres strictly to the grid and category requirements. While its text is nonsensical, the visual quality of the food and the cleanliness of the layout far surpass Stable Diffusion 3.5 Large, which produced a cluttered design with severe text distortions.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI judge analysis unavailable for this challenge.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent text rendering with perfect spelling and realistic chalk texture.
- + Authentic handwriting style that matches the prompt's request for natural variations.
- + High-quality composition with a realistic cafe background in soft focus.
- − The 'Today's Specials' title is in a clean print-cursive hybrid rather than elaborate elegant cursive.
Stable Diffusion 3.5 Large
- + Successfully captures a wider view of a cozy cafe environment.
- + Good use of chalk-style decorative borders around menu sections.
- − Numerous spelling errors including 'TODAAY' and 'Muglrrom'.
- − Incorrect date (2024 instead of 2026).
- − The text is very messy and difficult to read compared to the other model.
Verdict: Qwen Image 2.0 followed the prompt instructions near-perfectly, delivering clear, legible, and correctly spelled text with a convincing chalk texture. Stable Diffusion 3.5 Large struggled significantly with text accuracy, including spelling mistakes in the header and the body of the menu, and failed to follow the specific date requested.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent clarity and vibrant cinematic lighting.
- + High level of detail in the horse's anatomy and the astronaut's suit.
- + Creative addition of floating water droplets that enhance the surreal theme.
- − The horse's skin has an unusual scale-like texture near the neck that looks a bit digital.
- − The composition feels a bit static compared to the movement in the other model.
Stable Diffusion 3.5 Large
- + Stronger sense of motion and scale with the star dust and planetary background.
- + Better integration of the subjects into the environment through atmospheric lighting.
- + Atmospheric 'fog' and stardust create a more dreamy, surreal aesthetic.
- − Lower clarity on the astronaut's face and suit details compared to Qwen.
- − The horse's anatomy, specifically the hind legs, is somewhat obscured by the 'dust' effect.
Verdict: Both models failed the negative constraint of placing the 'horse on top' of the astronaut, instead opting for the standard astronaut riding a horse. Qwen Image 2.0 provides a sharper, more detailed image with high-contrast lighting, while Stable Diffusion 3.5 Large offers a superior sense of scale and atmosphere that feels more cinematic and surreal. Stable Diffusion 3.5 Large is the winner as it better captures the 'surreal' and 'cinematic' keywords of the prompt.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent adherence to the full prompt including the passenger.
- + Realistic textures for the capybara's fur and the leather steering wheel.
- + Natural lighting and composition that creates a cohesive cinematic scene.
- − The capybara's paws look slightly more like humanoid hands with fur.
Stable Diffusion 3.5 Large
- + High resolution and sharp details on the capybara's face.
- + Good clothing textures on the jacket.
- − Completely missed the passenger character requested in the prompt.
- − The perspective is awkward, with the capybara sitting too low and far back for the steering wheel.
- − The capybara has human-proportioned legs and hands extending from its body.
Verdict: Qwen Image 2.0 followed the complex prompt instructions much better, including both the capybara driver and the bored businesswoman passenger in the background. Stable Diffusion 3.5 Large completely omitted the passenger and had significant anatomical issues where the capybara appeared to have human legs and hands.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent text rendering with no spelling errors in any of the required fields.
- + Superior adherence to all prompt details, including the specific date and location.
- + Balanced gothic composition that feels like a cohesive invitation.
- − The jack-o-lantern is placed over the background trees rather than being integrated into the scene's depth.
Stable Diffusion 3.5 Large
- + Good gothic aesthetic with high contrast and a nice parchment texture.
- + Large, atmospheric moon and silhouettes create a strong spooky mood.
- − Failed to include the specific event details like date, time, and location.
- − Text rendering is inconsistent, with some garbled characters at the bottom of the scroll.
- − Composition is slightly cluttered with secondary jack-o-lanterns that distract from the central focus.
Verdict: Qwen Image 2.0 is the clear winner as it successfully rendered every piece of text requested, including the specific date, time, and location, which Stable Diffusion 3.5 Large completely omitted. While both models captured the aesthetic well, Qwen Image 2.0's polished execution of the invitation layout and perfect spelling makes it much more functional for the prompt's requirements.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent text rendering and alignment with the requested top-center position.
- + Clean, professional composition that perfectly follows the 'minimal garnish' instruction.
- + High-clarity materials that look both appetizing and stylized.
- − The style leans more toward realistic product photography than '3D cartoon' or 'miniature diorama'.
Stable Diffusion 3.5 Large
- + Stronger adherence to the 'miniature 3D cartoon' and 'diorama' aesthetic.
- + Includes more detailed sushi variety and a physical miniature flag.
- + Good texture work on the rice and wooden base.
- − Failed to place the text in the background as requested, putting it on a small sign instead.
- − The scene is a bit cluttered, ignoring the 'minimal garnish' instruction.
- − The flag icon is a physical object rather than a graphic element as implied by the prompt structure.
Verdict: Qwen Image 2.0 followed the layout and text instructions perfectly, resulting in a very clean and professional graphic. While Stable Diffusion 3.5 Large captured the 'miniature' feel and '3D cartoon' style more effectively, it failed to place the text correctly and created a much busier composition than requested. Qwen Image 2.0 is the winner for its superior prompt adherence and typography.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
Qwen Image 2.0
- + Successfully included all four requested animals interacting in one cluster.
- + Excellent rendering of lighting and 'god rays' through the trees.
- + Highly detailed fur texture and clear, sharp focus on all characters.
- − The fox kit is a bit squashed at the bottom and has a slightly awkward pose.
Stable Diffusion 3.5 Large
- + Captures a strong sense of joy and movement with the animals running forward.
- + Beautiful bokeh effect and soft, whimsical lighting.
- + The kitten and fox are very distinct and cute.
- − The tabby kitten appears more like a plain ginger kitten, missing specific tabby markings.
- − Anatomy on the puppy's back paw is slightly blurry/indistinct.
Verdict: Qwen Image 2.0 followed the prompt more accurately by showing the animals 'tumbling together' and having distinct breeds like the tabby kitten, while Stable Diffusion 3.5 Large opted for a more traditional 'running towards camera' composition. Qwen's technical execution of the complex lighting and fur details is slightly superior in terms of hyper-realism.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent typography and spelling of the specific name requested.
- + High-quality vector illustration style with clean lines and professional shading.
- + Perfect adherence to the 'Est. 1720' banner requirement.
- − The 'steam' element is integrated into the cloche surface rather than hovering above it.
- − Slightly less 'minimalist' than model B due to the amount of shading and detail.
Stable Diffusion 3.5 Large
- + Successfully captures a minimalist aesthetic with clean silhouettes.
- + Includes subtle paper texture on the background as requested.
- + Creative representation of steam both inside and above the cloche.
- − Incorrect spelling of the brand name adding an extra 'e' to 'Cafféé'.
- − The layout is a bit scattered with too much vertical whitespace between elements.
- − Visual artifacts present in the 'steam' above the cloche where it meets the knob.
Verdict: Qwen Image 2.0 is the clear winner as it produces a professional, ready-to-use vector logo with perfect typography. While Stable Diffusion 3.5 Large follows the minimalist prompt well, it fails on key details like spelling and cohesive layout, resulting in a less polished final product.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent typography with nearly perfect spelling for all required stages.
- + Accurately follows the requested vertical six-step flow with appropriate iconography for each.
- + Clean vector aesthetic with the specific NASA-inspired color palette.
- − Includes a minor typo in 'Translunjar' (contains an extra 'j').
- − Stylization of the lunar module landing icon is a bit bulky compared to the other minimalist icons.
Stable Diffusion 3.5 Large
- + Intricate detail in the planetary illustrations.
- + Good adherence to the requested color palette.
- − Fails to follow the sequential six-step logic requested in the prompt.
- − Poor text rendering with illegible, garbled labels.
- − Inaccurate iconography, showing a Space Shuttle style craft instead of a Saturn V or Apollo Lunar Module.
Verdict: Qwen Image 2.0 successfully interprets the infographic structure, providing a logical sequence of the Apollo 11 mission with readable text and clean icons. In contrast, Stable Diffusion 3.5 Large fails on instruction following, producing a cluttered mess of illegible text and a non-chronological layout that ignores the specific mission steps requested.
Explore each model
Stability AI's 8.1-billion parameter Multimodal Diffusion Transformer (MMDiT) text-to-image model featuring improved image quality, typography, complex prompt understanding, and resource-efficiency