Alibaba's Qwen image model
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
Qwen Image
#35 of 62 in Text-to-Image
Stable Diffusion 3.5 Medium
#56 of 62 in Text-to-Image
Where the votes landed
Qwen Image
0%
win rate
Ties
0%
Stable Diffusion 3.5 Medium
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
Qwen Image
- + Excellent photographic quality with realistic depth of field.
- + Accurate spatial arrangement of the ball inside and book on top.
- + Beautiful rendering of light and reflections on the glass and wood surfaces.
- − The plant is behind the cube but doesn't show much distortion through the glass as requested.
Stable Diffusion 3.5 Medium
- + The blue sphere has a nice 'glass marble' texture.
- + The plant is clearly visible through the glass cube.
- − The sphere is floating unnaturally in the center without support.
- − Physical artifacts are present, such as the book's edge merging into the glass.
- − Lower overall image resolution and visual fidelity compared to the competitor.
Verdict: Qwen Image provides a much more polished and realistic photograph with correct physics, whereas Stable Diffusion 3.5 Medium struggles with the physical interaction of the objects, resulting in a floating sphere and messy edges where the book meets the glass. Qwen Image also handles the soft window lighting with much more subtlety and realism.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
Qwen Image
- + Excellent anatomical realism in the man's face and hands.
- + Smooth, natural-looking rain and reflections.
- + Strong adherence to the 50mm lens and shallow depth of field request.
- − Lack of motion blur on the passing cars as requested.
- − The bicycle geometry is slightly warped near the pedals.
Stable Diffusion 3.5 Medium
- + Captures a more 'candid' and cinematic street atmosphere.
- + Better application of the 'imperfect framing' requested in the prompt.
- + Included a subtle hint of motion blur with the vehicles.
- − The man's hands are severely malformed and blending into the bicycle.
- − The red bicycle frame has nonsensical structural connections.
- − Lower image clarity and more muddy textures compared to the other model.
Verdict: Qwen Image is the superior model because it maintains high anatomical and structural integrity, delivering a realistic face and hands which Stable Diffusion 3.5 Medium fails to do. While Stable Diffusion 3.5 Medium better captured the requested 'imperfect framing' and cinematic mood, its significant technical failures in rendering the subject's body and the bicycle make it a less successful image overall.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
Qwen Image
- + Excellent adherence to the 'beads' in hair requirement
- + Masterful use of warm torchlight and dynamic spark effects
- + Detailed engraving on armor and textured leather straps
- − The scars look a bit like surface paint or fresh bloody cuts rather than healed tissue
- − The fire source on the right has some distracting digital artifacts
Stable Diffusion 3.5 Medium
- + Extremely realistic skin texture and lifelike eyes
- + Braids are well-integrated and realistic
- + Sophisticated engraving on the breastplate
- − Failed to include the beads in the hair as requested
- − The torchlight effect is much more subtle and lacks the specified warmth
- − The bokeh sparks appear as generic dots rather than embers
Verdict: Qwen Image followed the specific prompt details much better, particularly regarding the beads in the hair and the impact of the warm torchlight. While Stable Diffusion 3.5 Medium achieved higher realism in the facial skin texture and eyes, it missed several key descriptive elements that give the scene its atmospheric character.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
Qwen Image
- + Excellent adherence to the 'modern minimalist' aesthetic with clean layouts.
- + Successful rendering of bold sans-serif headlines and organized columns.
- + Vibrant food photography that integrates well with the background colors.
- − Internal category text remains gibberish despite the header being clear.
- − Combined pizza and mains into one section rather than separate sections as requested.
Stable Diffusion 3.5 Medium
- + Includes a high density of food images as requested in the grid prompt.
- + Strong variety in the food photos presented.
- − The layout is cluttered and lacks the 'modern minimalist' feel requested.
- − Font choice is stylized and less legible, failing the 'bold sans-serif' requirement.
- − Visual artifacts and blurry rendering in the text and grid lines.
Verdict: Qwen Image perfectly captures the professional, minimalist aesthetic of a modern restaurant menu with clear hierarchy and clean design lines. Stable Diffusion 3.5 Medium provides more varied food imagery but suffers from a cluttered layout and low legibility that contradicts the 'minimalist' prompt requirement.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
Qwen Image
- + Excellent typography with a glowing neon effect that matches the prompt perfectly.
- + Dynamic composition with components flying around, fulfilling the 'exploded' aspect of the prompt.
- + Vibrant lighting and high-quality rendering of textures like the patty and fresh vegetables.
- − The main burger body isn't fully 'exploded' or separated enough into distinct layers.
- − Slightly cartoonish feel in the flying sauce droplets compared to the realistic burger.
Stable Diffusion 3.5 Medium
- + Strong 'fiery background' with realistic flame integration and glowing embers.
- + The text rendering for the price and title is clear and formatted correctly as requested.
- + Photorealistic texture on the bun and patty.
- − Completely failed the 'exploded burger' and 'suspended in mid-air' components portion of the prompt.
- − The 'starburst' for the price is more of a line-art icon than a integrated glowing element.
- − The text is flat and lacks the 'fiery, glowing effect' requested compared to Model A.
Verdict: Qwen Image followed the complex layout instructions much better, delivering a dynamic exploded view and highly effective glowing typography. Stable Diffusion 3.5 Medium failed to separate the burger components despite the request for an exploded view, resulting in a static image that missed the core creative intent.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
Qwen Image
- + Excellent legibility with almost perfect spelling of complex menu items.
- + Consistent and realistic chalk font that matches the 'handwritten' request.
- + Accurate rendering of prices and date formatting.
- − The year is incorrectly rendered as '20026' instead of '2026'.
- − The text looks slightly too clean and digital, lacking some of the requested chalk 'smudge' realism.
Stable Diffusion 3.5 Medium
- + Beautiful chalk texture with realistic smudges and varied pressure.
- + Creative layout that feels authentic to a café atmosphere.
- − Poor text rendering with severe misspellings and 'gibberish' words.
- − Failed to include the specific year '2026' accurately.
- − Messy composition that makes it hard to read the actual menu items.
Verdict: Qwen Image followed the prompt's text requirements with high accuracy, producing a clear and readable menu, despite a minor error in the year. Stable Diffusion 3.5 Medium captured the artistic texture of chalk much better, but failed significantly on legibility and prompt adherence regarding specific text. Qwen Image is the preferred choice for a functional text-based request.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
Qwen Image
- + High resolution and smooth cinematic lighting
- + Clean, realistic space environment textures
- − Completely failed the negative constraint to put the horse on top of the astronaut
- − Standard, un-surreal interpretation of the prompt
Stable Diffusion 3.5 Medium
- + Attempted a more unusual composition with the horse's legs appearing to emerge through the astronaut
- + Higher contrast starfield
- − Failed the specific spatial instruction for the horse to be on top
- − Significant anatomical artifacts where the horse legs meet the clouds
- − Overall messy composition and low clarity
Verdict: Both models failed the specific logic puzzle in the prompt, which requested the horse to be on top of the astronaut. Qwen Image produced a much higher quality, aesthetically pleasing image, whereas SD 3.5 Medium produced a distorted image with poor leg anatomy and confusing object placement.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
Qwen Image
- + Excellent photorealism in the car interior and lighting.
- + Correctly followed the instruction for the passenger to be looking at her phone.
- + The capybara's anatomy and fur texture are highly realistic, including the requested paws on the wheel.
- − The hands on the steering wheel look more like primate hands than capybara paws.
- − The taxi sign on top of the car is misspelled as 'YOXI'.
Stable Diffusion 3.5 Medium
- + Creative POV looking through the front windshield.
- + Distinct capybara claws/paws visible on the wheel.
- + Good depth of field with the city lights.
- − Failed the negative constraint: the passenger is looking forward at the camera instead of her phone.
- − The car anatomy is a bit confusing with a steering wheel that looks like a flat dashboard.
- − Low facial detail on the human passenger.
Verdict: Qwen Image is the superior model as it correctly followed all complex prompt instructions, including the specific behavior of the passenger on her phone. While Stable Diffusion 3.5 Medium has a unique perspective, it failed to include the phone and the passenger's expression as requested, and Qwen's overall lighting and texture quality appear more lifelike.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
Qwen Image
- + Excellent layout following all prompt instructions including the scroll banner.
- + High text legibility for specific details like the date and location.
- + Cinematic lighting and high visual quality in the central jack-o-lantern.
- − One minor spelling error in the title text ('Halle Party').
- − The border is a bit heavy on thorns, slightly obscuring the inner artwork.
Stable Diffusion 3.5 Medium
- + Successfully captured the twisted trees and moody sky background.
- + Vibrant jack-o-lantern designs on both sides of the poster.
- − Numerous spelling errors in almost all text fields ('Halloweeen', 'Timme', 'Loccation', 'Aches').
- − Failed to include a central jack-o-lantern, placing them on the sides instead.
- − Did not include the requested scroll banner.
Verdict: Qwen Image is the clear winner because it adheres to the specific layout requirements, including the scroll banner and correct event details. While Stable Diffusion 3.5 Medium captures the aesthetic well, its text is riddled with misspellings and it failed to follow the instruction for a 'central' pumpkin.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
Qwen Image
- + Excellent typography rendering with clean, accurate text
- + Perfectly captures the isometric miniature diorama aesthetic
- + High-quality 3D clay-like textures and soft lighting
- − The flag icon beside the text is slightly distorted compared to the larger flag
Stable Diffusion 3.5 Medium
- + Realistic textures on the sushi roe and seaweed
- + Accurate solid blue background and square format
- − Failed to include 'Japan' in large bold text (it is small and off-center)
- − Does not include the requested diorama base or flag icon
- − The word 'SUSHI' has a strange dot artifact above the 'I'
Verdict: Qwen Image significantly outperformed Stable Diffusion 3.5 Medium by following every aspect of the prompt, including the complex text layout, the diorama base, and the specific isometric style. Stable Diffusion 3.5 Medium failed on most composition requirements, missing the large text hierarchy and the miniature scene elements.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
Qwen Image
- + Successfully included all four requested animals (dog, cat, bunny, fox).
- + Excellent rendering of translucent 'god rays' and dew sparkles.
- + Anatomy is consistent and the interaction feels more dynamic with paws in motion.
- − The fox's face leans slightly more towards a cartoonish aesthetic than a 'hyper-photorealistic' one.
Stable Diffusion 3.5 Medium
- + Very vibrant color palette with high-contrast golden lighting.
- + Good texture on the fur of the golden retriever.
- − Failed to include the bunny as requested in the prompt.
- − The kitten's anatomy is distorted with oversized, fox-like ears.
- − The layout is more static and less representative of the 'tumbling together' action.
Verdict: Qwen Image followed the prompt exactly, including all four specific animals and capturing the 'tumbling' motion well. Stable Diffusion 3.5 Medium failed to generate the baby bunny and produced a kitten with anatomical issues that made it look too similar to the fox kit. Qwen Image's lighting and particle effects (dew sparkles) also felt more integrated into the scene.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
Qwen Image
- + Clean minimalist vector aesthetic.
- + Accurate text and date rendering.
- + Perfectly captures the requested cloche dome with steam.
- − Layout of the word 'Florian' is slightly awkward and cluttered.
- − Steam element is very simple, almost cartoonish.
Stable Diffusion 3.5 Medium
- + Elegant hand-drawn vintage engraving style.
- + Warm monochromatic color palette is very appealing.
- + Sophisticated composition with ornate banner work.
- − Spelling errors in 'Florrian' and 'Est 170'.
- − Failed to include the specific year '1720'.
- − Less minimalist than requested.
Verdict: Qwen Image followed the specific instructions for text and date perfectly, delivering a clean vector logo that is ready for use, though the typography layout is a bit cramped. Stable Diffusion 3.5 Medium produced a much more beautiful and artistic illustration, but failed significantly on the text accuracy and the specific date requested. Qwen Image is the winner for its functional adherence to the prompt requirements.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
Qwen Image
- + Successfully captures a clean vector infographic style with logical spatial flow.
- + Accurately represents the crew members names and distinct steps requested.
- + The color palette perfectly adheres to the NASA-inspired theme.
- − Includes instructional text '('Stop at landing')' and prompt typos ('Wicon') as literal text in the image.
- − Minor spelling errors in text like 'ApolL' and 'Sarth Orbit'.
Stable Diffusion 3.5 Medium
- + Appropriately uses a deep space navy and muted red palette.
- + Clean line work for the orbital trajectories.
- − Fails to provide distinct icons for all 6 requested steps.
- − Text is mostly illegible gibberish and does not follow the requested labels.
- − Layout is cluttered and lacks the clarity of an infographic.
Verdict: Qwen Image is the clear winner as it successfully interprets the prompt as an infographic, providing a coherent sequence of events with recognizable icons and legible (though slightly flawed) text. Stable Diffusion 3.5 Medium produces a more abstract image with nonsensical text and fails to follow the specific '6 steps' structure requested in the prompt.
Explore each model
Stability AI's 2.5-billion parameter Multimodal Diffusion Transformer with improvements (MMDiT-X) text-to-image model optimized for consumer hardware, featuring improved image quality, typography, and complex prompt understanding