Alibaba's Qwen image model
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
Qwen Image
#35 of 62 in Text-to-Image
Stable Diffusion 3.5 Large
#30 of 62 in Text-to-Image
Where the votes landed
Qwen Image
0%
win rate
Ties
0%
Stable Diffusion 3.5 Large
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
Qwen Image
- + Perfect adherence to all spatial relationships in the prompt
- + Clean, attractive photorealistic aesthetic
- + Exceptional handling of reflections and transparency in the glass
- − The plant is slightly less 'behind' the cube than beside it, though still visible through the glass
Stable Diffusion 3.5 Large
- + Highly realistic textures on the glass, including dust and fingerprints
- + Correct color palette for all objects
- − Failed to place the red book on top of the cube
- − Confused the spatial relationship between the sphere and the book, placing the sphere on the book
- − The plant is mostly above/behind the scene rather than viewed through the cube
Verdict: Qwen Image followed every instruction in the prompt perfectly, including the specific spatial requirement of placing the red book on top of the glass cube. Stable Diffusion 3.5 Large failed these spatial instructions, placing the book at the bottom and the sphere on top of the book, despite its high level of photographic texture detail.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
Qwen Image
- + Natural and realistic skin texture without over-sharpening
- + Excellent rendering of light reflections and a wet pavement look
- + Composition feels more candid and aligns with the 50mm lens look
- − Anatomical issues with the man's hands and fingers
- − Structural issues with the bicycle pedals and frame
- − Missing the requested motion blur on the passing cars
Stable Diffusion 3.5 Large
- + Stronger adherence to the 'light rain' prompt with visible rain droplets
- + The man's skin and hair texture look more weathered and realistic for his age
- + Includes a more complex urban background
- − Physical interaction between the person and the bicycle is messy/blended
- − The car in the background lacks the requested motion blur
- − The 'imperfect framing' requested looks a bit too intentional and slightly off-balance
Verdict: Qwen Image delivers a cleaner, more photographic aesthetic that captures the atmospheric reflections very well, though it suffers from significant anatomical errors in the hands. Stable Diffusion 3.5 Large provides a much gritier, more detailed subject with better environmental effects (visible rain), but the image feels slightly more 'AI-rendered' in its coloring. Qwen's interpretation of a 'candid street photo' is more successful overall.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
Qwen Image
- + Excellent execution of warm torchlight and bokeh sparks
- + Highly detailed and readable textures on leather straps and metal engravings
- + Strong adherence to the 'beads' in the hair prompt
- − The scars look like surface-level digital slashes rather than healed tissue
- − The torch flame has a slightly artificial, 'sparkler' look
Stable Diffusion 3.5 Large
- + Extremely realistic skin texture and authentically faint scars
- + Intricate armor engraving and highly detailed mail/cloth underlayers
- + More naturalistic lighting and cinematic depth of field
- − Missed the requirement for small beads in the braided hair
- − The torchlight effect is less pronounced compared to the other model
Verdict: Qwen Image followed the specific details of the prompt more closely, notably including the beads and the specific torchlight effects requested. However, Stable Diffusion 3.5 Large produced a much more lifelike and high-fidelity image with superior skin textures and armor realism, despite missing the hair beads. Qwen is preferred for exact prompt adherence, while Stable Diffusion is preferred for overall photographic quality.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
Qwen Image
- + Successfully combines images and text categories onto a single page.
- + High contrast and vibrant colors in the food grid enhance the 'casual dining' feel.
- + Font choice is bold and minimalist as requested.
- − Text is largely gibberish despite having high-quality rendering.
- − Pizza section and Mains section are merged into one oddly named header.
Stable Diffusion 3.5 Large
- + Excellent food photography with realistic textures and lighting.
- + Clean and elegant typography that feels like a professional menu template.
- + Clearer structural separation for menu items.
- − The layout is cropped poorly, making it look like a website banner rather than a full menu page.
- − Fails the specific 'white background with food photos in grid' prompt by placing photos in columns on the sides.
- − Frequent spelling errors in category headers.
Verdict: Qwen Image delivers the most accurate interpretation of a physical menu layout, successfully integrating a grid of photos with pricing text on a clean white background. While Stable Diffusion 3.5 Large has superior image quality for the food itself, its composition feels fragmented and fails to provide a cohesive page layout.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
Qwen Image
- + Excellent typography with clean, readable text and perfect spelling.
- + Strong adherence to the 'exploded' request with components flying around.
- + Well-integrated UI elements like the starburst price tag.
- − The burger itself is slightly less exploded than a true vertical deconstruction.
- − Photorealism on the flying cucumber is a bit weak.
Stable Diffusion 3.5 Large
- + High level of photorealistic texture on the meat and bun.
- + Great atmosphere and lighting with the fire interacting with the burger.
- − Completely failed to include any of the requested text.
- − Failed the 'exploded' instruction as the burger is fully assembled.
- − Overall failed the 'ad' layout request in favor of a standard close-up.
Verdict: Qwen Image is the clear winner as it followed every instruction, including complex text rendering, a starburst price tag, and the specific 'exploded' layout. Stable Diffusion 3.5 Large produced a high-quality, realistic image of a burger, but it missed all of the text and promotional design elements requested in the prompt.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
Qwen Image
- + Excellent text legibility and spelling accuracy, correctly rendering complex names like 'Truffle Mushroom Risotto'.
- + Superior chalk texture with realistic smudges and strokes consistent with a hand-drawn board.
- + Follows the layout instructions perfectly by listing the specific items and prices as requested.
- − The year in the date is rendered as '20026' instead of '2026'.
- − Missing the '$' sign on the third item's price compared to the first two.
Stable Diffusion 3.5 Large
- + Creates a very pleasant and wide café environment with good depth of field and lighting.
- + The handwriting appears small and authentic for a large wall-mounted board.
- − Poor text accuracy with numerous spelling errors like 'TODAAY' and 'Cholcalte'.
- − Failed to follow the date instruction, rendering '2024' instead of '2026'.
- − The layout contains many more boxes and items than the three specifically requested in the prompt.
Verdict: Qwen Image is the clear winner due to its exceptional text rendering capabilities and adherence to the specific menu items requested. While it made a minor error with the year '20026', Stable Diffusion 3.5 Large struggled significantly with spelling, layout, and fundamental prompt details like the specific year and requested menu items.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
Qwen Image
- + Clean, sharp image resolution with no noise
- + Accurate depiction of an astronaut suit and horse anatomy
- − Failed to provide the 'surreal' style requested, leaning more toward a simple composite
- − The lighting on the horse and astronaut does not quite match the background light source
Stable Diffusion 3.5 Large
- + Excellent interpretation of the 'surreal' and 'cinematic' keywords
- + Dynamic composition with stardust and nebulas enhancing the theme
- + Consistent lighting and texture across the entire scene
- − Higher amount of grain/noise throughout the image
- − A few anatomical oddities in the horse's rear leg structure under the stardust
Verdict: Both models failed the negative constraint to have the 'horse on top' of the astronaut, both depicting the astronaut riding the horse. However, Stable Diffusion 3.5 Large captured the requested surreal and cinematic atmosphere much more effectively than Qwen Image, which felt like a static collage by comparison.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
Qwen Image
- + Excellent adherence to the passenger prompt including her bored expression and phone usage.
- + Features a clearly recognizable and professional taxi driver cap.
- + Hands/paws are placed reasonably well on the steering wheel.
Stable Diffusion 3.5 Large
- + Very high detail on the capybara's fur and whiskers.
- + Dynamic lighting with vibrant bokeh in the background.
- − Completely failed to include the human passenger in the back seat.
- − The cap is a baseball cap rather than a traditional taxi driver cap.
- − The capybara's paws are not placed on the steering wheel correctly.
Verdict: Qwen Image followed all instructions, successfully including the bored businesswoman in the back seat and capturing the specific 'taxi driver' aesthetic. Stable Diffusion 3.5 Large failed to include the passenger entirely, resulting in a composition that missed half of the prompt's requirements.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
Qwen Image
- + Excellent adherence to the requested event details including date and location
- + Highly accurate rendering of the complex gothic title text
- + Beautiful composition with a central glowing jack-o-lantern and clean border work
- − Minor spelling error at the very top of the image in the small decorative text
- − The gothic font for the title has a slight overlap in the 'll' area
Stable Diffusion 3.5 Large
- + Strong vintage parchment aesthetic and intricate border detail
- + Dynamic lighting with the large glowing moon
- + Correct banner text placement and content
- − Failed to include the specific event details (Date, Time, Location) requested
- − The font choice for 'Halloween Party' is less 'elegant gothic' and more modern/clean
- − Image aspect ratio is slightly taller than the requested square format
Verdict: Qwen Image is the superior choice because it followed all instructions, particularly the rendering of specific event details like the date and location, which Stable Diffusion 3.5 Large omitted entirely. Qwen Image also achieved a much better gothic aesthetic for the typography while maintaining a clean, professional invitation layout.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
Qwen Image
- + Excellent typography rendering and placement according to instructions.
- + Perfect adherence to the 3D cartoon/miniature aesthetic.
- + Clean, minimalist composition that fits the 'diorama' request.
- − The sushi elements are very simplified, bordering on toy-like rather than realistic PBR textures.
Stable Diffusion 3.5 Large
- + Highly detailed textures for the sushi fish and rice.
- + Good 3D isometric perspective.
- − Failed to place the text at the top-center, instead placing it on a small sign.
- − The scene is cluttered with too many elements, disregarding the 'minimal garnish' request.
- − Text on the sign is slightly distorted and dark.
Verdict: Qwen Image followed the layout instructions perfectly, placing the bold text at the top-center on a solid background as requested. Stable Diffusion 3.5 Large ignored the layout constraints for the text and created a much busier scene that deviated from the requested minimalist diorama style. Qwen Image is the clear winner for its superior prompt adherence and clean design.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
Qwen Image
- + Excellent character rendering with very distinct animal features.
- + High clarity and sharp focus on all four subjects.
- + Great inclusion of dew sparkles and clear god rays as requested.
- − The composition feels a bit static, more like a posed group photo than 'tumbling together'.
- − The tabby kitten's anatomy at the paws is slightly awkward.
Stable Diffusion 3.5 Large
- + Dynamic composition that perfectly captures the 'tumbling' and 'running' motion requested.
- + Superior lighting effects with a more natural, hazy sunrise atmosphere.
- + Very expressive and joyful facial expressions on the animals.
- − The 'tabby' kitten lacks distinct tabby markings, appearing more like a generic ginger cat.
- − Slightly more motion blur, which reduces the 'ultra-detailed' texture of the fur compared to the other model.
Verdict: While Qwen Image provides sharper details and better adherence to the 'tabby' specific instruction, Stable Diffusion 3.5 Large is the overall winner for its superior composition and energy. Stable Diffusion 3.5 Large better captured the 'joyful' and 'tumbling' aspects of the prompt, creating a more cinematic and immersive scene.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
Qwen Image
- + Successfully captures the vector emblem aesthetic with bold, clear lines.
- + Accurately includes all text elements like 'Est. 1720' and 'Caffè Florian'.
- + Good use of the requested warm brown and cream color palette.
- − The typography is messy with overlapping letters, making 'Florian' difficult to read.
- − Composition of the text is somewhat cramped compared to the icon.
Stable Diffusion 3.5 Large
- + Elegant layout with a more sophisticated vintage feel.
- + Excellent background texture that adheres well to the 'subtle texture' prompt.
- + Clear, legible typography for the main name.
- − Spelling error in the main title ('Cafféé' instead of 'Caffè').
- − The cloche icon is oddly split in the middle with confusing internal lines.
Verdict: Qwen Image follows the vector emblem style more effectively and gets the spelling correct, though its typographic arrangement is cluttered. Stable Diffusion 3.5 Large offers a more graceful and premium 'vintage' composition with superior textures, but it suffers from a spelling error and a disjointed central icon.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
Qwen Image
- + Successfully followed the narrative steps requested in the prompt
- + Clean, modern flat-vector aesthetic with clear icons
- + Text is legible and mostly accurate to the requested mission details
- − Included parenthetical instructions like '(Stop at landing)' as actual text
- − Some minor typos in names and labels like 'Sarth Orbit'
Stable Diffusion 3.5 Large
- + Highly detailed technical aesthetic that looks professional
- + Accurate NASA-inspired color palette usage
- − Failed to follow the requested numbered steps entirely
- − Text is mostly illegible gibberish
- − The rocket depicted is a shuttle hybrid rather than the requested Saturn V icon
Verdict: Qwen Image followed the complex prompt instructions much better, delivering the specific numbered steps and iconography requested in a clear infographic format. Stable Diffusion 3.5 Large produced a visually dense image but failed on almost all semantic requirements, including text legibility and following the six-step sequence.
Explore each model
Stability AI's 8.1-billion parameter Multimodal Diffusion Transformer (MMDiT) text-to-image model featuring improved image quality, typography, complex prompt understanding, and resource-efficiency