Alibaba's Qwen Image 2.0 model with enhanced text rendering, supporting both Chinese and English prompts with up to 6 images per request
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
Qwen Image 2.0
#34 of 62 in Text-to-Image
Stable Diffusion 3.5 Large Turbo
#61 of 62 in Text-to-Image
Where the votes landed
Qwen Image 2.0
0%
win rate
Ties
0%
Stable Diffusion 3.5 Large Turbo
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent adherence to the complex spatial relationship requirements.
- + Convincing glass reflections and realistic lighting from the window.
- + Very high photographic quality with natural textures on the wood and book cover.
- − Internal reflections of the sphere appear slightly chaotic or duplicated.
Stable Diffusion 3.5 Large Turbo
- + Clean, sharp aesthetic with high contrast.
- + Good rendering of the glass transparency and the plant behind it.
- − Failed the spatial prompt: the red book is inside/under the cube rather than on top of it.
- − The blue sphere is on top of the book, which was not requested.
- − The lighting feels more synthetic and clinical compared to the requested soft window light.
Verdict: Qwen Image 2.0 followed all instructions perfectly, placing the red book on top of the glass cube and the sphere inside it, while maintaining a very realistic photographic style. Stable Diffusion 3.5 Large Turbo failed to properly arrange the objects, placing the book at the base and the sphere on top of it, resulting in a less accurate interpretation of the prompt. Qwen Image 2.0 is the clear winner for its superior composition and prompt adherence.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent adherence to aesthetic prompts like imperfect framing and natural skin texture
- + Highly realistic textures on the wet pavement and the rusty bicycle
- + Authentic candid feeling that matches the 50mm lens photography style
- − Anatomy of the second person on the right is cut off awkwardly
- − Minor inconsistencies in the bicycle chain and pedal connection
Stable Diffusion 3.5 Large Turbo
- + Clear interpretation of the red bicycle and rainy environment
- + Good lighting on the subject
- − Failed the 'no stylization' prompt, appearing very plastic and CGI-like
- − Poor anatomical rendering of the hands and face
- − The rain effect is rendered as simple white vertical lines
Verdict: Qwen Image 2.0 successfully captured the requested 'candid street photo' aesthetic with impressive realism, natural skin textures, and wet pavement reflections. In contrast, Stable Diffusion 3.5 Large Turbo produced a stylized, almost cartoonish image with significant anatomical issues and a lack of the requested photographic realism.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent adherence to the 'beads in hair' and 'battle-worn' descriptor with realistic skin texture.
- + Superior cinematic lighting with convincing firelight reflections and detailed bokeh sparks.
- + Highly detailed materials including engraved plate, leather, and tattered cloth underlayers.
- − The eyes have an orange tint that might look slightly supernatural for a standard paladin.
- − The depth of field is a bit busy in the background compared to a true shallow focus.
Stable Diffusion 3.5 Large Turbo
- + Ornate armor design with clear engraving and golden accents.
- + Clean composition with a focused portrait style.
- − Missed the 'beads in hair' instruction entirely.
- − Skin textures and blood look more like digital paint than lifelike 'faint scars and dirt'.
- − The hair and lighting effects look flat and artificial compared to the cinematic quality requested.
Verdict: Qwen Image 2.0 followed every detail of the prompt, capturing the specific request for hair beads, battle-worn textures, and complex lighting. Stable Diffusion 3.5 Large Turbo produced a much more generic, 'plastic' looking image that failed to include several key prompt elements like the beads and realistic skin details.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent adherence to the grid-based food photo layout requested in the prompt.
- + Very high-quality food photography with realistic lighting and appetizing details.
- + Clear, bold headings for sections that match the requested typography style.
- − Text under the images (dish names) is gibberish/distorted.
- − Missing specific text-list sections for the menu items, relying only on photo captions.
Stable Diffusion 3.5 Large Turbo
- + Includes designated text-based menu sections for Pizza and 'Mians'.
- + Creative top-down composition that feels like a marketing flat lay.
- − Failed the prompt's request for a minimalist grid layout, resulting in a cluttered and confusing composition.
- − Food photos look artificial and plastic-like compared to realistic culinary photography.
- − Significant typos in the main headings, such as 'Mians' instead of 'Mains'.
Verdict: Qwen Image 2.0 followed the specific layout requirements much more effectively, producing a clean, professional-looking grid that looks like a real minimalist menu. While the fine text is garbled, the overall aesthetic and quality of the food photography far surpass Stable Diffusion 3.5 Large Turbo, which produced cartoonish food and a chaotic layout that ignored the grid instruction.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent typography rendering with the requested fiery, glowing effect.
- + Highly photorealistic food textures and lighting.
- + Followed all instructions, including the starburst for the price and the secondary message.
- − The 'exploded' effect is present but relatively subtle compared to the potential of the prompt.
Stable Diffusion 3.5 Large Turbo
- + Features a dynamic, glowing, and fiery background.
- + The burger is suspended in mid-air as requested.
- − Completely failed to include any of the requested text elements.
- − The image style is a bit more 'digital art' and less photorealistic than requested.
- − The burger is not 'exploded' or separated into its core components.
Verdict: Qwen Image 2.0 followed every instruction in the prompt, including complex text rendering with effects and specific composition elements like the starburst. Stable Diffusion 3.5 Large Turbo failed to include any text and did not achieve the specific 'exploded' burger look, resulting in a generic burger image.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent text rendering with perfect spelling and realistic chalk texture.
- + Authentic handheld handwriting style with natural variations as requested.
- + Realistic environment with convincing lighting and depth of field.
- − The 'Today's Specials' title is more of a print-script mix than high-end elegant cursive.
Stable Diffusion 3.5 Large Turbo
- + Clean and balanced composition with modern interior design elements.
- + Good lighting effects on the chalkboard surface.
- − Very poor text rendering with numerous spelling errors and gibberish.
- − The text looks like a digital font rather than authentic chalk handwriting.
- − Failed to follow the requested layout and date.
Verdict: Qwen Image 2.0 is the clear winner as it followed every detail of the prompt, including complex spelling and the specific requirement for a 'chalk texture' rather than digital fonts. Stable Diffusion 3.5 Large Turbo failed significantly on text accuracy and realism, producing a board that feels artificial and illegible.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent texture on the astronaut's suit and the horse's coat.
- + Strong cinematic lighting with realistic reflections on the visor.
- + Captures a surreal atmosphere with the floating droplets and starry backdrop.
- − The horse's anatomy is slightly fused with its front legs looking unnatural.
- − The 'not vice versa' instruction was interpreted normally, missing the potential surreal prompt check.
Stable Diffusion 3.5 Large Turbo
- + Dynamic sense of movement with the shadow and motion blur.
- + Clean, high-contrast composition that emphasizes the space setting.
- − The horse's back right leg and tail area become blurry and distorted.
- − The visor of the astronaut is opaque and lacks detail compared to the competitor.
- − The small planet in the top left looks like a flat, low-resolution sticker.
Verdict: Qwen Image 2.0 provides a much more detailed and visually rich image with superior textures and lighting. Stable Diffusion 3.5 Large Turbo struggles with anatomical coherence in the horse's legs and lacks the cinematic depth found in the Qwen output.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent adherence to the 'bored' expression for the passenger
- + High-quality fur texture and realistic lighting consistent with a night scene
- + Includes specific details like the smartphone and the chauffeur-style taxi cap
- − The passenger is sitting in the front passenger seat instead of the requested back seat
- − Anatomy of the capybara's paws on the wheel is slightly distorted
Stable Diffusion 3.5 Large Turbo
- + Strong composition that places the passenger in the back seat as requested
- + Clean, cinematic lighting and high contrast
- + Clear text on the cap and correct steering wheel grip
- − The passenger is not looking at her phone
- − The capybara's face looks slightly less photorealistic and more stylized
- − The passenger is out of focus and lacks detail
Verdict: Qwen Image 2.0 captures the surreal humor of the prompt much better, with a perfectly bored businesswoman on her phone, though it fails to place her in the back seat. Stable Diffusion 3.5 Large Turbo follows the spatial instructions better by putting the passenger in the back, but misses the core detail of her looking at her phone. Qwen Image 2.0 is the preferred choice for its superior detail and character acting which brings the prompt's concept to life.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
Qwen Image 2.0
- + Perfect text rendering for all requested strings including fine details.
- + Exceptional adherence to the gothic aesthetic and atmospheric lighting.
- + Well-integrated composition with a realistic parchment texture and detailed border.
- − The jack-o-lantern appears slightly more like a 3D model against a 2D background than the rest of the image.
Stable Diffusion 3.5 Large Turbo
- + Strong border design with stylized thorns and webs.
- + Clear, high-contrast central jack-o-lantern design.
- − Failed to include a significant amount of the requested text, including event details.
- − Text that was included lacks the requested 'elegant gothic' style.
- − The scroll banner element is completely missing.
Verdict: Qwen Image 2.0 followed the prompt instructions perfectly, rendering all the textual information correctly and capturing a cohesive vintage gothic mood. Stable Diffusion 3.5 Large Turbo failed to include the majority of the text and banner elements, resulting in a generic graphic rather than a functional invitation.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent text rendering with clean, bold typography.
- + Realistic appetising textures on the sushi and wooden board.
- + Correct Japanese flag representation and clean layout.
- − Missed the 'isometric' perspective requested, opting for a standard photographic angle.
- − Lacks the '3D cartoon miniature' aesthetic.
Stable Diffusion 3.5 Large Turbo
- + Successfully captured the 45-degree isometric diorama style requested.
- + Clean 3D render aesthetic with soft lighting.
- + Creative interpretation of the scene as a physical miniature.
- − Failed the text rendering, misspelling 'SUSHI' as 'SIIHI'.
- − The flag is generic and does not accurately represent the Japanese flag.
- − Text placement is on the objects rather than centered at the top as requested.
Verdict: Qwen Image 2.0 followed the text and flag instructions perfectly but failed to provide the specific isometric 3D miniature style requested. Stable Diffusion 3.5 Large Turbo nailed the isometric diorama composition and style, but suffered from poor text spelling and an incorrect flag. Qwen is the better image due to higher clarity and correct information, despite the perspective mismatch.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
Qwen Image 2.0
- + Successfully included all four requested animals (dog, cat, rabbit, fox).
- + High level of realism in fur texture and lighting effects.
- + Complex, dynamic interaction between the animals that fits the prompt.
- − The fox's face/mouth has a slight structural artifact where it meets the paws.
Stable Diffusion 3.5 Large Turbo
- + Cute, stylized aesthetic with clean lines.
- + Vibrant lighting and colors.
- − Failed to include the bunny and fox kit entirely.
- − Lacks the requested 'hyper-photorealistic' style, appearing more like a digital painting.
- − The dog's paws are fused and anatomically incorrect.
Verdict: Qwen Image 2.0 followed the prompt precisely, including all four specific animals and achieving a high-quality photorealistic look. In contrast, Stable Diffusion 3.5 Large Turbo failed to include half of the requested animals and produced a stylized, cartoonish result with significant anatomical flaws.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
Qwen Image 2.0
- + Perfect text accuracy and spelling.
- + Excellent vector emblem style with clean lines and balanced composition.
- + Correct interpretation of a cloche dome.
- − The steam is stylized in a way that looks slightly like a flame.
Stable Diffusion 3.5 Large Turbo
- + Strong vintage aesthetic and rich color palette.
- + Sophisticated detail in the banner and shading.
- − Spelling error in the main name ('Caffeé Florin').
- − Cluttered composition where the dome is merged with a cup handle, making the object shape confusing.
- − Visible artifacts and messy rendering on the banner ends.
Verdict: Qwen Image 2.0 followed the prompt instructions precisely, producing a clean, professional vector logo with perfect typography. Stable Diffusion 3.5 Large Turbo had better textured 'vintage' vibes, but it failed on basic spelling and the composition of the cloche dome was muddled and illogical.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
Qwen Image 2.0
- + Successfully included all six requested infographic steps in order.
- + Excellent text legibility and mostly correct spelling including names.
- + Applied the requested NASA-inspired color palette effectively.
- − Simple vertical layout lacks the dynamic composition of a professional infographic.
- − Spelling error in 'Translunjar'.
- − Iconography is somewhat inconsistent in scale and detail.
Stable Diffusion 3.5 Large Turbo
- + Excellent visual composition and 'modern vector' aesthetic.
- + Sophisticated use of negative space and layout design.
- + Artistic interpretation feels more like a high-end poster.
- − Failed to include several of the requested steps (Descent, Lunar Orbit).
- − Garbled, nonsensical text throughout the infographic section.
- − Iconography does not match the specific requested items like the Saturn V.
Verdict: Qwen Image 2.0 followed the instructions significantly better, providing all six requested chronological steps with mostly correct text and clear icons. While Stable Diffusion 3.5 Large Turbo produced a more visually striking and artistically balanced layout, it failed on a functional level as an infographic by omitting several steps and using 'lorem ipsum' style garbled text.
Explore each model
Distilled version of SD 3.5 Large that generates high-quality images in just 4 steps, offering faster inference and reduced costs