Improved version of Alibaba's Qwen image model with better text rendering, finer natural textures, and more realistic human generation.
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
Qwen Image 2512
#30 of 62 in Text-to-Image
Stable Diffusion 3.5 Medium
#57 of 62 in Text-to-Image
Where the votes landed
Qwen Image 2512
0%
win rate
Ties
0%
Stable Diffusion 3.5 Medium
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
Qwen Image 2512
- + Excellent photorealistic rendering of the book texture and wooden table.
- + Sophisticated handling of reflections and refractions within the glass cube.
- − The glass has a strong turquoise tint rather than being clear.
- − The sphere is resting on the bottom rather than appearing suspended, which is typical for this type of prompt, though it followed all instructions.
Stable Diffusion 3.5 Medium
- + Successfully placed the plant behind the cube as requested.
- + The sphere appears more glass-like and is centered within the volume.
- − The book is very thin and lacks realistic page detail.
- − The physics of the hanging sphere are unclear, and the overall image resolution appears lower/fuzzier than Model A.
Verdict: Qwen Image 2512 produces a much higher quality, photorealistic image with convincing textures on the red book and wooden table. While Stable Diffusion 3.5 Medium follows the scene layout well, its execution is hindered by lower clarity and a poorly rendered book that looks compressed.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
Qwen Image 2512
- + Exceptional skin texture and facial realism that looks like a genuine photograph.
- + The bicycle mechanics and structure are much more coherent and realistic.
- + Successfully captures the requested shallow depth of field and 'candid' eye contact.
- − Failed to include motion blur on the passing cars, which appear frozen in time.
- − Minor anatomical distortion on the subject's left hand resting on the seat.
Stable Diffusion 3.5 Medium
- + Beautiful bokeh and lighting in the background creates a strong cinematic atmosphere.
- + Captures the 'light rain' atmosphere and wet pavement reflections more vividly than Model A.
- + Composition feels more like a spontaneous street photo with natural movement.
- − Severe anatomical issues with the man's hands, which are blended into the bicycle frame.
- − The bicycle geometry is nonsensical, particularly with the front wheel and basket integration.
- − The man's face lacks the 'natural skin texture' requested, appearing somewhat muddy/smeared.
Verdict: Qwen Image 2512 is the clear winner due to its superior anatomical accuracy and photographic realism. While Stable Diffusion 3.5 Medium captures the 'cinematic' lighting and rain effects beautifully, it fails significantly on the structural details of the bicycle and the human hands. Qwen Image 2512 produces a much more believable and high-quality image, even if it missed the specific detail for motion blur on the cars.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
Qwen Image 2512
- + Excellent depiction of ornate engraving on the plate armor
- + Highly realistic facial features including subtle scars and skin texture
- + Strong adherence to the bead-in-braid requirement
- − The lighting in the background (torch) is a bit distracting in its proximity
Stable Diffusion 3.5 Medium
- + Dynamic warm lighting and effective use of bokeh sparks/fire in the background
- + Great leather and cloth texture visible under the armor
- + Intense, lifelike eye rendering
- − Missed the request for beads in the hair braids
- − The 'dirt' on the skin looks more like a skin condition or heavy freckling rather than battle grime
Verdict: Qwen Image 2512 is the superior image as it followed every detail of the prompt, including the specific request for beads in the hair. While Stable Diffusion 3.5 Medium has high visual quality and great textures, it failed on the specific bead requirement and the skin marks are less convincing than the scars in the Qwen version.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
Qwen Image 2512
- + Excellent grid layout with a high-impact header.
- + Bold use of color accents and vibrant high-quality food photography.
- + Consistently modern sans-serif fonts that create a clear visual hierarchy.
- − Text is largely nonsensical/gibberish despite look.
- − Small layout spacing issues on the right column categories.
Stable Diffusion 3.5 Medium
- + Natural and realistic food photography.
- + Clean white space that follows a conservative minimalist aesthetic.
- − Layout is disjointed with too much empty space in the center.
- − Text and lines are blurry or bleeding, lacking the requested bold professional finish.
- − Failed to create clear, bold headers for 'Appetizers' and 'Mains' as requested.
Verdict: Qwen Image 2512 is the clear winner as it successfully interprets the 'modern minimalist/bold' design style with a professional layout and vibrant colors. Stable Diffusion 3.5 Medium lacks the graphic design polish required, resulting in a cluttered, blurry layout with poor text rendering.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
Qwen Image 2512
- + Excellent typography with a fiery, glowing effect that perfectly matches the prompt.
- + Superior dynamic composition with clear separation of ingredients in mid-air.
- + Highly detailed food textures and realistic floating debris/embers.
Stable Diffusion 3.5 Medium
- + Effective vibrant orange fire background.
- + Correct price and currency symbol rendering.
- − Failed the 'exploded' and 'suspended' burger component instruction; the burger is mostly assembled.
- − Text lacks the requested fiery, glowing style.
- − Starburst element looks like a simple line icon rather than a cohesive ad element.
Verdict: Qwen Image 2512 followed every instruction in the prompt, creating a truly dynamic exploded view with high-quality, stylized typography. In contrast, Stable Diffusion 3.5 Medium produced a mostly static burger and failed to apply the fiery glows and complex layout requested.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
Qwen Image 2512
- + Excellent text rendering with almost perfect spelling and legibility.
- + The chalk texture and handwriting style appear authentic and consistent.
- + Composition is balanced and mimics a real café chalkboard perfectly.
- − Small spelling error in 'Risitto' (should be Risotto).
Stable Diffusion 3.5 Medium
- + Features a very realistic chalk dust and smudge texture on the board.
- + Captures a wider range of handwriting styles.
- − Serious spelling and legibility issues with almost all words.
- − Failed to follow the date and price formatting instructions accurately.
- − Layout is cluttered and confusing.
Verdict: Qwen Image 2512 followed the prompt with high precision, producing legible text and a coherent layout that looks like a real menu, despite one minor spelling mistake. Stable Diffusion 3.5 Medium struggled significantly with spelling and layout, resulting in nonsensical text that is difficult to read. Qwen Image 2512 is the clear winner for its superior text rendering and adherence to the requested content.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
Qwen Image 2512
- + Excellent anatomical rendering of both the astronaut and the horse
- + High cinematic quality with realistic textures and lighting
- + Strong adherence to the 'horse on top' request without physical glitches
- − The astronaut's face inside the helmet looks slightly distorted
- − The lighting on the horse's back seems a bit disconnected from the space background
Stable Diffusion 3.5 Medium
- + Good use of the 'surreal' prompt with the floating stance
- + Wide cinematic composition with a beautiful starfield
- − Anatomical failure with the astronaut having extra, mismatched legs growing from the horse's side
- − The horse's legs are unnaturally long and thin with poor hoof structure
- − Lower overall detail resolution compared to Model A
Verdict: Qwen Image 2512 is the clear winner as it provides a high-quality, anatomically correct image that perfectly follows the prompt. Stable Diffusion 3.5 Medium suffers from significant structural issues, including the astronaut having multiple sets of legs and the horse's proportions being severely warped.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
Qwen Image 2512
- + Excellent adherence to the passenger's behavior, showing her looking at a phone with a bored expression.
- + High-quality texture on the capybara's fur and clothing.
- + The composition feels grounded and cinematic with realistic lighting.
- − The capybara's paws look more like human-monkey hybrid hands than natural capybara feet.
Stable Diffusion 3.5 Medium
- + The capybara's paws are more anatomically accurate for the species.
- + Very sharp focus on the central subject with vibrant colors.
- − The passenger is not looking at her phone as requested.
- − The perspective is slightly distorted, making the background seating look tiny compared to the capybara.
- − The passenger's expression is more vacant than 'bored/normal'.
Verdict: Qwen Image 2512 is the superior image because it follows the complex multi-subject prompt much more accurately, specifically capturing the businesswoman looking at her phone. While Stable Diffusion 3.5 Medium has a charming central character, it fails to include the interaction with the phone and the spatial composition feels less realistic.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
Qwen Image 2512
- + Excellent typography with high accuracy and a fitting gothic font.
- + High visual quality with atmospheric lighting and a polished cinematic feel.
- + Successfully captures all prompt elements including the specific border, banner, and central pumpkin.
- − Contains a minor spelling error in the main title ('Hallowern').
Stable Diffusion 3.5 Medium
- + Successfully uses a parchment paper texture as requested.
- + Good use of color and high contrast between the pumpkins and the background.
- − Significant text errors throughout the entire image, including spelling and grammar issues.
- − Composition is cluttered, and the jack-o-lanterns are not 'central' as requested.
- − Design looks more like a cartoon than a polished cinematic gothic poster.
Verdict: Qwen Image 2512 is the clear winner as it produces a professional, atmospheric invitation that accurately follows the prompt's layout and style. While it has one small typo, Stable Diffusion 3.5 Medium fails significantly on text legibility, composition, and the requested 'cinematic' aesthetic.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
Qwen Image 2512
- + Perfectly follows the specific layout request with bold text and a Japanese flag icon.
- + Excellent miniature diorama feel with a detailed 3D base and high-quality textures.
- + Clean isometric composition with a clear 45-degree top-down perspective.
- − The text 'JAPAN' has slightly inconsistent coloring in the letter fills.
Stable Diffusion 3.5 Medium
- + Clean, minimalist aesthetic with nice lighting on the white plate.
- + Accurate text rendering for both 'JAPAN' and 'SUSHI'.
- − Missing several prompt elements including the small flag icon and the raised diorama base.
- − The 3D cartoon miniature style is less pronounced compared to the competitor.
- − The composition is slightly off-center with the text partially obscured by transparency.
Verdict: Qwen Image 2512 is the clear winner as it strictly adheres to all layout and stylistic instructions, including the flag icon and the miniature diorama base. Stable Diffusion 3.5 Medium fails to include the flag and the requested base, resulting in a much simpler composition that lacks the 'miniature scene' quality requested.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
Qwen Image 2512
- + Successfully included all four requested animals (dog, cat, rabbit, fox).
- + Excellent anatomical details and fur textures on each animal.
- + Beautifully rendered lighting with clear 'god rays' effects as requested.
Stable Diffusion 3.5 Medium
- + Captured the 'tumbling' and 'chasing' action better than the static pose of the other model.
- + Vibrant colors and a high-contrast, cheerful aesthetic.
- − Failed to include the requested baby bunny.
- − The animals have slightly distorted or overly stylized anatomy that looks less photorealistic.
- − One butterfly is cut off at the edge of the frame.
Verdict: Qwen Image 2512 is the clear winner as it followed the prompt perfectly, including all four specific animals while maintaining a high degree of photorealism and fine detail. In contrast, Stable Diffusion 3.5 Medium missed the rabbit entirely and the overall image quality feels more like a digital illustration than the requested 'hyper-photorealistic' masterpiece.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
Qwen Image 2512
- + Excellent typography with perfect spelling and accent on 'Caffè'.
- + Highly detailed and aesthetic cross-hatching and shading for a vintage look.
- + Accurate representation of all prompt elements including the banner and steam.
- − The 'minimalist' instruction was somewhat ignored in favor of a highly detailed illustration style.
Stable Diffusion 3.5 Medium
- + Successfully captured the warm brown and cream tones.
- + Classic woodcut/engraving texture style is appropriate for the theme.
- − Failed significantly with text rendering, misspelling the name as 'Florrian' and the date as '170'.
- − The cloche dome shape is distorted and looks more like a circular badge or tent.
- − The banner is messy and contains unintelligible gibberish text.
Verdict: Qwen Image 2512 is the clear winner as it delivered a professional, high-quality logo with perfect spelling and a very cohesive visual style. Stable Diffusion 3.5 Medium failed at the fundamental task of rendering the requested text accurately and produced a distorted central icon.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
Qwen Image 2512
- + Excellent adherence to the six-step sequence requested in the prompt.
- + Follows the NASA-inspired color palette perfectly with clean, modern vector aesthetics.
- + High-quality illustrations of the Saturn V and Lunar Module that maintain consistency.
- − Several spelling errors in the labels (e.g., 'Translaurtcoit', 'Desceeint').
- − Includes duplicate numbering and some confusing layout choices with the step labels.
Stable Diffusion 3.5 Medium
- + Strong minimalist vector style that feels like modern graphic design.
- + Creative use of red accents and orbital lines that fit the NASA-inspired theme.
- − Failed to follow the requested six-step sequence correctly.
- − The icons do not match the specific instructions (e.g., no clear Saturn V icon).
- − Severe text distortion and illegible gibberish throughout the image.
Verdict: Qwen Image 2512 is the clear winner as it successfully illustrative all six requested steps of the mission with the correct iconography, despite some spelling errors. Stable Diffusion 3.5 Medium failed to follow the specific step-by-step instructions and produced mostly nonsensical text and disconnected graphics.
Explore each model
Stability AI's 2.5-billion parameter Multimodal Diffusion Transformer with improvements (MMDiT-X) text-to-image model optimized for consumer hardware, featuring improved image quality, typography, and complex prompt understanding