OpenAI's legacy image generation model supporting generations, edits with masks (inpainting), and variations
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
DALL-E 2
#59 of 62 in Text-to-Image
Qwen Image 2.0
#34 of 62 in Text-to-Image
Where the votes landed
DALL-E 2
0.0%
win rate
Ties
0.0%
Qwen Image 2.0
100.0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
DALL-E 2
- + Features a wooden table surface and soft lighting.
- − Failed almost all spatial and object requirements.
- − The blue element is a massive pot instead of a small sphere inside the cube.
- − The red element is inside the cube rather than a book on top.
- − The plant is not clearly behind and visible through the glass.
Qwen Image 2.0
- + Perfect adherence to all spatial instructions and object placements.
- + High visual clarity and realistic textures for glass, paper, and foliage.
- + Accurately renders reflections and light coming from the left window.
- − Minor physics irregularity with the sphere appearing to float in the center of the cube.
Verdict: DALL-E 2 failed to handle the complex spatial relationships of the prompt, merging the colors and objects incorrectly into a confusing composition. Qwen Image 2.0 followed every instruction perfectly, placing the red book on top, the blue sphere inside, and the plant behind the cube with high-quality photographic realism.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
DALL-E 2
- + Successfully captures a 'broken' or imperfect framing style
- + Good reflection details on the wet pavement
- − Extreme blur makes the subject unrecognizable as an elderly Japanese man
- − Lacks almost all detail beyond vague shapes and colors
- − Composition is poor and confusing
Qwen Image 2.0
- + Excellent adherence to all prompt details including age, ethnicity, and activity
- + Photorealistic skin textures and fine details on the bicycle
- + Effective use of shallow depth of field and motion blur on the background car
- − The bike pedal/chain area has some slight structural physical errors
- − Legs/hands interface is slightly cramped visually
Verdict: DALL-E 2 failed to generate a coherent image, producing an extremely blurry shot where the subject is unidentifiable. Qwen Image 2.0 followed every instruction perfectly, delivering a high-quality, realistic photograph with impressive textures and an authentic candid feel.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
DALL-E 2
- + Features a distinct bokeh effect in the background.
- − Extreme lack of clarity and low resolution.
- − Does not follow the prompt details like braided hair or lifelike eyes.
- − Distorted perspective makes it difficult to distinguish the figure from the background.
Qwen Image 2.0
- + Excellent adherence to all prompt details including braided hair with beads and ornate armor.
- + High-quality skin textures with clear scars and dirt.
- + Effective lighting and color contrast between the metal and warm firelight.
- − The bokeh sparks appear as floating circles rather than integrated light trails.
- − Slight anatomical distortion in the hands holding the sword.
Verdict: Qwen Image 2.0 followed every specific detail of the prompt, including the complex hair requests and armor engraving, while maintaining high visual fidelity. DALL-E 2 produced a low-quality, blurry image that failed to represent the character's face, hair, or armor details effectively.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
DALL-E 2
- + Strong bold typography that feels modern.
- + Unique geometric layout for the abstract food elements.
- − Fails to follow the 'grid' instruction for food photos.
- − The food is chopped into messy, unrecognizable fragments.
- − Illegible gibberish text with no clear sections for appetizers or mains.
Qwen Image 2.0
- + Perfect adherence to the grid layout and section headers.
- + High-quality, appetizing food photography for each item.
- + Clear categorization of Appetizers, Pizza, and Mains as requested.
- − Secondary text under images is largely illegible.
- − Repeated pricing numbers reduce the sense of realism.
Verdict: Qwen Image 2.0 followed the prompt instructions precisely, delivering a clean, professional menu grid with high-quality food assets and logical sections. DALL-E 2 struggled significantly, producing abstract, fragmented imagery and a layout that does not function as a cohesive restaurant menu.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
DALL-E 2
- + Strong sense of vibrant, fiery motion
- + Abstract interpretation captures a 'magic' essence
- − Text is nonsensical and misspelled as 'MARGIC BAGUEC'
- − Food items are poorly defined and look burnt or unappetizing
- − Low resolution with significant digital artifacts
Qwen Image 2.0
- + Excellent text rendering with correct spelling and requested effects
- + High photorealistic detail in the food textures and ingredients
- + Perfect adherence to all prompt constraints including specific price and starburst
- − The 'exploded' effect is relatively subtle compared to the potential of the prompt
- − Slightly more commercial/stock-photo aesthetic than artistic
Verdict: Qwen Image 2.0 followed every instruction perfectly, producing clear, legible text and a high-quality photorealistic image. DALL-E 2 failed significantly on text rendering and image clarity, resulting in an unappetizing product that didn't meet the advertisement requirements.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
DALL-E 2
- + The text has a rough, thick chalk-like texture.
- − The text is completely illegible and gibberish.
- − It fails to follow any specific menu items requested in the prompt.
- − The background context of a 'cozy café' is missing.
Qwen Image 2.0
- + Near-perfect adherence to the text prompt, including specific menu items and prices.
- + Excellent visual quality with a realistic chalkboard surface and ambient café lighting.
- + The handwriting looks authentic with consistent chalk texture and natural variation.
- − The date '2026' is slightly compressed in width compared to the other text.
Verdict: Qwen Image 2.0 followed the prompt with impressive accuracy, correctly rendering complex menu names and prices in a realistic café setting. DALL-E 2 failed the challenge completely, producing illegible textures and nonsensical characters that do not follow the requested text.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
DALL-E 2
- + Successfully positions the astronaut on the moon's surface.
- + Consistent color palette across the subject and background.
- − Failed the primary logic constraint of having the horse on top.
- − Poor image quality with significant pixelation and low resolution.
- − Anatomical issues with the horse's legs and overall shape.
Qwen Image 2.0
- + High visual quality with sharp details and cinematic lighting.
- + Excellent rendering of the space suit and celestial background.
- + Clearer interpretation of the space theme with orbital views.
- − Failed the specific logic constraint of placing the horse on top of the astronaut.
- − Anatomical error with the horse having five visible legs.
Verdict: Both models failed the specific logic constraint of placing the 'horse on top' of the astronaut, choosing instead to generate a standard astronaut-riding-horse image. Qwen Image 2.0 is the clear winner due to its significantly higher resolution, cinematic lighting, and superior detail compared to the blurry and distorted output from DALL-E 2.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
DALL-E 2
- + Attempts to show the passenger interacting with a phone.
- − Extreme anatomical distortions in the human face and hands
- − The capybara looks like an unidentifiable orange statue or puppet
- − Background is mostly black and lacks the detail of Manhattan
- − Failed to place the passenger in the back seat properly
Qwen Image 2.0
- + Excellent photorealism and high visual quality
- + Accurately depicts all prompt elements including the capybara's hat, jacket, and expression
- + Perfectly captures the 'bored' expression of the businesswoman in the background
- − The passenger appears to be in the front passenger seat rather than the back seat
- − Minor blurring on the capybara's right paw where it meets the steering wheel
Verdict: Qwen Image 2.0 produced a high-quality, professional-looking image that follows almost every detail of the prompt with realistic textures and lighting. In contrast, DALL-E 2 struggled significantly, producing nightmare-like distortions, poor anatomy, and an unrecognizable main subject. Qwen Image 2.0 is the clear winner for its clear composition and adherence to the requested atmosphere.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
DALL-E 2
- + Features an authentic hand-drawn vintage aesthetic.
- + Strong color palette consistent with traditional Halloween posters.
- − Text is largely illegible and fails to follow specific instructions.
- − Lacks the central jack-o-lantern and thorn border requested.
Qwen Image 2.0
- + Excellent text rendering with total prompt adherence.
- + Composition perfectly integrates all requested elements including the thorn border and jack-o-lantern.
- − The digital polish makes it look less like an authentic 'vintage' parchment.
- − Lighting on the scroll feels slightly detached from the background.
Verdict: Qwen Image 2.0 followed every specific detail of the prompt, providing clear and accurate text for the invitation details and title. DALL-E 2 failed to render legible text or several of the key visual elements like the thorns and spiderwebs. Qwen Image 2.0 is the superior choice for a functional invitation design.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
DALL-E 2
- + Clean isometric perspective
- + Uniform light blue background
- − Failed to render all requested text, showing only 'Sush'
- − Poor food representation with unidentifiable, abstract blobs
- − Missing 'JAPAN' text and flag icon
Qwen Image 2.0
- + Perfect adherence to text and icon requirements
- + High-quality realistic PBR textures on the sushi and wood base
- + Clean, professional composition that matches the 'diorama' aesthetic
- − The perspective is more of a front-angle than a strict 45-degree top-down isometric view
- − Slight misalignment in the wood grain at the corner
Verdict: Qwen Image 2.0 followed every instruction in the prompt, including the specific text and flag icon, while providing highly detailed and appetizing sushi. DALL-E 2 failed significantly on both text rendering and the visual quality of the food items, producing an abstract and incomplete scene.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
DALL-E 2
- + Features the requested golden retriever and kitten.
- + Bright, high-key lighting correlates with the sunrise request.
- − Poor anatomical coherence with distorted animal limbs and faces.
- − Significant background blur and artifacts that look low-resolution.
- − Missing the fox kit as a distinct identifiable creature.
Qwen Image 2.0
- + Perfectly depicts all four requested animals: puppy, kitten, bunny, and fox kit.
- + Atmospheric lighting with clear god rays and dew sparkles as requested.
- + Sharp details in fur texture and eyes with high visual clarity.
- − The fox's anatomy is slightly tangled beneath the other animals, making it harder to see.
- − Butterflies are a bit static for a 'chase' scene.
Verdict: Qwen Image 2.0 followed every aspect of the prompt, successfully rendering all four specific animal species with high-quality textures and professional lighting. In contrast, DALL-E 2 produced a low-fidelity image with significant anatomical distortions and failed to include the fox kit or bunny clearly. Qwen Image 2.0 is the clear winner for its superior composition, detail, and prompt adherence.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
DALL-E 2
- + Successfully captures the requested warm brown and cream tones
- + Achieves a minimalist vector aesthetic
- − Text is nonsensical and fails to include the requested brand name
- − Steam element is poorly rendered and detached
- − Composition feels cluttered despite the minimalist intent
Qwen Image 2.0
- + Excellent typography with perfect spelling of 'Caffè Florian' and 'Est. 1720'
- + Highly coherent implementation of the cloche, steam, and banner elements
- + Strong vector logo composition with tasteful subtle texturing
- − The shading on the cloche is slightly more illustrative than a strictly flat minimalist logo
Verdict: Qwen Image 2.0 followed every instruction perfectly, producing a professional-grade logo with accurate text rendering and a cohesive design. In contrast, DALL-E 2 failed significantly on text legibility and failed to include the established date or the correct brand name.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
DALL-E 2
- + Follows the requested color palette well.
- − Text consists of nonsensical gibberish and misspells 'Apollo'.
- − Layout is chaotic and fails to present any of the specific requested steps.
- − Contains significant artifacts and lacks a coherent infographic structure.
Qwen Image 2.0
- + Excellent prompt adherence, including all six specific steps of the mission.
- + Clean, modern vector style with legible and correctly spelled text.
- + Accurate iconography including the Saturn V and Lunar Module.
- − Includes a minor spelling error in 'Translunjar'.
- − The icons for descent and landing are very similar in design.
Verdict: Qwen Image 2.0 followed the prompt instructions near-perfectly, creating a structured and informative infographic that captures the specific mission steps requested. DALL-E 2 produced a chaotic image with indecipherable text and failed to follow the sequential logic of the prompt, making it unusable as an infographic.
Explore each model
Alibaba's Qwen Image 2.0 model with enhanced text rendering, supporting both Chinese and English prompts with up to 6 images per request