OpenAI's previous generation image model with higher quality than DALL-E 2 and support for larger resolutions
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
DALL-E 3
#38 of 62 in Text-to-Image
Wan 2.6
#28 of 62 in Text-to-Image
Where the votes landed
DALL-E 3
0.0%
win rate
Ties
0.0%
Wan 2.6
100.0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
DALL-E 3
- + High visual clarity and artistic detail in the wood and glass textures.
- + Good use of shadows and highlights to create atmosphere.
- − Failed the spatial instructions: both the book and the blue sphere are inside the cube frame.
- − The cube is rendered as a wooden display case rather than a glass cube.
Wan 2.6
- + Perfect adherence to all spatial instructions, including object placement.
- + Realistic handling of the window light and the refraction through the glass cube.
- + Follows the prompt literally with a glass cube, blue sphere inside, and red book on top.
- − The red book has slightly aged/worn edges that weren't specifically requested but add realism.
- − The blue sphere's reflection on the bottom glass is a bit sharp.
Verdict: While DALL-E 3 produces a more stylized and high-contrast image, it fails significantly on prompt adherence by placing the red book inside the wooden frame instead of on top. Wan 2.6 followed every specific instruction regarding object placement and materials, resulting in a much more accurate representation of the request.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
DALL-E 3
- + Excellent reflection in the puddle creating a strong artistic composition
- + Good adherence to the imperfect framing request with elements in the extreme foreground
- + Captures a wider street atmosphere with distinct Japanese lanterns and signage
- − Anatomical issues with the man's neck and posture appearing distorted
- − Faces significant AI artifacts in the background and foreground blur
- − Skin texture looks smoothed and painterly rather than natural
Wan 2.6
- + Superb natural skin texture and highly realistic rain droplets on clothing
- + Excellent adherence to the 'realistic' and 'no stylization' requirements
- + Stronger execution of motion blur from the passing car in the background
- − The red of the bicycle is slightly less vibrant than requested
- − Composition is a bit more centered and conventional compared to the requested 'imperfect framing'
Verdict: Wan 2.6 is the clear winner due to its exceptional realism and adherence to technical photography prompts like 'natural skin texture' and 'no stylization.' While DALL-E 3 creates a more artistic composition with the puddle reflection, it suffers from anatomical distortions and a distinctly 'generated' aesthetic that fails the realism requirement.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
DALL-E 3
- + Excellent high-contrast warm lighting
- + Intricate detail on the armor engravings and leather straps
- + Compelling lifelike eyes with realistic reflections
- − Failed to include the specific request for braided hair with beads
- − The scars look like clean, stylized lines rather than organic battle wounds
Wan 2.6
- + Perfect adherence to the hair braids and beads requirement
- + Superior depiction of dirt and realistic battle wear
- + Stunningly detailed cloth underlayers with frayed edges and realistic textures
- − Lighting is slightly flatter compared to the intense glow in Model A
- − The armor engraving is a bit less defined in the shadowed regions
Verdict: While DALL-E 3 captures a very striking cinematic aesthetic, Wan 2.6 is the superior image for its strict adherence to all prompt elements, specifically the braided hair and beads which DALL-E 3 missed. Wan 2.6 also offers a more realistic interpretation of 'battle-worn' through gritty textures on both the skin and the layered clothing.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
DALL-E 3
- + Excellent variety of layouts showing different design concepts
- + Vibrant color accents and artistic food photography
- + Captures a clean, professional aesthetic
- − Text is largely nonsensical garble
- − Does not follow the grid structure for food photos as strictly in some panels
- − The multi-page mock-up view is slightly cluttered
Wan 2.6
- + Legible English headers for requested categories
- + Clean and orderly grid of food photos
- + Excellent font choice and hierarchy for a casual dining menu
- − Prices are nonsensical (e.g., $1.95 for a pizza, $0.09 for an item)
- − Body text transitions into gibberish at smaller scales
- − Food photos are somewhat repetitive in color and subject
Verdict: Wan 2.6 provides a much more functional and realistic menu layout with legible headers and a cohesive grid that follows the prompt's structural requirements. While DALL-E 3 offers more artistic flair, its text is completely unreadable, whereas Wan 2.6 successfully integrates the specific section names requested.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
DALL-E 3
- + Excellent photorealistic texture on the burger patty and cheese.
- + Creative use of flying ingredients beyond the main stack to enhance the 'exploded' feel.
- + Clean composition with a glowing ground effect.
- − Several spelling errors in the text including 'AGIC BURGR' and 'Limiited'.
- − The price tag is in a square box rather than the requested starburst.
Wan 2.6
- + Perfect adherence to all text requirements with zero spelling errors.
- + Accurately followed the 'starburst' instruction for the price tag.
- + Superior execution of the fiery, glowing effect on the text and background.
- − The food components are less 'exploded' and more slightly detached compared to Model A.
- − The starburst graphic looks a bit like a flat sticker rather than a fully integrated 3D element.
Verdict: While DALL-E 3 produced a very high-quality food render with great textures, it failed significantly on the text rendering and the specific starburst shape request. Wan 2.6 followed every specific instruction perfectly, including complex text rendering with effects and the specific starburst price tag, making it the superior choice for a ready-to-use advertisement.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
DALL-E 3
- + Excellent artistic lighting and atmosphere
- + Includes elegant decorative chalk flourishes
- − Extreme spelling errors for all menu items
- − Price formatting is nonsensical with a giant $234 label
- − Layout is cluttered and hard to read
Wan 2.6
- + Perfect text adherence with zero spelling errors
- + Convincing chalk texture with smudges and dust on the ledge
- + Clear and legible handwriting-style composition
- − Slightly simpler composition compared to the artistic frame in Model A
Verdict: Wan 2.6 is the clear winner as it followed the complex text prompt perfectly, including the exact menu items and prices without a single typo. DALL-E 3 failed significantly on the literacy aspect, producing several misspelled words and nonsensical pricing despite its pleasing aesthetic lighting.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
DALL-E 3
- + Excellent atmospheric lighting and cinematic glow
- + Beautifully rendered nebula and cloud-like horizon
- + Clean, minimalist composition
- − Failed the logical constraint; the astronaut is riding the horse, not the horse riding the astronaut
- − Lower level of intricate texture on the spacesuit and horse
Wan 2.6
- + Extremely high detail in the horse's coat and equipment
- + Vibrant colors and high-quality textures
- + Dynamic posing and realistic lighting on the astronaut's visor
- − Failed the logical constraint; the astronaut is riding the horse, not the horse riding the astronaut
- − Anatomically confused hooves and legs on the horse
Verdict: Both DALL-E 3 and Wan 2.6 failed the specific 'horse on top' prompt instruction, instead providing a standard 'astronaut riding a horse' image. Wan 2.6 is the better purely visual image due to its remarkable detail and vibrant colors, though DALL-E 3 captures a more ethereal and surreal atmosphere.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
DALL-E 3
- + Excellent fur detail and rim lighting on the capybara.
- + Logical interior perspective with a clear dashboard and steering wheel interaction.
- + Includes subtle creative details like a storefront sign reading 'CAPYBARA'.
- − Completely fails to include the passenger in the back seat.
- − The capybara's paw placement looks a bit stiff and anthropomorphized.
Wan 2.6
- + Successfully includes all elements of the prompt, including the bored passenger on her phone.
- + Captured the rainy, neon atmosphere of Manhattan very effectively through the windshield.
- + The capybara's expression and coat texture are highly realistic.
- − The roof of the taxi appears to have a second windshield/glass panel reflecting the front dashboard, which is nonsensical.
- − The steering wheel appears to be floating or disconnected from a steering column.
Verdict: While DALL-E 3 has higher technical image quality and cleaner details, it failed a major part of the prompt by omitting the passenger. Wan 2.6 adhered to all instructions, including the specific mood and action of the businesswoman, despite some anatomical/mechanical clipping issues with the car's roof and wheel.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
DALL-E 3
- + Ornate and intricate gothic 3D-style composition
- + Highly atmospheric and cinematic lighting
- + Includes all requested decorative elements like trees and webs
- − Text rendering is mostly gibberish or heavily distorted
- − Failed to include the specific event details accurately
- − The small scroll banner is missing or unrecognizable
Wan 2.6
- + Excellent text legibility and accuracy for all requested fields
- + Perfect adherence to the specific message on the scroll banner
- + Clearer visual of the jack-o-lantern and moody night sky
- − The composition is a bit more generic than the 3D depth of its competitor
- − Slightly less 'polished' look in the blending of the border and background
Verdict: Wan 2.6 is the clear winner because it successfully followed the complex text instructions, rendering all names, dates, and the specific banner message perfectly. DALL-E 3 produced a more visually striking and atmospheric gothic piece, but the text is unreadable, making it fail as an actual invitation.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
DALL-E 3
- + Excellent 3D miniature clay-like aesthetic with soft, refined textures.
- + Vibrant lighting and high-quality rendering of stylized food.
- + Includes all thematic elements like flags and text on the diorama base.
- − Failed to place 'JAPAN' and 'SUSHI' at the top-center; instead integrated it into the base.
- − Missed the word 'SUSHI' entirely in the text rendering.
Wan 2.6
- + Perfect adherence to text placement instructions (TOP-CENTER and BOLD).
- + Highly accurate isometric perspective and 45-degree angle.
- + Clean, minimalist composition that feels professional and spacious.
- − The 'JAPAN' and 'SUSHI' text is overlaid graphics rather than integrated into the 3D scene.
- − The lighting on the sushi itself is slightly flat compared to Model A.
Verdict: Wan 2.6 followed the complex layout instructions much more accurately, correctly placing the requested text at the top-center with a small flag icon. While DALL-E 3 produced a more charming and richly textured 3D model, it ignored the specific text placement and omitted one of the required words.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
DALL-E 3
- + Features a very warm and vibrant lighting effect with prominent god rays.
- + Excellent soft, fluffy texture on the animals' fur.
- + Whimsical interpretation with cute, stylized butterfly/animal hybrids.
- − Has a very 'AI-cartoonish' or 3D-render look rather than the requested hyper-photorealism.
- − The butterfly hybrids are a creative deviation but not biologically realistic as implied by photorealism.
- − The composition feels a bit crowded and static.
Wan 2.6
- + Achieves a much higher level of realism in the fur, anatomy, and environment.
- + Captures the 'playfully chasing' and 'tumbling' aspect of the prompt with dynamic poses.
- + Includes realistic dew sparkles and a natural-looking golden sunrise.
- − The fox's eyes appear slightly asymmetrical.
- − One or two of the small flying seeds look a bit like digital artifacts.
Verdict: While DALL-E 3 creates a charming, storybook-style illustration with great lighting, Wan 2.6 far better follows the 'hyper-photorealistic' instruction. Wan 2.6 captures the energy of the baby animals actually playing in a meadow with realistic textures and lighting, avoiding the plastic-like look of the DALL-E 3 output.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
DALL-E 3
- + Excellent texture and retro aesthetic
- + Strong vector emblem design with professional layout
- + All requested elements like steam and the date are present
- − Failed to include the primary text 'Caffè Florian', replacing it with 'Coffee House'
Wan 2.6
- + Perfect adherence to all text requirements including 'Caffè Florian'
- + Clean minimalist composition that follows the vector style prompt
- + Correctly features the banner, steam, and date
- − The banner is very small and lacks prominence compared to the main text
- − The background texture is a bit generic compared to the image design
Verdict: DALL-E 3 produced a far more visually compelling and authentic vintage emblem, but it completely missed the main brand name requested in the prompt. Wan 2.6 followed the prompt instructions perfectly, capturing the specific name, the cloche, the banner, and the vintage minimalist aesthetic with high accuracy.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
DALL-E 3
- + Excellent adherence to the color palette and vector style.
- + Includes complex iconography and sequential storytelling as requested.
- + High visual density and aesthetic appeal consistent with space infographics.
- − Displays space shuttle-style crafts which are historically inaccurate for Apollo 11.
- − Text elements are mostly placeholder gibberish.
- − The layout is split into three panels rather than a single unified poster.
Wan 2.6
- + Clean legible text for names and title.
- + Correct use of the navy and red color palette.
- − Failed to include any of the 6 requested infographic steps.
- − Lacks the Saturn V, Moon, and Lunar Module icons requested.
- − Composition is extremely sparse and looks like a tea towel rather than a professional infographic.
Verdict: DALL-E 3 successfully captured the complex layout and stylistic requirements of an infographic, despite some historical inaccuracies in the spacecraft shapes. Wan 2.6 failed almost every prompt instruction, providing a very simple design that lacked the sequence of mission steps and the requested icons.
Explore each model
Alibaba's multimodal generation model from the Wan AI suite, supporting text-to-video, image-to-video, reference-to-video with audio, and text-to-image, in both Chinese and English