OpenAI's previous generation image model with higher quality than DALL-E 2 and support for larger resolutions
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
DALL-E 3
#40 of 62 in Text-to-Image
Wan 2.5 (Preview)
#27 of 62 in Text-to-Image
Where the votes landed
DALL-E 3
0%
win rate
Ties
0%
Wan 2.5 (Preview)
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
DALL-E 3
- + High visual detail in the book and sphere textures
- + Artistic lighting and depth of field
- − Failed the spatial instructions by placing the book inside the cube instead of on top
- − The cube has a thick wooden frame not requested in the prompt
Wan 2.5 (Preview)
- + Perfect adherence to all spatial instructions
- + Accurately rendered the red book sitting on top of the glass cube
- + Correctly depicted the plant behind and visible through the glass
- − Lighting is a bit harsh with some lens flare/dust artifacts
- − The reflections on the table are slightly inconsistent with the glass base
Verdict: Wan 2.5 (Preview) followed the prompt with 100% accuracy, correctly positioning the book on top of the cube and the sphere inside. DALL-E 3 failed the spatial logic by putting the book inside the cube and adding a wooden frame that wasn't requested.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
DALL-E 3
- + Excellent atmosphere with strong cinematic lighting
- + Creative use of foreground framing to create depth
- + Strong reflection details on the wet pavement
- − Anatomical errors in the man's feet and skin texture
- − High level of digital stylization despite prompt requesting none
- − The bicycle geometry is distorted and physically impossible
Wan 2.5 (Preview)
- + Natural skin texture and convincing facial detail
- + Highly accurate bicycle anatomy and tools
- + Strong adherence to the 'no stylization' and 'natural' prompt requirements
- − Lacks the requested motion blur from passing cars
- − The rain effect appears somewhat static and uniform
- − Composition is a bit safe compared to the 'imperfect framing' request
Verdict: Wan 2.5 (Preview) produces a far more realistic image with convincing skin textures and a physically accurate bicycle, whereas DALL-E 3 suffers from significant anatomical distortions and a plastic, AI-stylized look. Wan 2.5 is the winner for its superior technical execution and grounded, lifelike quality that better matches the request for a non-stylized street photo.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
DALL-E 3
- + Excellent high-contrast dramatic lighting with glowing bokeh sparks.
- + Superior engraving details on the plate armor that looks aged and worn.
- + Strong skin texture and lifelike eye detail.
- − Failed to include the requested braided hair with beads.
- − The scars look like surface paint or superficial marks rather than healed wounds.
Wan 2.5 (Preview)
- + Perfectly captured the technical hair requirement including braids and small silver beads.
- + Realistic dirt and grime application on the face.
- + Expertly rendered textile textures including the wool underlayer, leather straps, and chainmail.
- − The lighting is a bit flat across the face compared to the dramatic armor lighting.
- − The character looks a bit young for a 'battle-worn' descriptor.
Verdict: Wan 2.5 (Preview) and DALL-E 3 both produced high-quality images, but Wan 2.5 followed the specific prompt instructions much more closely by including the braided hair and beads which DALL-E 3 ignored. While DALL-E 3 has more dramatic cinematic lighting, Wan 2.5 excels in texture variety with its detailed rendering of chainmail, leather, and frayed cloth.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
DALL-E 3
- + Features a comprehensive grid layout with multiple menu variations.
- + Captures a realistic photography style for the food items.
- + Includes realistic shadows and textures that give it a professional mockup feel.
- − The text is largely illegible and uses nonsensical symbols.
- − The composition feels more like a portfolio spread than a functional menu design.
Wan 2.5 (Preview)
- + Excellent adherence to the grid layout and section prompts (Appetizers/Pizza/Mains).
- + Higher legibility with bold, clean sans-serif fonts for the headers.
- + Better use of vibrant accents and a clean, minimalist white background.
- − Some spelling errors in titles such as 'Menue' and 'Mainns'.
- − Food plating is somewhat repetitive across different categories.
Verdict: Wan 2.5 (Preview) produced a much more functional and aesthetically accurate response to the prompt, successfully incorporating the requested menu sections and minimalist design language. While DALL-E 3 created a decent mock-up style image, its layout is chaotic and the text is entirely gibberish, whereas Wan 2.5 (Preview) provided a clear, usable structural design.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
DALL-E 3
- + Excellent photorealistic texture on the meat patty and bun
- + Dynamic ground-burst fire effect adds a strong sense of scale
- + Vibrant colors and glowing rim lighting
- − Multiple spelling errors in the text including 'AGIC BURGR' and 'Limiited'
- − The price is inside a box rather than the requested starburst
Wan 2.5 (Preview)
- + Perfect adherence to all text requirements with zero spelling errors
- + Accurately rendered price in a fiery starburst as requested
- + Clean composition with a sophisticated melting-sauce effect on the typography
- − The lighting on the bottom bun feels slightly disconnected from the glowing elements
- − Less intense background atmosphere compared to Model A
Verdict: While DALL-E 3 (Model A) delivers more dramatic lighting and highly detailed textures, it fails significantly on text legibility and spelling. Wan 2.5 (Model B) followed all instructions perfectly, including the specific starburst requirement and complex typography, making it much more suitable for a professional ad design.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
DALL-E 3
- + Excellent chalk texture and artistic flourishes
- + Warm, ambient lighting consistent with a 'cozy café'
- − Numerous spelling errors including 'Trufle', 'Occtus', and 'Riototo'
- − Text layout is cluttered and contains gibberish characters
- − The price for the first item is rendered as a nonsensical $234
Wan 2.5 (Preview)
- + Near-perfect spelling of all complex menu items
- + Followed handwriting style instructions exactly with consistent slant and stroke
- + Clear, legible layout with realistic chalk smudge effects
- − The 'cursive' for the title is more of a casual print-script hybrid than elegant cursive
- − The perspective shift makes the bottom right of the board feel slightly compressed
Verdict: Wan 2.5 (Preview) significantly outperforms DALL-E 3 by successfully rendering the specific text requested with high accuracy and legibility. While DALL-E 3 captures a more 'artistic' atmosphere, its failure to spell the menu items correctly and the inclusion of nonsensical text makes it less effective for the prompt's requirements.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
DALL-E 3
- + Features a beautiful, ethereal galaxy backdrop with cinematic lighting.
- + Captures a dreamlike, surreal atmosphere with the horse moving through clouds in space.
- + Composition is balanced and follows the rule of thirds effectively.
- − Completely failed the negative constraint to have the horse on top of the astronaut.
- − The horse lacks fine muscle detail and appears somewhat flat.
Wan 2.5 (Preview)
- + High level of technical detail on the astronaut's suit and the horse's anatomy.
- + Sharp focus and vibrant colors create a very cinematic feel.
- + Good rendering of textures including the horse's mane and the space helmet reflections.
- − Completely failed the negative constraint to have the horse on top of the astronaut.
- − The lower part of the horse's legs are slightly blurry/deformed compared to the rest of the image.
Verdict: Both DALL-E 3 and Wan 2.5 failed the specific negative constraint to have the 'horse on top' of the astronaut, both defaulting to the standard 'astronaut on horse' trope. Wan 2.5 is the preferred image due to its superior clarity, high-resolution details on the space suit, and more realistic anatomical rendering of the horse compared to the softer, more stylized output of DALL-E 3.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
DALL-E 3
- + Excellent fur detail and lighting coherence
- + Realistic car interior texture
- + Strong side-profile composition that feels cinematic
- − Completely missed the passenger in the back seat
- − The capybara's hand/paw looks more like a human hand in hair than a paw
Wan 2.5 (Preview)
- + Successfully included all prompt elements, including the bored passenger
- + High level of photorealism and detail on the rainy exterior
- + Correct yellow taxi driver cap as requested
- − The paws on the steering wheel look slightly like human fingers
- − Composition is a bit crowded with the taxi sign blocking the view
Verdict: Wan 2.5 (Preview) and DALL-E 3 both produced high-quality images, but Wan 2.5 is the clear winner for following all instructions, specifically including the bored businesswoman in the back. DALL-E 3 failed to include the passenger entirely, whereas Wan 2.5 captured the requested humor and narrative perfectly.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
DALL-E 3
- + Excellent gothic atmosphere and dark parchment aesthetic
- + Intricate 3D-like border with high texture quality
- + moody cinematic lighting consistent with the prompt
- − Failed to render all requested text accurately
- − Text is largely gibberish and missing specific location details
- − Overall composition is a bit cluttered at the bottom
Wan 2.5 (Preview)
- + Perfect text adherence with clearly readable dates and location
- + Strong execution of the scroll banner and gothic title
- + Excellent balance of all prompt elements including thorns and clouds
- − Lighting is a bit vibrant and less 'vintage' than requested
- − Composition is somewhat generic compared to the first image
Verdict: While DALL-E 3 captures a more authentic 'vintage gothic' mood, it fails significantly on the text rendering requirements. Wan 2.5 (Preview) follows every instructional detail accurately, including the specific date, location, and secondary banner text, making it much more useful as a functional invitation.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
DALL-E 3
- + Excellent 3D isometric perspective on a diorama base.
- + Highly detailed textures for the salmon and rice grains.
- + Captures the 'miniature' aesthetic with complex surrounding garnishes.
- − Failed to place 'JAPAN' and 'SUSHI' at the top-center as requested.
- − The word 'SUSHI' is missing entirely.
- − Included multiple flags and objects not specified in the prompt.
Wan 2.5 (Preview)
- + Perfect adherence to text placement and content instructions.
- + Clean, professional graphic design and typography.
- + Accurately represents the 45-degree angle on a circular diorama base.
- − The lighting is a bit more diffused/flat compared to Model A.
- − The rice grains look more like rounded pellets than realistic sushi rice.
Verdict: While DALL-E 3 (Model A) produced a more visually intricate 3D model with better textures, it failed significantly on the layout requirements for text and titles. Wan 2.5 (Model B) followed the prompt instructions precisely, including all specific text elements and their requested positions, making it the better choice for this specific design brief.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
DALL-E 3
- + Features very soft, illustrative fur texture that fits the 'fluffy' descriptor.
- + Excellent use of god rays and warm golden lighting to create a magical atmosphere.
- + Includes charming, imaginative details like butterflies with soft, fuzzy bodies.
- − Lean heavily into a 3D animation/stylized aesthetic rather than 'hyper-photorealistic'.
- − The anatomy of the kitten is somewhat distorted and stylized.
Wan 2.5 (Preview)
- + Delivers a much more photorealistic look as requested in the prompt.
- + Captures the sense of movement and 'tumbling' more effectively through dynamic poses.
- + Accurately represents all four specific animals with realistic anatomical features.
- − The dew sparkles are rendered as floating, gravity-defying water droplets.
- − The fox's eyes have a slightly unnatural, glowing blue ring artifact.
Verdict: While DALL-E 3 creates a beautiful, whimsical scene, it ignores the 'photorealistic' requirement in favor of a stylized Pixar-like aesthetic. Wan 2.5 (Preview) better adheres to the prompt by providing a scene that looks like real photography, capturing the motion and species accuracy more effectively despite minor artifacts in the dew drops.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
DALL-E 3
- + Excellent vector emblem layout with intricate vintage detailing
- + Perfectly captures the requested 'warm brown and cream' color palette with subtle textures
- + High visual quality and professional balance
- − Completely failed to use the specified 'Caffè Florian' text, substituting it with 'Coffee House'
Wan 2.5 (Preview)
- + Successfully rendered the specific text 'Caffè Florian' correctly
- + Included the banner for 'Est. 1720' as requested
- + Clean and more minimalist vector look
- − The cloche dome is transparent, which is somewhat unconventional for that specific heraldic vessel
- − The crumple texture on the paper background feels a bit cheap compared to the illustration style
Verdict: DALL-E 3 produced a far more aesthetically pleasing and authentic vintage logo emblem, but it completely ignored the specific brand name in the prompt. Wan 2.5 (Preview) followed all text instructions perfectly, including the exact brand name and date banner, making it the more functional choice despite having a slightly simpler artistic execution.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
DALL-E 3
- + Strong artistic aesthetic with a cohesive vintage poster feel
- + Effective use of the NASA-inspired color palette
- + High level of graphical complexity and texture
- − Fails to follow the sequential 6-step logical flow
- − Includes incorrect spacecraft like the Space Shuttle
- − Text is entirely illegible gibberish
Wan 2.5 (Preview)
- + Perfect adherence to the 6-step infographic structure
- + Legible and accurate text including names and mission phases
- + Clean vector style that matches the prompt's request for functional icons
- − Iconography for the spacecraft is slightly repetitive
- − Composition is a bit sparse in the 'Descent' and 'Landing' area
Verdict: While DALL-E 3 (Image A) creates a visually striking set of posters, it fails completely as an infographic, including irrelevant imagery like the Space Shuttle and scrambled text. Wan 2.5 (Image B) successfully follows the technical requirements of the prompt, providing a clear, logical progression of the Apollo 11 mission with legible labels and correct iconography.
Explore each model
Alibaba's text-to-image and image-to-image generation model from the Wan AI suite, offering high-quality visual generation capabilities