Head to head
Esc

Models · slot A

to navigate to pick

DALL-E 2 OpenAI Qwen Image Max Alibaba

Settled by community votes across 8 shared challenges, with an AI judge weighing in on each.

DALL-E 2

15.8 arena score

#59 of 62 in Text-to-Image

Skill signature · Text-to-Image

Qwen Image Max

21.3 arena score

#35 of 62 in Text-to-Image

Vote tally

Where the votes landed

DALL-E 2

0%

win rate

Ties

0%

Qwen Image Max

0%

win rate

Shared challenges 8

Challenge by challenge

The strongest take from each model on every shared challenge, with the AI judge's read.

Geometric Composition

Text-to-Image

“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”

DALL-E 2
Qwen Image Max

AI Judge Analysis

DALL-E 2

  • + Handles reflection on the wooden surface realistically.
  • Fails to include the red book on top of the cube.
  • Confusion of shapes and colors; the sphere is missing, replaced by a red core in the cube, and a giant blue pot in the background.
  • Poor resolution and blurry details compared to modern standards.

Qwen Image Max

  • + Perfect adherence to all prompt elements including the sphere inside and the plant behind.
  • + High visual quality with realistic textures on the wood, book, and glass.
  • + Accurate lighting and shadow execution coming from the left window light.
  • The sphere appears to be floating unnaturally rather than resting on the bottom of the cube.

Verdict: DALL-E 2 struggled significantly with spatial relationships and object placement, missing the book entirely and misinterpreting the blue sphere as a background object. Qwen Image Max followed the prompt perfectly, providing a clear, high-resolution image with all requested objects in their correct positions.

Fantasy Warrior

Text-to-Image

“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”

DALL-E 2
Qwen Image Max

AI Judge Analysis

DALL-E 2

  • + Features a macro-like perspective with heavy texture.
  • Extremely low visual coherence, appearing more like a weathered statue or messy clay than a person.
  • Fails to render 'lifelike eyes', 'braided hair', or recognizable 'leather straps'.
  • Highly pixelated and blurry with messy artifacts.

Qwen Image Max

  • + Excellent adherence to all prompt details including braided hair with beads, faint scars, and engraved plate armor.
  • + Superior visual quality with lifelike eyes and realistic skin textures.
  • + Effective lighting and bokeh sparks that create a cinematic atmosphere.
  • The torch in the background is slightly distorted/blobby in shape.

Verdict: DALL-E 2 produced a low-quality, abstract image that failed to capture the human elements of the prompt. Qwen Image Max delivered a high-fidelity, cinematic portrait that perfectly translated complex details like the braided hair, intricate armor engravings, and weather-beaten skin into a cohesive character study.

Magic Burger Explosion: Fiery Photorealism Challenge

Text-to-Image

“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”

DALL-E 2
Qwen Image Max

AI Judge Analysis

DALL-E 2

  • + Successfully captures a chaotic, fiery atmosphere
  • + Includes a glowing effect on the central elements
  • Text is garbled and unreadable (e.g., 'MARGIC BAGUEC')
  • Food looks unappetizing and lacks photorealistic detail
  • Extremely low image resolution and clarity

Qwen Image Max

  • + Perfect text rendering for all requested elements
  • + High-quality photorealistic textures on the food
  • + Excellent layout and adherence to the 'exploded' and 'starburst' requirements
  • The burger is slightly less 'exploded' than typical deconstructed ads, appearing more as a tilted stack

Verdict: Qwen Image Max followed every instruction perfectly, producing a professional-grade advertisement with clear, correctly spelled text and appetizing textures. In contrast, DALL-E 2 produced a low-quality, blurry image with severe text errors and unrecognizable food components.

The Capybara Taxi Driver

Text-to-Image

“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”

DALL-E 2
Qwen Image Max

AI Judge Analysis

DALL-E 2

  • + Attempts the basic elements of the prompt.
  • Extreme anatomical distortion in the human face and hands
  • Capybara looks like a distorted puppet or mask
  • Fails to place the capybara in the driver's seat correctly
  • Lacks the requested photorealism

Qwen Image Max

  • + Excellent photorealistic quality and lighting
  • + Perfect adherence to all prompt details including the cap, jacket, and expression
  • + Strong composition with a clear background city scene and believable interior
  • None notable for this request

Verdict: Qwen Image Max followed every instruction perfectly, delivering a high-quality, photorealistic image with correct anatomy and composition. In contrast, DALL-E 2 produced a nightmare-inducing image with severe artifacts, failed anatomy, and poor spatial reasoning.

The Halloween Invitation

Text-to-Image

“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”

DALL-E 2
Qwen Image Max

AI Judge Analysis

DALL-E 2

  • + Captures a dark, moody parchment texture well
  • + The border has a unique, aged organic aesthetic
  • Text is illegible and contains multiple spelling errors
  • Fails to include key elements like the jack-o-lantern and bats
  • Low resolution and overall muddy visual quality

Qwen Image Max

  • + Perfect text rendering of all requested details and banner text
  • + High visual quality with clear, cinematic lighting
  • + Accurately includes all prompt elements like thorns, webs, and the central jack-o-lantern
  • The font for the event details is a bit plain compared to the title

Verdict: Qwen Image Max is the clear winner as it perfectly followed all instructions, including complex text rendering for the title, banner, and event details. DALL-E 2 produced an illegible image that missed almost all the specific visual requirements of the prompt.

Isometric Miniature Diorama Scenes

Text-to-Image

“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”

DALL-E 2
Qwen Image Max

AI Judge Analysis

DALL-E 2

  • + Clean isometric perspective
  • + Minimalist aesthetic fits the 'diorama base' request
  • Failed to render text correctly, displaying misspelled 'Sush' instead of the requested layout
  • Missing the 'JAPAN' text and flag icon entirely
  • The food items are poorly defined and do not look like sushi

Qwen Image Max

  • + Perfect adherence to text and icon requirements
  • + High-quality 3D rendering with realistic PBR materials on the fish and glass
  • + Excellent miniature diorama composition with a clear raised base
  • The 'JAPAN' text is at the top rather than being integrated into the 3D scene floor as some might interpret, though it fulfills the prompt literally

Verdict: Qwen Image Max followed every instruction in the prompt, including complex text rendering, the flag icon, and high-fidelity 3D modeling of the sushi. DALL-E 2 failed significantly, producing misspelled text, poor image quality, and missing several key elements of the request.

Adorable Baby Animals in Sunny Meadow

Text-to-Image

“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”

DALL-E 2
Qwen Image Max

AI Judge Analysis

DALL-E 2

  • + Features multiple animal types such as a puppy and a kitten.
  • Poor visual quality with blurry textures and significant anatomical distortions.
  • Objects like the butterflies and animal limbs are nonsensical and disconnected.
  • Lacks the requested 'hyper-photorealistic' 8K quality, appearing more like an early AI generation.

Qwen Image Max

  • + Excellent adherence to the 'god rays' and 'golden sunrise' lighting request.
  • + High-fidelity textures on fur, flowers, and butterflies.
  • + Strong composition that captures the 'joyful' and 'tumbling' aspects of the prompt.
  • Includes two golden retriever puppies instead of the specific variety of four different animals requested.
  • Missing the baby bunny and the red fox kit (though one animal has fox-like coloration, it lacks distinctive fox kit features).

Verdict: Qwen Image Max is the clear winner due to its superior image quality, realistic lighting, and intricate detail, even though it failed to perfectly count out the four distinct animal types. DALL-E 2 produced a low-quality, distorted image with significant artifacts that do not meet the '8K masterpiece' requirement. Qwen Image Max achieved.

Vintage Cafe Logo

Text-to-Image

“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”

DALL-E 2
Qwen Image Max

AI Judge Analysis

DALL-E 2

  • + Simple minimalist silhouette of the cloche dome.
  • + Consistent warm brown and cream color palette.
  • Text is completely garbled and illegible.
  • Missing the required steam and banner elements.
  • Lacks the requested 'Est. 1720' text entirely.

Qwen Image Max

  • + Perfect text rendering of 'Caffè Florian' and 'Est. 1720'.
  • + Excellently captures all prompt elements including the banner and steam.
  • + High-quality vector emblem style with appropriate vintage texture.
  • Slightly more complex than 'minimalist' might imply, though fits the 'vintage' aesthetic perfectly.

Verdict: Qwen Image Max followed the instructions perfectly, rendering all text accurately and including every requested element like the banner and steam. In contrast, DALL-E 2 failed to produce legible text and ignored several key parts of the prompt.

Next steps

Explore each model

The Max series of Tongyi Qwen’s image generation model excels across a wide range of generation tasks. Compared with the Plus series, it significantly reduces the “AI-like” feel in generated images, enhancing their realism. It delivers more lifelike material textures for human subjects, finer and more detailed natural textures, and more visually appealing text rendering.