Head to head
Esc

Models · slot A

to navigate to pick

DALL-E 2 OpenAI GPT Image 2 OpenAI

Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.

DALL-E 2

15.8 arena score

#59 of 62 in Text-to-Image

Skill signature · Text-to-Image

GPT Image 2

27.7 arena score

#4 of 62 in Text-to-Image

Vote tally

Where the votes landed

DALL-E 2

0.0%

win rate

Ties

0.0%

GPT Image 2

100.0%

win rate

0.0% 0.0% ties 100.0%
Shared challenges 13

Challenge by challenge

The strongest take from each model on every shared challenge, with the AI judge's read.

Geometric Composition

Text-to-Image

“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”

DALL-E 2
GPT Image 2

AI Judge Analysis

DALL-E 2

  • + Features a wooden table with reflections.
  • + Includes the requested colors in the scene.
  • Failed to render a red book on top of the cube.
  • The blue sphere is rendered as a giant blue pot in the background.
  • Poor prompt adherence regarding spatial relationships.

GPT Image 2

  • + Excellent prompt adherence, placing every object exactly as described.
  • + High visual quality with realistic textures on the book and glass.
  • + Accurate lighting and refraction through the glass cube.
  • None observed; the image perfectly matches the requested scene.

Verdict: DALL-E 2 failed to follow the spatial instructions, merging the red book into the cube and turning the blue sphere into a background object. GPT Image 2 followed the prompt perfectly, rendering a high-quality, realistic scene with correct object placement and lighting.

Candid Street Photography

Text-to-Image

“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”

DALL-E 2
GPT Image 2

AI Judge Analysis

DALL-E 2

  • + Strong bokeh and shallow depth of field reflecting the lens request
  • + Realistic reflections in the wet pavement
  • Subject is almost completely out of focus, failing to show the man or the act of repairing
  • Lacks any sense of subject detail or 'natural skin texture' requested

GPT Image 2

  • + Excellent adherence to all prompt details including age, ethnicity, and activity
  • + Accurately rendered motion blur on passing cars and realistic wet ground textures
  • The 'imperfect framing' request resulted in some distracting foreground elements like the partial sign
  • Slightly less shallow depth of field than requested for a 50mm lens

Verdict: GPT Image 2 is the clear winner as it successfully interprets all aspects of the complex prompt, providing a clear subject with natural textures while maintaining the requested atmosphere. DALL-E 2 fails the challenge by producing an image where the primary subject is so blurry it is unrecognizable, ignoring the core narrative of the prompt.

Fantasy Warrior

Text-to-Image

“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”

DALL-E 2
GPT Image 2

AI Judge Analysis

DALL-E 2

  • + Features a very shallow depth of field as requested
  • + Captures an extremely weathered, gritty texture
  • Anatomical failure where the face and helmet are incoherent and distorted
  • Low resolution and lacks fine details like beads or eyes
  • Composition is cramped and confusing

GPT Image 2

  • + Excellent adherence to all prompt details including braided hair with beads, scars, and ornate engraving
  • + Highly lifelike eyes and realistic skin texture with dirt
  • + Superior lighting and composition with clear bokeh sparks in the background
  • The 'faint scars' are a bit subtle, appearing more like dirt or freckles

Verdict: DALL-E 2 produced a low-quality, incoherent image that fails to represent most of the prompt's specific details. In contrast, GPT Image 2 delivered a professional-grade portrait that perfectly executed every requirement from the leather straps to the beaded braids, with exceptional visual clarity.

Modern Clean Menu

Text-to-Image

“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”

DALL-E 2
GPT Image 2

AI Judge Analysis

DALL-E 2

  • + Strong bold sans-serif typography
  • + High contrast color palette
  • Nonsense text and illegible characters
  • Food images are abstract and unappealing slices rather than professional photography
  • Fails to create distinct sections for appetizers, pizza, and mains

GPT Image 2

  • + Excellent adherence to all prompt instructions including specific category sections
  • + High-quality, professional food photography in a clean grid layout
  • + Perfectly legible text, pricing, and realistic menu items
  • Slightly more 'commercial' than 'minimalist' but still very clean

Verdict: DALL-E 2 produced an abstract, virtually unusable design with illegible text and confusing food imagery. In contrast, GPT Image 2 followed every detail of the prompt perfectly, delivering a professional-grade, functional menu design with clear sections, beautiful food photos, and perfect typography.

Magic Burger Explosion: Fiery Photorealism Challenge

Text-to-Image

“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”

DALL-E 2
GPT Image 2
0% wins 0% ties 100% wins

AI Judge Analysis

DALL-E 2

  • + Successfully captures a dynamic, fiery aesthetic or mood.
  • + Elements are suspended as requested.
  • Text is nonsensical and does not follow the prompt's requirements.
  • Overall image quality is blurry and lacks photorealistic detail.
  • Component rendering is abstract and doesn't look like appetizing food.

GPT Image 2

  • + Excellent text rendering that perfectly matches the fiery/glowing prompt requirements.
  • + High level of photorealistic detail in the food textures like the patty and lettuce.
  • + Precise adherence to all prompt instructions, including the specific price in a starburst.
  • The composition is a bit crowded with all the secondary text elements.
  • The bottom bun seems to be dripping sauce upwards which is physically inconsistent even for an 'exploded' view.

Verdict: GPT Image 2 is the clear winner as it follows every instruction in the prompt, including the complex text and pricing requirements. While DALL-E 2 produces a blurry and indecipherable image with garbled text, GPT Image 2 creates a professional-looking advertisement with crisp details and vibrant colors.

Chalkboard Menu

Text-to-Image

“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”

DALL-E 2
GPT Image 2

AI Judge Analysis

DALL-E 2

  • + Attempts a chalk-like texture.
  • Text consists of unintelligible gibberish.
  • Fails to follow any of the specific menu item instructions.
  • Poor resolution and artistic quality.

GPT Image 2

  • + Perfect text rendering of the exact prompt requirements.
  • + Exceptional chalk texture and handwriting style with natural variations.
  • + High-quality composition including café background elements and chalk holder.
  • None identified for this specific prompt.

Verdict: DALL-E 2 fails significantly, producing illegible 'vaguely-text-shaped' marks that do not follow the prompt's content. GPT Image 2 achieves near-perfect results, rendering complex specific text with authentic chalk textures and a highly realistic cozy café atmosphere.

The Reversed Rodeo

Text-to-Image

“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”

DALL-E 2
GPT Image 2
0% wins 0% ties 100% wins

AI Judge Analysis

DALL-E 2

  • + Features a classic painterly and hazy aesthetic.
  • + Correctly places an astronaut and a horse in space.
  • Fails the prompt instructions by putting the astronaut on top of the horse.
  • Low resolution with significant digital noise and lack of detail.

GPT Image 2

  • + Perfectly adheres to the specific inversion instruction with the horse riding the astronaut.
  • + High visual quality with sharp details, realistic textures, and a clear NASA logo.
  • + Strong composition with a surreal and humorous interpretation.
  • The horse's hooves hold reins in a way that is anatomically impossible, though consistent with the surreal prompt.

Verdict: DALL-E 2 completely failed the core logical challenge of the prompt, providing a standard 'astronaut on horse' image with low visual fidelity. GPT Image 2 successfully followed the difficult 'horse on top' instruction while delivering a high-resolution, cinematic, and surreal image that matches all requested criteria.

The Capybara Taxi Driver

Text-to-Image

“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”

DALL-E 2
GPT Image 2

AI Judge Analysis

DALL-E 2

  • The image is completely irrelevant to the prompt.
  • The output depicts a black leather bag instead of a taxi scene.
  • Fatal failure in prompt adherence.

GPT Image 2

  • + Excellent adherence to all prompt details including the capybara's outfit and the passenger's expression.
  • + High-quality photorealistic rendering with realistic lighting from outside the car.
  • + Perfectly captures the composition of a professional driver and a bored passenger.
  • The capybara's front paw on the right has a slightly anatomical blur where it meets the wheel.
  • Small digital artifact on the passenger's phone interface.

Verdict: DALL-E 2 completely failed the prompt, providing an image of a black handbag that has no relation to the request. GPT Image 2 successfully generated a high-quality, humorous, and photorealistic scene that perfectly matches every detail of the taxi driver capybara and the passenger.

The Halloween Invitation

Text-to-Image

“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”

DALL-E 2
GPT Image 2

AI Judge Analysis

DALL-E 2

  • + Successfully captures a weathered, antique paper aesthetic.
  • + Conveys a dark, moody atmosphere with high-contrast lighting.
  • Text is largely illegible with numerous misspellings and gibberish characters.
  • Missing several required elements including the jack-o-lantern and bats.
  • Image quality is blurry with unclear, messy borders.

GPT Image 2

  • + Excellent text rendering with near-perfect accuracy for all requested details.
  • + Highly detailed composition including every prompt element like thorns, webs, and bats.
  • + Superior visual clarity and professional-grade cinematic lighting.
  • The 'Arches' background element looks more like a bridge than a specific venue, though still thematic.

Verdict: GPT Image 2 is the clear winner as it followed every instruction, including specific text strings and diverse imagery elements like the glowing jack-o-lantern and gothic border. DALL-E 2 produced a very low-quality image with nonsensical text and failed to include the primary visual motifs requested.

Isometric Miniature Diorama Scenes

Text-to-Image

“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”

DALL-E 2
GPT Image 2

AI Judge Analysis

DALL-E 2

  • + Features a clean solid blue background as requested.
  • + Attempts an isometric perspective with clear shadows.
  • Failed significantly on text rendering, displaying 'Sush' instead of the requested text.
  • Missing 'JAPAN' text and the flag icon.
  • The 3D models are very basic and do not look like high-quality sushi or PBR materials.

GPT Image 2

  • + Excellent adherence to all prompt elements including text, flag, and diorama base.
  • + High-quality 3D renders with realistic textures on the sushi and wood.
  • + Perfect composition and layout following the 45° top-down isometric request.
  • None identified; it accurately followed every specific detail of the prompt.

Verdict: GPT Image 2 followed the prompt perfectly, including the specific text placement, the flag icon, and the isometric diorama style. In contrast, DALL-E 2 failed to include most of the text elements and produced low-quality, abstract results that barely resemble the intended subject.

Adorable Baby Animals in Sunny Meadow

Text-to-Image

“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”

DALL-E 2
GPT Image 2

AI Judge Analysis

DALL-E 2

  • + Includes a golden retriever puppy and kitten
  • + Vibrant green meadow colors
  • Severely lacks photorealism with significant artifacts
  • Anatomically incorrect and distorted animals
  • Fails to include the fox kit clearly and features broken butterfly wings

GPT Image 2

  • + Exceptional photorealism and fur texture
  • + Perfectly includes all requested animals (puppy, kitten, bunny, fox)
  • + Beautifully rendered lighting with god rays and dew sparkles
  • One butterfly appears slightly disconnected from the environment
  • A few flowers in the extreme foreground have soft blurring edges

Verdict: GPT Image 2 significantly outperforms DALL-E 2 in every category, providing a stunningly realistic and heartwarming scene that perfectly adheres to the complex prompt. While DALL-E 2 produces a distorted and low-quality image with severe anatomical errors, GPT Image 2 delivers sharp 8K quality, expressive animal faces, and masterful golden hour lighting.

Vintage Cafe Logo

Text-to-Image

“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”

DALL-E 2
GPT Image 2
0% wins 0% ties 100% wins

AI Judge Analysis

DALL-E 2

  • + Simple minimalist icon approach.
  • + Follows the requested color palette.
  • Text is complete gibberish and unreadable.
  • Cloche dome is poorly rendered with strange artifacts.
  • Failed to include the requested banner and date clearly.

GPT Image 2

  • + Perfect text rendering for both name and date.
  • + Excellent use of texture and shading to create a vintage aesthetic.
  • + Expertly composed layout with the banner and cloche as requested.
  • Slightly more complex than a 'minimalist' prompt might imply but fits the 'vintage' style perfectly.

Verdict: GPT Image 2 followed every aspect of the prompt, delivering high-quality typography and a professional vintage layout. DALL-E 2 failed significantly on the text and produced a messy, incoherent graphic that did not incorporate the requested elements like the banner.

Apollo 11: Journey to Tranquility

Text-to-Image

“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”

DALL-E 2
GPT Image 2

AI Judge Analysis

DALL-E 2

  • + Adheres to the color palette requested.
  • + Captures a complex, abstract tech-diagram aesthetic.
  • Text is completely illegible and nonsensical.
  • Fails to show the specific 6-step chronological sequence requested.
  • Visuals are chaotic and do not look like a modern vector infographic.

GPT Image 2

  • + Follows the 6-step chronological sequence perfectly with matching icons.
  • + Excellent text rendering for headers, steps, and names.
  • + Clean, professional flat-vector design that matches the NASA aesthetic.
  • Very minor spelling error on 'Tranquility' (labeled as 'Tranquility' but the pin design is slightly cramped).
  • The eagle in the top right logo is a bit more detailed than the 'flat' style of the rest of the image.

Verdict: GPT Image 2 followed the prompt instructions near-perfectly, creating a structured 6-step infographic with high-quality vector illustrations and legible text. DALL-E 2 failed significantly, producing garbled text and an abstract layout that ignores the requested historical steps.

Next steps

Explore each model