Head to head
Esc

Models · slot A

to navigate to pick

DALL-E 3 OpenAI GPT Image 1 OpenAI

Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.

DALL-E 3

20.1 arena score

#42 of 62 in Text-to-Image

Skill signature · Text-to-Image

GPT Image 1

23.2 arena score

#29 of 62 in Text-to-Image

Vote tally

Where the votes landed

DALL-E 3

0%

win rate

Ties

0%

GPT Image 1

0%

win rate

Shared challenges 13

Challenge by challenge

The strongest take from each model on every shared challenge, with the AI judge's read.

Geometric Composition

Text-to-Image

“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”

DALL-E 3
GPT Image 1

AI Judge Analysis

DALL-E 3

  • + High artistic detail with intricate wood grain and interior micro-landscape
  • + Beautiful cinematic lighting and depth of field
  • + High resolution textures on the book and wooden frame
  • Failed to place the red book on top of the cube, placing it inside instead
  • The 'sphere' is complex rather than a simple 'small blue sphere'
  • The cube is more of a wooden frame than a pure glass cube

GPT Image 1

  • + Perfect adherence to spatial instructions (book on top, sphere inside)
  • + Accurately represents every object in the correct color and position
  • + Realistic lighting and shadows consistent with the window on the left
  • The glass cube has some minor structural clipping with the book
  • The sphere appears to be floating rather than resting on the bottom of the cube

Verdict: GPT Image 1 followed the complex spatial instructions of the prompt perfectly, placing the red book on top of the glass cube and the blue sphere inside. DALL-E 3 produced a more visually striking and detailed image, but failed significantly on the prompt adherence by putting the book inside the cube.

Candid Street Photography

Text-to-Image

“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”

DALL-E 3
GPT Image 1

AI Judge Analysis

DALL-E 3

  • + Excellent reflection on the wet pavement mimicking the subject's pose
  • + Effective use of the 'imperfect framing' prompt with foreground objects
  • + Strong atmosphere with cinematic lighting and environmental details
  • Anatomical issues with the man's feet appearing elongated and distorted
  • The car in the background lacks the requested motion blur, appearing static
  • Skin texture appears slightly painterly rather than 'natural'

GPT Image 1

  • + Outstanding natural skin texture and facial realism
  • + Realistic lighting and rain droplets visible on the bicycle seat
  • + More authentic 50mm lens feel with a credible shallow depth of field
  • Lacks the requested motion blur from passing cars
  • The bicycle chain and wheel structure are physically nonsensical/merged
  • Framing is very conventional rather than the 'imperfect' requested framing

Verdict: GPT Image 1 produces a significantly more realistic portrait with superior skin textures and lighting, capturing the rainy atmosphere more convincingly. DALL-E 3 follows the 'imperfect framing' and 'pavement reflection' prompts better, but suffers from significant anatomical distortions in the man's feet and a less realistic overall finish. GPT Image 1 is the winner for its photographic fidelity, despite the structural errors in the bicycle's mechanical parts.

Fantasy Warrior

Text-to-Image

“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”

DALL-E 3
GPT Image 1

AI Judge Analysis

DALL-E 3

  • + Excellent high-contrast lighting and warm torchlight reflections
  • + Intricate high-detail engraving on the plate armor
  • + Powerful close-up composition with vibrant colors
  • Missed the prompt requirement for braided hair with beads
  • Skin texture looks slightly plasticky or digitally smoothed

GPT Image 1

  • + Perfect adherence to specific details like braided hair with small beads
  • + Extremely realistic skin texture with lifelike dirt and faint scars
  • + Strong sense of 'battle-worn' character through facial expression and grime
  • Armor engravings are more subtle and less 'ornate' than requested
  • Overall image is a bit darker, making it harder to see the detailed leather straps

Verdict: While DALL-E 3 produces a more striking, high-fantasy aesthetic with beautiful lighting on the armor, GPT Image 1 is the superior choice for prompt adherence and realism. GPT Image 1 successfully included the specific hair beads and braided hair requested, and the texture of the skin is much more lifelike and detailed than the smoother finish in DALL-E 3.

Modern Clean Menu

Text-to-Image

“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”

DALL-E 3
GPT Image 1

AI Judge Analysis

DALL-E 3

  • + Provides a comprehensive four-page layout concept
  • + Uses a grid-based approach with varied photography
  • + Good use of vibrant color blocking/accents
  • Text is largely unintelligible gibberish
  • The 'grid' is somewhat chaotic and cluttered on some pages

GPT Image 1

  • + Clean, readable sans-serif typography with legible pricing
  • + High-quality, appetizing food photography
  • + Perfectly adheres to the 'modern minimalist' and 'white background' brief
  • Missing the 'mains' section specifically requested in the labels (though food is shown)
  • Only shows a single-page view rather than a full menu design concept

Verdict: DALL-E 3 provides a broader look at a full menu layout but suffers from total text illegibility and a slightly cluttered grid. GPT Image 1 follows the minimalist prompt much better, producing a clean, professional-looking design with readable text and high-quality food photography that feels ready for a casual dining environment.

Magic Burger Explosion: Fiery Photorealism Challenge

Text-to-Image

“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”

DALL-E 3
GPT Image 1

AI Judge Analysis

DALL-E 3

  • + Excellent dynamic lighting with high-contrast fiery effects
  • + Complex exploded view with many varied ingredients
  • + Great sense of ground impact and flying debris
  • Multiple spelling errors in the text ('MAGIC BURGR', 'Limiited')
  • The price tag lacks the starburst shape and the specific value of '€6.99' requested

GPT Image 1

  • + Perfect text accuracy for all components including the price and secondary message
  • + Successfully renders the fiery, glowing text effect as requested
  • + Clean and photorealistic burger textures
  • The composition is quite static compared to the 'dynamic' request
  • Background is somewhat plain with only minor embers
  • Missing the sense of motion and 'exploded' energy found in the other model

Verdict: While DALL-E 3 captures the 'dynamic' and 'exploded' energy much more effectively, it fails significantly on text accuracy and the specific price request. GPT Image 1 follows all textual instructions perfectly and produces a high-quality product shot, even if it feels slightly less energetic. GPT Image 1 is the preferred choice for an advertisement where correct spelling and pricing are critical.

Chalkboard Menu

Text-to-Image

“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”

DALL-E 3
GPT Image 1

AI Judge Analysis

DALL-E 3

  • + Excellent artistic flair with decorative chalk illustrations.
  • + Realistic lighting and environmental context of a cozy café.
  • Numerous spelling errors like 'Trufle', 'Occtus', and 'Grililled'.
  • The text layout is cluttered and includes nonsensical numbers and words.

GPT Image 1

  • + Perfect text rendering with zero spelling errors for all requested items.
  • + The chalk texture is highly realistic, showing convincing grain and pressure variations.
  • + Adheres strictly to the handwritten requirement without using digital-looking fonts.
  • The composition is a bit plain and lacks the decorative elements requested by 'elegant cursive'.
  • The bottom edge of the board cuts off the final line of text.

Verdict: While DALL-E 3 creates a more visually interesting café atmosphere, it fails significantly on text legibility and spelling. GPT Image 1 (DALL-E 3) follows the prompt perfectly regarding the specific menu content and maintains a very high-quality, realistic chalk handwriting texture without any spelling mistakes.

The Reversed Rodeo

Text-to-Image

“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”

DALL-E 3
GPT Image 1

AI Judge Analysis

DALL-E 3

  • + Excellent cinematic lighting and atmosphere
  • + The addition of clouds and the Milky Way creates a beautiful surreal aesthetic
  • + Clear and sharp rendering of the astronaut suit
  • Failed the negative constraint; the astronaut is on top of the horse
  • The horse's legs and hooves look somewhat distorted and unnatural

GPT Image 1

  • + High resolution and realistic texture on the horse's coat
  • + Clean composition with good detail on the astronaut's gear
  • + Stronger anatomical accuracy for the horse compared to Model A
  • Failed the negative constraint; the astronaut is on top of the horse
  • The surreal element is less creative, adhering to a very standard 'space horse' trope

Verdict: Both models failed to follow the specific spatial instruction to place the 'horse on top' of the astronaut, instead defaulting to the common trope of an astronaut riding a horse. DALL-E 3 produced a much more cinematic and visually interesting environment, whereas GPT Image 1 offered a more grounded and anatomically correct subject but lacked the requested surreal flair.

The Capybara Taxi Driver

Text-to-Image

“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”

DALL-E 3
GPT Image 1

AI Judge Analysis

DALL-E 3

  • + Excellent interior detail on the dashboard and upholstery
  • + Great textures on the capybara's fur and clothing
  • + Interesting lighting with the interior ceiling lamp
  • Failed to include the human passenger requested in the prompt
  • The perspective is more from the passenger's seat than a full scene including the back

GPT Image 1

  • + Successfully followed all instructions including the bored passenger
  • + High level of photorealism with shallow depth of field
  • + Natural cinematic lighting that fits the nighttime city setting
  • The 'paws' look slightly like human fingers in gloves
  • Interior detail of the car is less defined than the other model

Verdict: While DALL-E 3 produced a sharp image with great interior textures, it completely missed the requirement for a passenger in the back seat. GPT Image 1 captured the full narrative of the prompt, including the indifferent businesswoman and the specific 'taxi' branding on the cap, making it the superior choice for prompt adherence.

The Halloween Invitation

Text-to-Image

“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”

DALL-E 3
GPT Image 1

AI Judge Analysis

DALL-E 3

  • + Ornate and intricate gothic border design
  • + Highly stylized 3D-depth and cinematic lighting
  • + Creative interpretation of twisted trees and thorns integrated into the frame
  • Text is largely illegible gibberish at the bottom
  • Failed to include the specific prompt text correctly, creating words like 'THIGELSTS'

GPT Image 1

  • + Excellent text legibility and accuracy for most of the prompt
  • + Clearer layout that serves as a functional invitation
  • + Mood and lighting better match the vintage gothic request
  • Merged the 'Time' and 'Location' fields into a single confusing line
  • The border is a bit thin and less prominent than requested

Verdict: GPT Image 1 is the superior choice because it successfully renders the requested text in a readable format, which is essential for an invitation, whereas DALL-E 3 produces nonsensical characters. While DALL-E 3 captures a more complex artistic composition, GPT Image 1 follows the specific instructions for the scroll banner and event details much more effectively.

Isometric Miniature Diorama Scenes

Text-to-Image

“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”

DALL-E 3
GPT Image 1

AI Judge Analysis

DALL-E 3

  • + Excellent 3D toy-like aesthetic with high-gloss finish
  • + Very clean isometric perspective and lighting
  • + Creative integration of text on the side of the diorama base
  • Failed to place text at 'top-center' as requested
  • Missions the word 'SUSHI' in the text elements
  • The sushi design is repeated/monolithic rather than a varied dish

GPT Image 1

  • + Perfect adherence to text placement instructions (top-center, dual lines)
  • + Includes the small flag icon exactly where requested
  • + Features a more realistic variety of sushi types (nigiri and maki)
  • Texture is a bit more matte/clay-like than the 'refined PBR' look of Model A
  • The lighting is slightly darker and less vibrant than Model A

Verdict: While DALL-E 3 (Image A) produces a more visually stunning 3D render with beautiful lighting, GPT Image 1 (Image B) is the clear winner for prompt adherence. Image B followed every layout instruction perfectly, including the specific text positioning and header hierarchy, whereas Image A integrated the text as a label on the model itself and missed half of the required words.

Adorable Baby Animals in Sunny Meadow

Text-to-Image

“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”

DALL-E 3
GPT Image 1

AI Judge Analysis

DALL-E 3

  • + Successfully includes all four requested animals: puppy, kitten, bunny, and fox kit.
  • + Effective use of god rays and cinematic lighting to create a dreamlike atmosphere.
  • + Captures a very expressive and whimsical style that aligns with 'big expressive eyes'.
  • The 'butterflies' have strange mammalian bodies and wings, which looks unnatural and eerie.
  • The style leans heavily toward digital illustration rather than the requested 'hyper-photorealistic' scene.
  • The fox kit has unusual black markings on its paws that look like hearts, which is slightly cartoonish.

GPT Image 1

  • + Excellent adherence to the 'hyper-photorealistic' requirement with realistic fur and anatomical structures.
  • + Successfully captures the action of 'playfully chasing' and 'tumbling' with dynamic poses.
  • + Natural-looking lighting and environment that feels like a real meadow.
  • The bunny is missing its back legs or has them merged into the cat's anatomy in an anatomical error.
  • The kitten's tail is missing or obscured in a way that makes the body look slightly truncated.
  • The butterflies are less prominent and lack the 'ultra-detailed' feel requested compared to the animals.

Verdict: GPT Image 1 is the superior choice because it much more closely adheres to the 'hyper-photorealistic' style requested, whereas DALL-E 3 produced a digital painting/stylized illustration. While DALL-E 3 managed to fit all animals in a clear portrait, GPT Image 1 captured the actual movement and joyful energy of the prompt with far more realistic textures and lighting.

Vintage Cafe Logo

Text-to-Image

“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”

DALL-E 3
GPT Image 1

AI Judge Analysis

DALL-E 3

  • + Excellent visual appeal with high-quality stipple texture
  • + More complex and professional composition
  • + Perfectly centered and balanced design elements
  • Failed to include the primary text 'Caffè Florian', substituting it with 'Coffee House'
  • The design is quite busy for the 'minimalist' portion of the prompt

GPT Image 1

  • + Successfully included all requested text including 'Caffè Florian'
  • + Much better adherence to the 'minimalist' and 'vector emblem' style
  • + Correct placement of the 'Est. 1720' text on a banner
  • The black background contradicts the 'light background' request
  • The graphics are somewhat simple compared to Model A
  • The steam effect is a bit thin and less stylized

Verdict: Model B (GPT Image 1) followed the text prompt much more accurately, correctly rendering the name 'Caffè Florian' and the banner for the date, whereas Model A (DALL-E 3) substituted the main text. Although Model A has a more polished aesthetic with better textures, Model B's adherence to the minimalist logo concept makes it a more functional design for the specific request.

Apollo 11: Journey to Tranquility

Text-to-Image

“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”

DALL-E 3
GPT Image 1

AI Judge Analysis

DALL-E 3

  • + Strong artistic aesthetic with a vintage space-age feel.
  • + Sophisticated use of the requested color palette.
  • + High visual complexity and detailed textures within the moons.
  • Failed to follow the step-by-step instruction, instead creating a repetitive layout.
  • Inaccurate iconography, including depicting Space Shuttles which were not part of Apollo 11.
  • Text is mostly illegible gibberish.

GPT Image 1

  • + Excellent adherence to the specific 6-step infographic structure.
  • + Correct iconography including the Saturn V, Lunar Module, and Earth/Moon orbits.
  • + Clean, modern vector style with legible and mostly accurate text.
  • Composition is a bit cluttered and lacks the flow of a professional poster.
  • Minor spelling error in 'EARLLUNAR'.
  • The flat-vector style is somewhat simplistic compared to Model A.

Verdict: While Model A creates a more visually striking poster, it fails significantly on prompt adherence by including Space Shuttles and ignoring the requested 6-step logic. Model B follows the instructions perfectly, providing the specific icons and sequence requested in a clear, legible infographic format.

Next steps

Explore each model