OpenAI's legacy image generation model supporting generations, edits with masks (inpainting), and variations
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
DALL-E 2
#59 of 62 in Text-to-Image
GPT Image 1
#32 of 62 in Text-to-Image
Where the votes landed
DALL-E 2
0%
win rate
Ties
0%
GPT Image 1
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
DALL-E 2
- + Features a wooden-textured surface
- + Creative interpretation of light and reflection
- − Fails to follow almost all spatial instructions
- − The blue element is a pot instead of a sphere inside the cube
- − The red element is inside the cube rather than a book on top
- − Image resolution and clarity are poor
GPT Image 1
- + Perfect adherence to all spatial instructions and object descriptions
- + High visual clarity and realistic textures for glass, paper, and foliage
- + Accurate implementation of soft window light from the left
- − The blue sphere appears to be floating slightly rather than resting on the bottom of the cube
Verdict: DALL-E 2 failed to follow the logical and spatial requirements of the prompt, incorrectly placing a red block inside the cube and turning the blue sphere into a large background pot. GPT Image 1 followed every instruction perfectly, producing a crisp, realistic image that maintains the correct relationship between the cube, sphere, book, and plant.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
DALL-E 2
- + Successfully captures the imperfect framing and bokeh effects requested in the prompt.
- + Atmospheric and chaotic feel of a rainy street.
- − Fails to show the subject's face or 'natural skin texture' as requested.
- − Low resolution with significant digital noise and lack of detail on the focal point.
GPT Image 1
- + Excellent adherence to details like 'natural skin texture', 'elderly Japanese man', and 'red bicycle'.
- + High visual quality with realistic rain droplets and lighting.
- + Great composition that balances the subject with the background motion blur.
- − The bike's drivetrain and chain mechanics are slightly nonsensical upon close inspection.
- − The framing is a bit too 'perfect' despite the prompt asking for imperfect framing.
Verdict: GPT Image 1 is superior as it captures nearly all elements of the prompt, including the specific textures and subject description that DALL-E 2 almost entirely ignores. While DALL-E 2 creates an interesting abstract mood, its failure to show the man's face or skin texture makes it a poor match for the technical requirements.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
DALL-E 2
- + Features a very close-up perspective.
- + Includes a visible bokeh effect.
- − Extremely low visual quality with heavy noise and artifacts.
- − Anatomic features are blurry and indistinguishable.
- − Fails to show specific details like braided hair with beads or ornate engraving.
GPT Image 1
- + Exquisite detail on the engraved armor and skin textures.
- + Perfectly captures all prompt elements including braided hair with beads and lifelike eyes.
- + Professional-grade lighting, composition, and realistic textures.
- − The sparks in the background are slightly uniform in distribution.
Verdict: DALL-E 2 produced a low-resolution, messy image that fails to meet almost any of the prompt's descriptive requirements for detail or clarity. GPT Image 1, in contrast, followed every aspect of the prompt with high fidelity, creating a stunning portrait with realistic skin, intricate armor engraving, and beautiful lighting.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
DALL-E 2
- + Strong bold typography that feels modern
- + Vibrant colors in the food imagery
- − The food photos are fragmented into chaotic shapes rather than a menu grid
- − The text is largely gibberish and poorly rendered
- − Does not function as a legible menu layout
GPT Image 1
- + Excellent adherence to the grid layout request
- + High-quality, realistic food photography for all three requested categories
- + Clean, professional typography that mimics a real menu
- + Sections are clearly labeled with vibrant color accents
- − Small minor typos in descriptive subtext (e.g., 'Apperoiation descrigion')
- − The layout is very safe and conventional
Verdict: GPT Image 1 followed every part of the prompt, creating a highly functional and aesthetically pleasing menu layout with clear sections for appetizers, pizza, and mains. In contrast, DALL-E 2 produced a fragmented, abstract design where the food imagery and text were largely incomprehensible as a menu. GPT Image 1 is the clear winner for its superior composition, text rendering, and professional visual quality.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
DALL-E 2
- + Successfully captures a chaotic, fiery atmosphere.
- + Good sense of energy and motion in the burger components.
- − Text is nonsensical and does not follow the prompt.
- − Low visual clarity with messy, distorted textures on the food.
- − Failed to include the price starburst or the secondary message.
GPT Image 1
- + Perfect text rendering for all requested messages including the currency and price.
- + Excellent photorealistic detail on individual burger ingredients.
- + Strong adherence to the 'exploded' layout with a clean, professional composition.
- − Slightly missing the '6' in the price (renders as €.99 instead of €6.99).
Verdict: GPT Image 1 is vastly superior, following all prompt instructions including complex text integration and specific layout requirements. DALL-E 2 fails significantly on text legibility and image clarity, resulting in an abstract and messy visual that doesn't function as a professional ad.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
DALL-E 2
- + Captures a messy, authentic chalk texture.
- − Text is completely illegible gibberish.
- − Fails to follow any of the specific menu item prompts.
- − Extremely low visual quality and resolution.
GPT Image 1
- + Perfect text adherence for all requested menu items and dates.
- + Excellent rendering of chalk texture and handwriting style.
- + Clean, balanced composition that fits the 'cozy café' theme.
- − Handwriting is slightly too uniform, appearing somewhat like a digital font despite the chalk texture.
Verdict: GPT Image 1 followed every specific detail of the prompt, including complex menu items and dates, with perfect legibility and high visual quality. DALL-E 2 produced a low-resolution image with nonsensical text that completely ignored the prompt's content requirements. GPT Image 1 is the clear winner for its superior text rendering and prompt adherence.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
DALL-E 2
- + Successfully captures a dreamy, surreal aesthetic.
- + Creates an interesting texture and mood that feels artistic.
- − Failed the negative constraint; the astronaut is riding the horse.
- − Low image resolution and significant anatomical distortions in the horse's legs and tail.
- − Lacks the high detail and cinematic quality requested.
GPT Image 1
- + Excellent visual quality with high resolution and crisp textures.
- + Cinematic lighting and clear, detailed astronaut suit and horse anatomy.
- − Failed the negative constraint; the astronaut is riding the horse instead of vice versa.
- − The composition is a standard trope rather than the subverted concept requested.
Verdict: Both DALL-E 2 and GPT Image 1 completely failed to adhere to the core instruction 'horse on top, not vice versa,' instead generating standard images of astronauts riding horses. GPT Image 1 is the clear winner because it delivered on the 'highly detailed' and 'cinematic' style requirements, whereas DALL-E 2 produced a low-fidelity, distorted image.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
DALL-E 2
- + Features a close-up of a passenger with a phone
- − Failed to render a recognizable capybara (looks like a distorted statue)
- − Extreme anatomical distortions on the human face
- − Composition is chaotic and does not resemble a taxi interior
GPT Image 1
- + Excellent adherence to all prompt details including the jacket and cap
- + High visual quality with realistic fur and bokeh lighting
- + Perfectly captures the 'bored' expression of the passenger in the background
- − The passenger's phone is slightly blurry due to depths of field
Verdict: GPT Image 1 followed the instructions perfectly, delivering a high-quality, photorealistic image that accurately depicts a capybara driver and a bored passenger. DALL-E 2 failed significantly, producing a distorted, low-quality image where the capybara is unrecognizable and the human figure is nightmarish.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
DALL-E 2
- + Features an intricate, hand-drawn vintage aesthetic
- + Good border details with thorns and twisted elements
- − Text is illegible and contains multiple spelling errors
- − Missing several requested elements like the jack-o-lantern and bats
- − Low image clarity with significant artifacts
GPT Image 1
- + Excellent typography with legible title and banner text
- + Includes all requested elements: jack-o-lantern, bats, and webbed borders
- + High visual quality with moody cinematic lighting and a clear gothic style
- − Minor text error where 'TIME' and 'Location' details were merged into one line
- − The parchment texture is very subtle compared to the requested dark parchment
Verdict: GPT Image 1 is the clear winner as it successfully rendered most of the requested text and all the visual elements including a glowing jack-o-lantern and bats. DALL-E 2 produced an illegible mess of text and failed to include several key components of the prompt, making it unusable as an invitation.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
DALL-E 2
- + Strong isometric perspective and lighting shadows.
- − Failed to include the word 'JAPAN'.
- − Misspelled 'SUSHI' as 'Sush'.
- − The food items are abstract and do not clearly look like sushi.
GPT Image 1
- + Perfect adherence to text requirements, including 'JAPAN', 'SUSHI', and the flag icon.
- + Beautiful 3D cartoon style with soft refined textures as requested.
- + Clean composition on a raised diorama base.
- − The chopsticks are floating slightly oddly off the edge of the plate/base.
Verdict: GPT Image 1 followed every instruction in the prompt, including complex text and icon placement, while maintaining a consistent and appealing 3D miniature aesthetic. DALL-E 2 struggled significantly with the text rendering, misspelling words and omitting others, and failed to produce recognizable sushi.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
DALL-E 2
- + Includes butterflies as requested
- + Bright, sunny lighting
- − Severe anatomical distortions in the kitten and butterfly
- − Poor image quality with significant artifacts and smearing
- − Missing the fox kit and clear bunny features
GPT Image 1
- + Excellent adherence to all prompt elements including four specific animals
- + High visual quality with realistic fur textures and cinematic lighting
- + Perfectly captures the 'god rays' and 'joyful wholesome vibe'
- − The kitten's tail is somewhat lost in the background blending
Verdict: GPT Image 1 significantly outperforms DALL-E 2 by providing a high-fidelity, coherent scene that includes all four requested baby animals with exceptional detail. DALL-E 2 fails to render a recognizable fox or bunny and suffers from major anatomical errors and poor resolution.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
DALL-E 2
- + Matches the requested light background color
- + Conveys a minimalist aesthetic
- − Text consists of illegible gibberish
- − The 'steam' effect is poorly defined and visually messy
- − Fails to include the 'Est. 1720' banner requested
GPT Image 1
- + Perfect text rendering of 'Caffè Florian' and 'Est. 1720'
- + Includes all requested elements like the cloche, steam, and banner
- + Clean, professional vector-style composition
- − Ignored the request for a light background, providing a black one instead
- − The brown-on-black contrast is slightly lower than ideal
Verdict: GPT Image 1 is the superior choice because it accurately renders the specific text and includes every requested design element, such as the banner and the cloche. DALL-E 2 fails significantly on typography and adherence to the specific prompt details, producing nonsensical characters.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
DALL-E 2
- + Captures a complex NASA-style technical blueprint aesthetic
- + Proper color palette is present
- − Text is entirely gibberish and mispelled
- − Fails to provide the specific 6-step icons requested
- − Messy layout with visual noise and low-resolution artifacting
GPT Image 1
- + Excellent adherence to the clean, modern flat-vector style
- + High-quality, legible text and names including the crew
- + Follows the icon requests well (Saturn V, Earth orbit, Translunar arc)
- − Misses some of the step labels at the bottom (EARLLUNAR is a typo)
- − Alignment of labels and icons is slightly disorganized
Verdict: DALL-E 2 fails significantly on text legibility and following the specific multi-step prompt, producing a cluttered and nonsensical graphic. GPT Image 1 follows the stylistic and content instructions very closely, providing clear vector iconography and accurate historical names despite a few minor typos in the footer.
Explore each model
OpenAI's previous image generation model that accepts both text and image inputs and produces image outputs