OpenAI's legacy image generation model supporting generations, edits with masks (inpainting), and variations
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
DALL-E 2
#59 of 62 in Text-to-Image
GPT Image 1 Mini
#13 of 62 in Text-to-Image
Where the votes landed
DALL-E 2
0.0%
win rate
Ties
0.0%
GPT Image 1 Mini
100.0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
DALL-E 2
- + Features a small glass cube with interesting light refraction.
- + Captures the wooden table texture well.
- − Fails almost all prompt instructions regarding object placement and identity.
- − Confuses the blue sphere with a massive blue background object.
- − The red book is depicted as a red core inside the cube rather than on top.
GPT Image 1 Mini
- + Follows all prompt instructions perfectly, including spatial relationships.
- + High visual clarity and realistic lighting representing the 'soft window light' from the left.
- + Excellent material rendering of glass, paper, and wood.
- − The plant is slightly out of focus, though this meets the 'partially visible' requirement.
Verdict: DALL-E 2 completely failed the spatial reasoning and object identification of the prompt, blending the colors into a single abstract object. GPT Image 1 Mini followed every instruction accurately, producing a coherent and high-quality scene that perfectly matches the requested composition.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
DALL-E 2
- + Matches the 'imperfect framing' requirement well
- + Captures a very realistic 50mm shallow depth of field effect
- + Good wet pavement reflections
- − Subject is almost completely out of focus/blurred
- − Fails to clearly show a 'Japanese man' due to focus issues
- − Low resolution and grainy texture
GPT Image 1 Mini
- + Excellent adherence to the 'elderly Japanese man' subject with natural skin texture
- + Highly detailed rendering of the wet bicycle and rain drops
- + Strong cinematic composition while remaining realistic
- − The bokeh in the background is static; lacks the requested 'motion blur from passing cars'
- − Framing feels a bit too perfect and centered despite the prompt's request for 'imperfect framing'
Verdict: GPT Image 1 Mini is the clear winner as it provides a coherent, high-quality image that satisfies the core subject requirements (man, red bicycle, rain). While DALL-E 2 attempted to lean into the 'imperfect photography' aspects of the prompt, it resulted in a muddy, unrecognizable subject, whereas GPT Image 1 Mini delivered professional-grade cinematic realism.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
DALL-E 2
- + Successfully captures a tight macro-style focus.
- + Uses high-contrast lighting that conveys a gritty atmosphere.
- − Severely lacks image coherence and clarity, appearing as a messy collage of textures.
- − Failed to render identifiable eyes, braided hair, or recognizable armor features.
- − Composition is confusing and lacks a clear focal point.
GPT Image 1 Mini
- + Excellent adherence to all prompt details including braided hair, scars, and ornate engraving.
- + High visual quality with realistic skin textures and lifelike eyes.
- + Beautiful lighting and bokeh that create a cinematic depth of field.
- − The 'small beads' requested in the hair are subtle to the point of being nearly invisible.
Verdict: GPT Image 1 Mini outperformed DALL-E 2 in every metric, delivering a clear and detailed portrait that followed the complex prompt precisely. DALL-E 2 failed to produce a coherent image, resulting in a distorted mess of textures that lacked a recognizable human face or the specific requests like braided hair.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
DALL-E 2
- + Strong artistic interpretation of the 'minimalist' and 'bold' keywords.
- + Creates a premium, high-fashion editorial feel.
- − Text is nonsensical gibberish.
- − The food photos are fragmented and abstracted rather than displayed in a clear grid for a functional menu.
- − Fails to clearly define the requested sections for Appetizers, Pizza, and Mains.
GPT Image 1 Mini
- + Perfect adherence to the prompt's layout requirements, including specific menu sections.
- + Text is correctly spelled and legible.
- + Clean, professional grid layout that actually functions as a menu template.
- − The design is a bit generic and lacks artistic flair.
- − The spacing on the left side feels slightly empty without filler text.
Verdict: GPT Image 1 Mini is the clear winner as it followed all instructions, including the specific categories of food and the grid layout. While DALL-E 2 produced a visually interesting aesthetic, it failed the functional requirements of the prompt and produced illegible text, whereas GPT Image 1 Mini created a usable, professional-looking menu design.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
DALL-E 2
- + Energetic sense of motion
- + Good use of embers and fire colors
- − Text is garbled and contains spelling errors like 'MARGIC' and 'BAGUEC'
- − Image lacks photorealism and looks more like a messy painting
- − Missing the price and starburst required by the prompt
GPT Image 1 Mini
- + Excellent adherence to all layout and text requirements
- + High photorealistic quality with clear suspended ingredients
- + Perfectly rendered glowing text and price starburst
- − Slightly static composition compared to a true 'exploded' view
Verdict: GPT Image 1 Mini followed every instruction in the prompt, including complex text and specific price formatting, with high clarity. DALL-E 2 produced a messy, unrecognizable burger with misspelled and missing text. GPT Image 1 Mini is the clear winner for its professional advertising aesthetic and accurate prompt adherence.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
DALL-E 2
- + Captures a messy, authentic chalk texture
- − Text is completely illegible and gibberish
- − Fails to follow specific menu item instructions
- − The image is low resolution and poorly framed
GPT Image 1 Mini
- + Excellent text rendering with 100% accuracy to the prompt
- + Captures a realistic chalk grain texture on the letters
- + Clean and balanced composition with a professional wooden frame
- − The handwriting style is a bit too uniform, leaning towards a digital font feel despite the grain
Verdict: GPT Image 1 Mini followed the prompt perfectly, rendering the specific menu items and date with perfect spelling and legibility. In contrast, DALL-E 2 produced an illegible mess of pseudo-letters that failed every requirement of the text-to-image challenge.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
DALL-E 2
- + Features a grainy, vintage cinematic aesthetic
- − Failed the negative constraint; the astronaut is riding the horse
- − Anatomical distortions in the horse's legs
- − Low resolution and blurry details
GPT Image 1 Mini
- + High visual clarity and atmospheric lighting
- + Detailed textures on the spacesuit and horse's coat
- − Failed the negative constraint; the astronaut is riding the horse
- − The harness/reins are visually floating or disconnected
Verdict: Both DALL-E 2 and GPT Image 1 Mini failed to follow the specific spatial instruction to place the 'horse on top' of the astronaut, instead providing the common interpretation of an astronaut riding a horse. GPT Image 1 Mini is the winner due to significantly higher visual quality, better composition, and sharper details compared to the distorted and blurry output of DALL-E 2.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
DALL-E 2
- − Failed completely to follow the prompt.
- − Generated an image of a black handbag instead of a taxi scene.
- − Irrelevant to the user request.
GPT Image 1 Mini
- + Excellent adherence to all prompt details including the capybara, businessman, and setting.
- + High visual quality with realistic textures on the fur and leather jacket.
- + Successfully captured the lighting and atmosphere of a New York taxi at night.
- − The passenger's hand holding the phone looks slightly distorted.
- − The capybara has its right paw on the wheel but the left is less clearly visible.
Verdict: DALL-E 2 suffered a total failure, generating a close-up of a black handbag that has no relation to the prompt provided. GPT Image 1 Mini followed the complex prompt near-perfectly, delivering a high-quality, cinematic image that accurately captures the surreal concept with photorealistic detail.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
DALL-E 2
- + Strong vintage parchment texture
- + Atmospheric and chaotic border design
- − Text is illegible and contains gibberish
- − Lacks clear jack-o-lantern and required event details
- − Poor visual clarity
GPT Image 1 Mini
- + Perfect text rendering of all requested details
- + Accurate depiction of jack-o-lantern, bats, and thorns
- + Clean and professional composition
- − Slightly less 'gritty' texture than the prompt's mention of dark parchment might imply
- − Central lighting is a bit standard
Verdict: GPT Image 1 Mini followed the prompt instructions perfectly, including specific text strings and event details that were completely mangled in DALL-E 2. While DALL-E 2 captures a more 'ancient' feel, GPT Image 1 Mini provides a functional, high-quality, and aesthetically pleasing invitation.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
DALL-E 2
- + The lighting creates sharp, interesting shadows.
- + The background color matches the requested solid light blue.
- − Failed to include 'JAPAN' text and the flag icon.
- − Misspelled 'SUSHI' as 'Sush'.
- − The 3D model of the sushi is very primitive and lacks the requested refined textures.
GPT Image 1 Mini
- + Perfect adherence to all text and icon requirements in the requested layout.
- + Excellent 3D miniature aesthetic with soft, high-quality textures.
- + Composition is perfectly centered and balanced on a diorama-style base.
- − None notable; it followed every specific detail of the prompt accurately.
Verdict: GPT Image 1 Mini followed the prompt with near-perfect accuracy, delivering high-quality 3D renders of the sushi and including all requested text and icons. DALL-E 2 failed on the typography, spelling, and general visual quality, producing a very sparse and unattractive scene.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
DALL-E 2
- + Natural lighting on the puppy's fur
- + Good sense of movement
- − Severe anatomical distortions and blurred subjects
- − Failed to render distinct animal types clearly
GPT Image 1 Mini
- + Excellent clarity and adherence to all animal types
- + Superb golden hour lighting with god rays
- − Somewhat repetitive poses for the animals
Verdict: GPT Image 1 Mini captured the prompt's complexity perfectly, rendering all four requested animals with high fidelity and beautiful lighting. DALL-E 2 struggled with the multiple subjects, resulting in significant artifacts and a lack of coherent detail.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
DALL-E 2
- + Matches the light background requested in the prompt.
- + Minimalist aesthetic is preserved.
- − Text is completely garbled and nonsensical.
- − The cloche dome and steam are poorly defined and look like ink blots.
- − Failed to include the 'Est. 1720' banner.
GPT Image 1 Mini
- + Perfect text rendering for both 'Caffè Florian' and 'Est. 1720'.
- + Excellent vector-style illustration of the cloche dome with clear steam lines.
- + Good use of texture and vintage typography.
- − Ignored the request for a light background, utilizing a black background instead.
- − The banner style is a bit busy compared to the 'minimalist' request.
Verdict: GPT Image 1 Mini is the clear winner because it successfully rendered all the requested text correctly and created a cohesive, professional-looking emblem. While DALL-E 2 followed the background color instructions, its complete failure to generate legible text or a recognizable logo icon makes it unusable.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
DALL-E 2
- + Matches the requested color palette well.
- − Text consists entirely of illegible gibberish including the title.
- − The layout is chaotic and fails to follow the 6-step infographic structure.
- − Image is cropped poorly and lacks clear icons for the specific mission stages.
GPT Image 1 Mini
- + Perfectly follows the requested 6-step infographic structure with matching icons.
- + Text is highly legible and correctly spells mission-critical words.
- + Adheres strictly to the flat-vector style and NASA-inspired color theme.
- − The 'Translunar' icon is slightly abstract compared to the others.
- − The rocket icon for 'Launch' is a bit simplified.
Verdict: GPT Image 1 Mini followed the prompt instructions near-perfectly, creating a structured, legible, and professional infographic that accurately portrays all six requested mission stages. DALL-E 2 failed significantly, producing an image with nonsensical text and a disorganized layout that did not deliver on the specific list of steps requested.
Explore each model
OpenAI's cost-effective image generation model for when image quality isn't the top priority