OpenAI's previous generation image model with higher quality than DALL-E 2 and support for larger resolutions
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
DALL-E 3
#42 of 62 in Text-to-Image
GPT Image 2
#3 of 62 in Text-to-Image
Where the votes landed
DALL-E 3
0.0%
win rate
Ties
0.0%
GPT Image 2
100.0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
DALL-E 3
- + High artistic detail with intricate textures.
- + Good use of warm and cool lighting contrast.
- + Visually striking composition.
- − Failed spatial prompt adherence: the book is inside the cube rather than on top.
- − Failed spatial prompt adherence: the sphere is on top of the book rather than just inside the cube.
- − The cube has a wooden frame, which was not requested.
GPT Image 2
- + Perfect adherence to all spatial instructions in the prompt.
- + Realistic glass refraction and depth.
- + Clean, minimalist composition that follows the lighting instructions accurately.
- − Simple textures compared to the other model.
- − The sphere has a very matte finish that looks slightly flat.
Verdict: GPT Image 2 followed every detail of the prompt, correctly placing the book on top of the cube and the sphere inside it, while DALL-E 3 failed the spatial arrangements by placing the book inside the cube. GPT Image 2 also provided a more accurate interpretation of the 'glass cube' without adding an unprompted wooden frame.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
DALL-E 3
- + Excellent atmospheric lighting and cinematic mood.
- + Strong puddles and reflection effects on the pavement.
- + Effective use of 'imperfect framing' with the foreground elements.
- − Anatomical issues with the man's feet (too many toes, strange shape).
- − Skin texture and clothing look somewhat illustrative/digital rather than 'no stylization'.
- − The man is bare-chested and barefoot in a rainy street, which feels unrealistic.
GPT Image 2
- + Photorealistic skin textures and believable modern clothing.
- + Accurate bike mechanics and tool kit detail.
- + Captured the 'motion blur' on passing cars very effectively while maintaining a sharp subject.
- − The 'imperfect framing' is less intentional than in Model A.
- − Rain is barely visible, appearing more like a post-rain dampness than a 'light rain' falling.
- − Large white foreground object on the right is a bit distracting.
Verdict: GPT Image 2 is the superior choice because it adheres more strictly to the 'no stylization' and 'realistic' parts of the prompt, producing a scene that looks like a genuine street photograph. While DALL-E 3 captures a more cinematic mood with better reflections, it suffers from significant anatomical errors and an overly stylized, almost filtered appearance.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
DALL-E 3
- + Excellent depiction of warm torchlight reflecting off the metal surfaces.
- + Highly detailed engraving on the plate armor and texture on the fabric.
- + Strong adherence to the shallow depth of field and bokeh sparks request.
- − Missed the prompt instruction for hair braided with small beads.
- − The scars look slightly stylized rather than realistic.
GPT Image 2
- + Perfectly included the hair braids with small beads as requested.
- + Exceptional skin texture including dirt, freckles, and fine pores.
- + Very realistic and natural-looking eyes.
- − The warm torchlight effect is much more subtle compared to the other image.
- − The armor engraving is less ornate than Model A.
Verdict: While DALL-E 3 captures the dramatic lighting and intricate armor details beautifully, it completely misses the specific 'braids with beads' requirement. GPT Image 2 follows all prompt instructions meticulously, including the hair details and realistic 'battle-worn' textures, while maintaining high visual quality.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
DALL-E 3
- + Provides multiple layout variations in a single image
- + Captures a high-end designer aesthetic with blocks of color and white space
- − Text is nonsensical and contains gibberish characters
- − The food photos in the grid lack clarity and professional appeal
- − Fails to create a functional, readable menu layout
GPT Image 2
- + Excellent text rendering with perfectly legible fonts and specific menu items
- + Strict adherence to the prompt including all three requested sections (Appetizers, Pizza, Mains)
- + High-quality, appetizing food photography that integrates well into the grid
- − Layout is slightly more traditional than 'minimalist' art-style menus
Verdict: GPT Image 2 is the clear winner as it produces a fully functional, professional-grade menu with legible text and high-quality food photography. DALL-E 3 creates a conceptual mood board with layout ideas, but the text is unreadable and it fails to meet the technical requirements of a graphic design task.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
DALL-E 3
- + Excellent photorealistic texture on the burger patty and buns.
- + The composition feels very dynamic with the vertical explosion effect.
- + Clean, professional lighting that highlights the food beautifully.
- − Spelling errors in the text, such as 'MAGC BURGR' and 'Limiited'.
- − The price is contained in a simple box rather than the requested starburst.
- − Text does not feature the requested fiery, glowing effect.
GPT Image 2
- + Perfect adherence to all text requirements, including 'MAGIC BURGER', 'LIMITED TIME ONLY', and the price.
- + Successfully integrated the starburst shape and the fiery glowing effect for all text elements.
- + Highly detailed food rendering with realistic sauce splashes and fresh-looking vegetables.
- − The composition is a bit crowded with large text overlapping the background elements.
- − The angle of the top bun is slightly awkward relative to the rest of the stack.
Verdict: While DALL-E 3 produced a beautiful and clean image, it failed significantly on text spelling and the specific 'fiery text' style requirements. GPT Image 2 followed the prompt's complex text instructions perfectly, including the starburst shape and the glowing fire effect, while maintaining high visual quality.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
DALL-E 3
- + Ornate and artistic composition with beautiful chalk illustrations
- + Strong use of lighting and shadows to create atmosphere
- − Numerous spelling errors including 'Trufle', 'Occtus', and 'Grililled'
- − Prices are nonsensical and text becomes illegible gibberish in several areas
- − Failed to follow specific menu item list requested
GPT Image 2
- + Perfect adherence to the requested text and menu items with zero spelling errors
- + Highly realistic chalk texture that truly looks like authentic handwriting
- + Clean and readable layout that perfectly matches the requested cafe aesthetic
- − Slightly simpler composition compared to the decorative style of Image A
Verdict: GPT Image 2 is the clear winner as it followed the complex text-rendering instructions perfectly, including the specific date and price points. While DALL-E 3 created a more visually ornate image, it failed significantly on prompt adherence by misspelling words and generating illegible text throughout the board.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
DALL-E 3
- + Excellent cinematic lighting and space background
- + Highly detailed ornate armor on the horse
- + Dynamic composition with a sense of motion
- − Completely failed the semantic prompt to put the horse on top of the astronaut
GPT Image 2
- + Perfectly followed the difficult 'horse on top' instruction
- + Included clever details like the horse holding the reins/straps
- + Good texture on the lunar surface and astronaut suit
- − The astronaut hands/gloves are anatomical nightmares with too many fingers
- − The harness connecting the horse to the astronaut is physically confusing
Verdict: While DALL-E 3 produced a far more beautiful and high-quality image, it ignored the specific surreal instruction for the horse to be riding the astronaut. GPT Image 2 successfully translated the literal meaning of the prompt and captured the requested surrealism, making it the winner despite significant anatomical errors in the astronaut's hands.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
DALL-E 3
- + Excellent fur detail and lighting on the capybara.
- + Clear city lights bokeh in the background.
- + Strong adherence to the businesswoman's bored expression.
- − The capybara is wearing a yellow jacket and tie instead of the requested dark jacket.
- − The paws are not visible on the steering wheel.
- − The scale of the capybara relative to the car seat is slightly off.
GPT Image 2
- + Perfect adherence to all prompt details, including the dark jacket and paws on the wheel.
- + Highly realistic taxi interior and perspective from outside the door.
- + The capybara's expression is very calm and professional as requested.
- − The lighting is a bit more muted and less 'cinematic' than Model A.
- − The background city details are a bit more cluttered.
Verdict: GPT Image 2 is the clear winner as it followed every specific instruction in the prompt, including the dark jacket and the positioning of the paws on the steering wheel, which DALL-E 3 failed to do. While DALL-E 3 produced a sharp image, its capybara was wearing the wrong outfit and was not shown driving, whereas GPT Image 2 captured the exact scene requested with high realism.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
DALL-E 3
- + Exquisite ornate framing and 3D depth
- + Unique artistic interpretation of the scroll and thorns
- + Consistent lighting across all elements
- − Text is largely nonsensical and illegible beyond the main header
- − The jack-o-lantern is quite small and lacks impact
- − Failed to include several required text details
GPT Image 2
- + Perfect text rendering for all requested details including dates and locations
- + Excellent composition with a strong central jack-o-lantern
- + Highly detailed background including theNYC 'Arches' bridge and gothic architecture
- − Slightly less 'parchment' texture than requested
- − Text is very clean, bordering on digitally overlaid rather than weathered
Verdict: GPT Image 2 is the clear winner as it successfully rendered all the specific event details requested (Date, Time, Location) with perfect legibility, whereas DALL-E 3 produced mostly gibberish text. GPT Image 2 also provided a much more cinematic and atmospheric composition that felt both spooky and polished.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
DALL-E 3
- + Excellent 3D cartoon aesthetic with soft, pillowy textures.
- + Presents a very clean and cohesive miniature diorama feel.
- + Unique stylized interpretation of sushi ingredients.
- − Failed to place the text 'JAPAN' and 'SUSHI' at the top-center as requested.
- − Missed the 'SUSHI' text entirely.
- − The rice grains look more like foam balls than rice.
GPT Image 2
- + Perfect adherence to text placement and content instructions.
- + Extremely high material realism with realistic PBR textures for the fish and wood.
- + Accurate isometric 45-degree perspective and diorama composition.
- − The sushi variety causes it to look slightly cluttered compared to the 'minimal' request.
- − The 'cartoon' aspect is very subtle, leaning much more towards realism.
Verdict: While DALL-E 3 captured a charming cartoon aesthetic, it failed significantly on the specific text placement and content requirements. GPT Image 2 followed every instruction in the prompt perfectly, including the specific text hierarchy and the inclusion of the flag and diorama base, while still maintaining high visual quality.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
DALL-E 3
- + Successfully captures all four requested animal species in a high-fantasy style.
- + Includes creative, stylized butterflies with furry bodies that add to the whimsical theme.
- + Vibrant lighting with clear 'god rays' effects.
- − The style is heavily CGI/3D animation rather than the requested photorealistic look.
- − The animals look like plush toys rather than real biological creatures.
- − The proportions of the cat and fox are unnaturally small compared to the puppy.
GPT Image 2
- + Achieves a much higher level of realism in animal anatomy and fur texture.
- + Better captures the 'playfully chasing' action described in the prompt.
- + Superior lighting that feels natural while still maintaining the 'magical' sunrise vibe.
- − The fox kit's anatomy looks slightly distorted on its front legs.
- − The 'tumbling together' aspect is more individualistic than a group huddle.
Verdict: GPT Image 2 is the clear winner as it successfully balances the whimsical prompt with a believable, photorealistic aesthetic, whereas DALL-E 3 produced a stylised 3D render that lacks realism. GPT Image 2 also does a better job of conveying motion and movement in the meadow, which aligns with the request for the animals to be chasing and tumbling.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
DALL-E 3
- + Excellent vector illustration style with clean lines
- + High contrast and appealing use of negative space
- + Accurate rendering of the requested 'Est. 1720' date
- − Failed the primary text prompt, displaying 'COFFEE HOUSE' instead of 'Caffè Florian'
- − The cloche looks slightly like a bread basket due to the vertical line pattern
GPT Image 2
- + Perfect adherence to text prompts, including the name and dates
- + Beautiful vintage typography and sophisticated framing
- + Excellent application of subtle texture on a light background
- − The steam trails are slightly less integrated with the cloche than model A
Verdict: GPT Image 2 is the clear winner as it followed all textural and labeling instructions perfectly, including the specific name 'Caffè Florian'. While DALL-E 3 produced a high-quality vector graphic, it failed the core requirement of using the requested restaurant name, defaulting to generic text.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
DALL-E 3
- + Captures a more artistic, vintage 'Swiss-style' graphic design aesthetic
- + Accurate color palette adherence across multiple layout variations
- − Failed to provide a clear 6-step linear infographic as requested
- − Contains many non-NASA-related icons like space shuttles that weren't part of Apollo 11
- − Text is largely illegible gibberish
GPT Image 2
- + Perfectly follows the 6-step logical sequence requested in the prompt
- + Excellent text rendering for headers, steps, and crew names
- + Consistent flat-vector iconography that matches the mission phases accurately
- − The 'Descent' and 'Landing' lunar modules are slightly more complex than a 'minimalist flat vector' style
Verdict: GPT Image 2 is the clear winner as it perfectly adheres to the structural requirements of the instructional prompt, providing a clear 6-step infographic with accurate terminology and legible text. DALL-E 3 produced more abstract artistic posters that failed to follow the specific step-by-step logic and included historically inaccurate icons like the Space Shuttle.
Explore each model
OpenAI's state-of-the-art image generation model with arbitrary resolution up to 4K and strong instruction following