OpenAI's previous generation image model with higher quality than DALL-E 2 and support for larger resolutions
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
DALL-E 3
#40 of 62 in Text-to-Image
Imagen 4.0 Generate 001
#55 of 62 in Text-to-Image
Where the votes landed
DALL-E 3
0%
win rate
Ties
0%
Imagen 4.0 Generate 001
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
DALL-E 3
- + High level of intricate detail within the sphere
- + Effective use of lighting and shadows to create depth
- − Failed the core positional prompt by placing the book INSIDE the cube instead of on top
- − The cube has an opaque wooden lid which contradicts the 'on top' instruction
- − Added complex textures to the sphere not requested
Imagen 4.0 Generate 001
- + Perfect adherence to the spatial relationship of all objects
- + Clean, photorealistic rendering of glass and reflections
- + Accurately depicts the plant behind and through the glass
- − The sphere appears to be floating rather than resting, which may be a physics slight oversight
- − The lighting is a bit flat compared to the more dramatic lighting in image A
Verdict: Imagen 4.0 followed the prompt instructions perfectly, correctly placing the red book on top of the cube and the blue sphere inside it. DALL-E 3 failed the spatial logic of the prompt by putting the red book inside the cube and adding a wooden frame and lid that wasn't requested, making it impossible for a book to sit on top of the 'glass' directly.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
DALL-E 3
- + Excellent atmospheric lighting and puddles reflections
- + Good environmental storytelling with the Japanese lanterns and street depth
- + Creative framing with the foreground bicycle elements
- − The man is barefoot on wet pavement, which feel unrealistic
- − Anatomical issues with the man's neck and fingers
- − The car in the background lacks sufficient motion blur requested in the prompt
Imagen 4.0 Generate 001
- + Highly realistic skin textures and rain droplets on his clothing
- + Accurate representation of an elderly man with appropriate winter attire for rain
- + Effective bokeh and shallow depth of field
- − Missed the 'motion blur' aspect for the passing traffic
- − The bicycle mechanics look slightly distorted near the chain and pedals
- − The composition is a bit more centered than the 'imperfect framing' request implied
Verdict: While DALL-E 3 creates a more cinematic and moody atmosphere with its reflection work, Imagen 4.0 is much more successful at rendering realistic human features and textures without the anatomical distortions found in the other image. Imagen 4.0 also feels more authentic to a real-world scenario by dressing the man appropriately for the weather, whereas DALL-E 3's barefoot subject breaks immersion.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
DALL-E 3
- + Excellent skin texture with visible pores and realistic scabbing in the scars.
- + Sophisticated lighting with strong rim light and realistic metallic reflections.
- + Rich detail on the cloth underlayer and leather straps as requested.
- − Failed to include the specific 'braided hair with small beads' instruction, opting for loose hair.
- − The scars look like symbols or brandings rather than natural battle wounds.
Imagen 4.0 Generate 001
- + Perfect adherence to the hair braiding and bead detailing prompt.
- + Clearer representation of bokeh sparks and a visible torch source.
- + Ornate engraving on the plate armor is crisp and follows the contours of the metal.
- − The skin has a slightly plastic or airbrushed quality compared to Model A.
- − The leather straps are somewhat repetitive and lack the worn texture seen in the alternative.
Verdict: Imagen 4.0 followed more of the specific prompt instructions, particularly the braided hair and beads which DALL-E 3 ignored. However, DALL-E 3 produced a much more photorealistic portrait with superior skin texture and more cinematic lighting. Imagen 4.0 is the winner for total prompt adherence, while DALL-E 3 is visually more convincing as a lifelike photograph.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
DALL-E 3
- + Provides a complete multi-page design concept.
- + Strong use of color blocking and vibrant accents.
- + Good photographic variety of food items.
- − Text is largely illegible and contains various artifacts.
- − The layout feels more like a mood board than a usable menu design.
Imagen 4.0 Generate 001
- + Excellent typography with clean, bold sans-serif fonts.
- + Strict adherence to the grid layout requested.
- + High-quality, professional food photography that fits the minimalist aesthetic.
- − Occasional gibberish in the smaller sub-text items.
- − Layout is a bit rigid compared to a real-world multi-page menu.
Verdict: Imagen 4.0 significantly outperformed DALL-E 3 by delivering a design that looks like a finished professional product. While DALL-E 3 captured the color palette well, its text was unreadable and the layout was messy; Imagen 4.0 followed all prompt instructions regarding layout, sections, and font styles with much higher clarity.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
DALL-E 3
- + Excellent dynamic lighting with high contrast and glow effects.
- + Impressive sense of motion and 'exploded' energy with pieces flying outward.
- + Highly photorealistic textures on the charred patty and fresh vegetables.
- − Multiple spelling errors in the text, including 'MAGIC BURGR' and 'Limiited'.
- − Price tag is poorly integrated and lacks the requested starburst design.
Imagen 4.0 Generate 001
- + Perfect text rendering for all requested phrases including the price.
- + Accurately follows requested design elements like the fiery starburst and specific price.
- + Very clean composition that looks like a professional advertising asset.
- − Lighting is relatively flat compared to the dramatic backlighting requested.
- − The 'exploded' effect feels more static and vertical rather than dynamic and energetic.
Verdict: Imagen 4.0 Generate 001 is the clear winner due to its superior text rendering, accurately capturing all spelling and the specific 'starburst' requirement. While DALL-E 3 produced a more visually stunning and dynamic lighting environment, the numerous typos in the primary text make it unsuitable for an advertisement.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
DALL-E 3
- + Excellent visual atmosphere with natural warm lighting and detailed chalk texture
- + Artistic and elegant composition that fits the 'cozy café' aesthetic
- − Numerous spelling errors including 'Trufle', 'Occtus', and 'Clutee Cokies'
- − The pricing format is confusing and doesn't match the prompt's simplicity
Imagen 4.0 Generate 001
- + Near-perfect adherence to the specific text and menu items requested
- + High legibility for the main menu items and prices
- + Includes the final incomplete prompt item as 'Brown Butter Chocolate Chip Cookies'
- − Includes instructional text from the prompt (like 'Tittle', 'Menu', 'Footer') as actual text on the board
- − The handwriting looks more like a digital font than actual chalk on a board
- − Significant typos in the filler/instructional text
Verdict: Imagen 4.0 Generate 001 followed the text instructions much more accurately, correctly identifying the menu items and pricing despite its robotic literalism in printing the prompt instructions. However, DALL-E 3 produced a far more visually convincing and atmospheric 'chalkboard' image with authentic textures, even though its ability to spell specific complex words was significantly worse.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
DALL-E 3
- + Excellent dreamlike atmosphere with soft galactic lighting.
- + Strong surrealist composition with the horse emerging from a sea of clouds.
- − Failed the negative constraint to have the 'horse on top' of the astronaut.
- − Lower level of technical detail on the spacesuit and horse tack compared to the competitor.
Imagen 4.0 Generate 001
- + High level of intricate detail on the astronaut's suit and horse's mane.
- + Creative inclusion of nebula patterns on the horse's coat.
- − Failed the negative constraint; the astronaut is riding the horse instead of the inverse.
- − The rainbow light streaks feel slightly cluttered and less cinematic than the atmospheric lighting of Model A.
Verdict: Both DALL-E 3 and Imagen 4.0 struggled with the specific spatial constraint 'horse on top, not vice versa,' with both models defaulting to the standard trope of an astronaut riding a horse. Imagen 4.0 provides a much sharper, more detailed image with complex textures, whereas DALL-E 3 offers a more cohesive and painterly surrealist aesthetic.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
DALL-E 3
- + Excellent texture on the capybara's fur and whiskers
- + Atmospheric lighting with realistic interior reflections from the street lights
- + Highly detailed dashboard and taxi interior
- − Failed to include the human passenger requested in the prompt
- − Composition is very tight, focusing almost exclusively on the driver
Imagen 4.0 Generate 001
- + Successfully included all prompt elements, including the human passenger on her phone
- + Great composition showing the relationship between the driver and passenger
- + Clean, sharp rendering of the city lights through the windows
- − The capybara's 'paws' look a bit more like talons or claws than real capybara feet
- − The 'TAXI' sign is floating or appearing inside the car roof erroneously
Verdict: While DALL-E 3 produced a more high-fidelity, atmospheric texture for the capybara itself, it completely failed to include the human passenger specified in the prompt. Imagen 4.0 followed all instructions, including the bored businesswoman and the specific framing, making it the superior choice for prompt adherence despite minor anatomical oddities on the paws.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
DALL-E 3
- + Excellent atmospheric lighting and moody 'vintage gothic' aesthetic.
- + Highly detailed and intricate borders with a realistic parchment texture.
- + Creative use of 3D elements like the jack-o-lantern and hanging bats.
- − Text rendering is poor with several typos and illegible words ('THIGELSTS', 'APARIERY').
- − The requested date and location details are disorganized or missing.
Imagen 4.0 Generate 001
- + Near-perfect text rendering for all requested information, including date and location.
- + Clear and legible typography that fits the gothic theme well.
- + Solid adherence to layout elements like the scroll banner and thorn border.
- − Visual style feels more like a vector illustration than a 'cinematic' or 'vintage' poster.
- − The large vertical parchment roll on the left feels slightly out of place in the composition.
Verdict: While DALL-E 3 creates a much more atmospheric and visually stunning gothic piece, Imagen 4.0 is far superior for this specific task due to its ability to render complex text accurately. Imagen 4.0 followed all instructions for the specific wording, date, and location, making it a functional invitation whereas DALL-E 3's text was mostly gibberish.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
DALL-E 3
- + Excellent adherence to the '3D cartoon scene' and 'isometric' style requested.
- + Precise text rendering of 'JAPAN' on the base.
- + Includes the small flag icon as requested.
- − Failed to place the text at the 'top-center' of the image.
- − Missing the secondary 'SUSHI' text.
- − The rice texture is excessively simplified and bead-like.
Imagen 4.0 Generate 001
- + Photorealistic PBR material quality with highly realistic textures.
- + Accurate sushi anatomy and food rendering.
- + Clean composition on a raised diorama base.
- − Completely ignored the request for text ('JAPAN', 'SUSHI').
- − Ignored the request for a 'cartoon scene' style, opting for realism.
- − Missing the flag icon and solid light blue background.
Verdict: DALL-E 3 followed the stylistic and layout instructions much better, delivering a true isometric cartoon diorama with text and the requested background color, even though it missed half the text. Imagen 4.0 produced a realistic image of high quality but failed almost every specific instruction regarding text, style, and color, making it a poor match for the prompt's requirements.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
DALL-E 3
- + Strong magical atmosphere with prominent god rays and a soft, warm glow.
- + High levels of fluffiness and expressive, large eyes that convey a wholesome vibe.
- + Excellent use of lighting to create depth and focus on the characters.
- − Anatomical surrealism where butterflies have animal faces.
- − The kitten is very small and lacks distinct tabby markings compared to the prompt.
- − Fails the 'photorealistic' requirement in favor of a stylized Pixar-like aesthetic.
Imagen 4.0 Generate 001
- + Successfully includes all four animals with accurate markings, including a clear tabby kitten.
- + Better adherence to 'photorealistic' textures while maintaining the playful action requested.
- + Highly detailed meadow with clear dew sparkles on the grass and flowers.
- − The fox kit's proportions and pose feel slightly stiff or taxidermy-like.
- − The lighting is a bit flat compared to the dramatic 'god rays' seen in the other model.
- − Composition is slightly cluttered with large flowers in the foreground blocking the scene.
Verdict: While DALL-E 3 creates a more emotionally evocative and 'magical' image, it fails on technical details by giving the butterflies animal heads and ignoring the tabby pattern. Imagen 4.0 Generate 001 provides a much more accurate interpretation of the prompt's specific subjects and the requested photorealistic style, despite having slightly less dramatic lighting.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
DALL-E 3
- + Excellent vintage aesthetic with cross-hatching and stippled texture
- + Complex and balanced vector emblem composition
- + Strong adherence to the warm brown and cream color scheme
- − Failed to include the specific name 'Caffè Florian', using generic 'COFFEE HOUSE' instead
- − Typography is slightly cluttered inside the banner
Imagen 4.0 Generate 001
- + Perfect text rendering of the brand name 'Caffè Florian'
- + Clean minimalist vector style that adheres to modern logo standards
- + Accurately represents the cloche with steam and 'Est. 1720' banner
- − Lacks the 'vintage' and 'subtle texture' requested in the prompt
- − Composition is a bit generic compared to the intricate emblem style requested
Verdict: DALL-E 3 captured the vintage texture and emblem style much better, creating a logo that feels authentic to the requested era, but it failed to use the specific brand name. Imagen 4.0 followed the text instructions perfectly and produced a clean, functional logo, though it is less 'vintage' and more 'modern minimalist'. Imagen 4.0 is the overall winner for following the primary branding instruction while still maintaining a cohesive design.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
DALL-E 3
- + Expertly captures the NASA-inspired color palette and vintage vector aesthetic.
- + High level of complexity and visual interest in the layout.
- + Strong stylistic consistency across the elements.
- − Fails on informational structure, providing three separate layouts with garbled text.
- − Includes incorrect spacecraft like the Space Shuttle instead of the Saturn V.
- − Doesn't follow the specific sequential steps requested in the prompt.
Imagen 4.0 Generate 001
- + Excellent adherence to the sequential steps and iconography requested.
- + Legible, crisp text and clean vector style.
- + Accurate representation of the Saturn V rocket for the Apollo mission.
- − Missing the final two steps (Descent and Landing) requested in the prompt.
- − Information logic is a bit confused, such as labeling a Earth icon as 'Launch' and a Moon icon as 'Translunar'.
- − Composition is slightly unbalanced with significant empty space.
Verdict: Model B (Imagen 4.0) is the superior choice because it actually attempts to create an infographic with legible text and the specific iconography requested, despite missing the final two steps. Model A (DALL-E 3) produces a visually stunning artistic poster, but it fails as an infographic by including irrelevant spacecraft like the Space Shuttle and rendering all information-bearing text as illegible symbols.
Explore each model
Google's latest Imagen 4.0 text-to-image generation model with significantly better text rendering and overall image quality