OpenAI's legacy image generation model supporting generations, edits with masks (inpainting), and variations
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
DALL-E 2
#59 of 62 in Text-to-Image
OmniGen v2
#57 of 62 in Text-to-Image
Where the votes landed
DALL-E 2
0%
win rate
Ties
0%
OmniGen v2
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
DALL-E 2
- + Features a wooden table texture
- + Includes a glass cube
- − Fails to include the blue sphere inside the cube
- − Missed the red book entirely
- − The large blue object in the background is not part of the prompt
- − Composition is poor with extreme out-of-focus elements
OmniGen v2
- + Perfect adherence to all spatial instructions in the prompt
- + High visual clarity and realistic refractions in the glass
- + Correct soft window lighting from the left
- − The sphere appears to be floating slightly, lacking a distinct contact shadow
Verdict: OmniGen v2 followed every detail of the complex spatial prompt perfectly, placing the sphere inside the cube and the book on top. DALL-E 2 failed significantly, missing multiple objects and including a large blue pot that was not requested. OmniGen v2 is the clear winner for its superior prompt adherence and high image quality.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
DALL-E 2
- + Successfully captures an 'imperfect framing' and 'candid' feel
- + Excellent wet pavement reflections
- + Follows the shallow depth of field request
- − The subject's face is completely obscured, failing the 'natural skin texture' check
- − The image is too blurry to see the 'repairing' action clearly
OmniGen v2
- + Clearly depicts an elderly Japanese man with a red bicycle
- + Includes visible rain and background motion blur from cars
- + Good reflection and high image clarity
- − The man is just standing with the bike rather than 'repairing' it
- − The image looks heavily staged and stylized, ignoring the 'no stylization' and 'imperfect framing' prompts
Verdict: DALL-E 2 attempted a more artistic, candid photography style that met the framing constraints, but unfortunately blurred out the primary subject entirely. OmniGen v2 produced a much clearer, high-quality image that captured more prompt keywords (man, red bike, rain), though it failed to show the man actually repairing the bike and looked a bit too clean/artificial compared to the requested candid aesthetic.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
DALL-E 2
- + Strong tactile texture on the armor surfaces
- + Effective close-up side profile composition
- − Extremely noisy/grainy visual quality with significant artifacts
- − Fails to clearly show hair braids, beads, or lifelike eyes
- − Anatomical clarity is poor
OmniGen v2
- + Excellent adherence to all prompt details including braids with beads, scars, and dirt
- + High visual clarity and lifelike eye rendering
- + Beautiful lighting with bokeh sparks and ornate engraving
- − The character appears a bit too 'clean' for specifically being 'battle-worn'
- − Skin texture is slightly smoothed compared to the armor
Verdict: OmniGen v2 followed the prompt with high precision, successfully rendering details like the braids with beads and the lifelike eyes which DALL-E 2 completely missed. DALL-E 2 produced a low-resolution image with heavy distortion and failed to capture the primary human features requested.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
DALL-E 2
- + Features very bold and aesthetically pleasing modern typography
- + High-contrast visual style aligns with a minimalist 'artsy' brand
- − The 'grid' of food photos is overly abstract and fragmented, making the food unrecognizable
- − Does not follow the layout requirements for specific sections like appetizers or mains
- − The text is entirely gibberish and poorly scaled
OmniGen v2
- + Accurately creates a professional multi-section layout with columns and headers
- + Food photography is realistic, vibrant, and presented in a clear grid format
- + Provides visual accents and a clean white background as requested
- − Several spelling errors in the headers (e.g., 'APPTETIZES', 'PIZZZZAN')
- − The font used for headers is slightly stylized rather than a standard clean sans-serif
Verdict: OmniGen v2 is the clear winner as it creates a functional and recognizable menu layout that adheres to all prompt requirements, including specific food sections and a grid layout. DALL-E 2 produces an abstract, artistic image that fails the structural requirements of a restaurant menu.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
DALL-E 2
- + Successfully captures the exploded/deconstructed burger layout
- + Creates a dynamic sense of motion with suspended particles
- + Atmospheric glowing embers effect matches the prompt well
- − Text is heavily garbled and contains spelling errors like 'MARGIC' and 'BAGUEC'
- − Food items lack photorealistic detail and appear melted or poorly defined
- − Missing the starburst element and specific price formatting
OmniGen v2
- + Excellent text rendering for 'MAGIC BURGER' and the starburst price
- + High visual clarity and photorealistic textures on the lettuce and bun
- + Clean composition that functions well as a professional ad
- − Fails the 'exploded' instruction by showing a fully assembled burger
- − Lacks the sense of motion and suspended components requested
- − Some secondary text like 'LIMITED TIME' is slightly cut off at the edge
Verdict: DALL-E 2 followed the complex layout instructions for an exploded burger but failed significantly on image quality and text accuracy. OmniGen v2 produced a much higher quality, professional-looking advertisement with legible text, even though it ignored the 'exploded' instruction in favor of a standard assembled burger.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI judge analysis unavailable for this challenge.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
DALL-E 2
- + Successfully captures a cinematic, gritty space aesthetic
- + Effective use of scale with the planet background
- − Failed the negative constraint; the astronaut is riding the horse
- − The anatomy of the horse legs and the astronaut's form is muddy and distorted
- − Low overall image resolution and clarity
OmniGen v2
- + High image clarity and clean vectors
- + Good rendering of textures on the horse's mane and the space suit
- − Failed the crucial negative constraint; the astronaut is riding the horse
- − The composition is very static and lacks a 'cinematic' feel
- − The extra moons in the background look like clip-art
Verdict: Both DALL-E 2 and OmniGen v2 failed the difficult spatial reasoning prompt 'horse on top, not vice versa,' instead defaulting to the common trope of an astronaut riding a horse. OmniGen v2 produced a much sharper and cleaner image, but DALL-E 2 better captured the 'cinematic' and 'surreal' mood requested, despite its lower technical quality.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
DALL-E 2
- + Attempts a unique cinematic perspective from the dashboard.
- − Extreme anatomical distortions on the human face and hands.
- − The animal does not resemble a capybara and looks like a distorted statue.
- − Failed to follow the prompt regarding the yellow cap and professional expression.
OmniGen v2
- + Successfully renders a realistic capybara wearing a yellow taxi cap and dark jacket.
- + High visual clarity and photorealistic textures in the animal and car.
- + The businesswoman has a bored, normal expression as requested.
- − The capybara's hands holding the steering wheel are human hands instead of front paws.
- − The human passenger is seated in the front passenger seat instead of the back seat.
Verdict: DALL-E 2 produced a highly distorted image with severe artifacts and failed significantly on the capybara's appearance. OmniGen v2 successfully created a coherent, photorealistic scene that followed the majority of the prompt, including the specific attire and city background.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
DALL-E 2
- + Successfully captures a weathered, vintage parchment texture
- + Moody lighting and gothic frame provide a strong atmospheric feel
- − Text is largely unintelligible and includes multiple spelling errors
- − Fails to include a central glowing jack-o-lantern
- − Image is blurry with low level of detail
OmniGen v2
- + Excellent adherence to all prompt elements including the jack-o-lantern, bats, and trees
- + High visual clarity and polished, cinematic lighting
- + Most requested text is legible and correctly placed
- − Text contains minor typos such as 'Friets', '7pvm', and 'The Arcas'
- − The scroll banner text is garbled compared to the secondary text below it
Verdict: OmniGen v2 produced a far superior result by following every detail of the prompt, including specific imagery like the jack-o-lantern and bats which DALL-E 2 omitted. While OmniGen v2 has some minor spelling errors in the fine print, DALL-E 2's output is completely illegible and lacks the requested central subject, making it unusable as an invitation.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
DALL-E 2
- + Features a 3D isometric perspective.
- + Maintains a clean aesthetic with a light blue background.
- − Failed significantly on text rendering, displaying 'Sush' instead of 'JAPAN' and 'SUSHI'.
- − Low visual fidelity with artifacts in the food models.
- − Missing the flag icon and the small raised diorama base.
OmniGen v2
- + Excellent adherence to the text prompt requirements, including both 'JAPAN' and 'SUSHI' with a flag icon.
- + High visual quality with clean textures and a clear 3D cartoon style.
- + Perfectly captures the isometric diorama base requested.
- − The flag icon does not accurately represent the Japanese flag.
- − Minor repetitive pattern on the rice grains.
Verdict: OmniGen v2 is the clear winner as it followed every instruction in the prompt, including the specific text and layout requirements. DALL-E 2 failed to render the correct text and provided a lower quality, less coherent image.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
DALL-E 2
- + Attempts a more realistic, non-stylized photographic approach
- + Captures an active sense of movement and 'tumbling' as requested
- − Severe anatomical distortions and artifacts, particularly on the kitten and butterflies
- − Missing several of the requested animals
- − Poor image clarity and resolution
OmniGen v2
- + Successfully includes the golden sunrise lighting and god rays
- + High visual clarity and clear subject definitions
- + Cheerful composition with expressive eyes as requested
- − Looks like a digital illustration or 3D render rather than the requested 'photorealistic' scene
- − Combines animal features (the fox and bunny appear merged into one stylized creature)
- − Missing the fourth distinct animal (fox kit or bunny)
Verdict: DALL-E 2 struggles significantly with image coherence, producing many artifacts and anatomical errors, though it attempts a more realistic texture. OmniGen v2 produces a much cleaner, more aesthetically pleasing image that captures the lighting and mood well, even though it leans heavily into a stylized, cartoonish look and fails to render four distinct animals. OmniGen v2 is the preferred output due to its significantly higher quality and lack of disturbing visual glitches.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
DALL-E 2
- + Matches the requested warm brown and cream color scheme
- + Captures a rough, distressed texture consistently
- − Text is nonsensical and does not follow the prompt
- − Logo elements like the cloche and steam are poorly defined and messy
- − Fails to include the 'Est. 1720' banner
OmniGen v2
- + Follows the composition requests with a cloche, steam, and banner
- + Text is legible and largely accurate, including the 'Est. 1720' date
- + Clean vector-style execution with appropriate vintage typography
- − Minor spelling error in the name ('CAFFFLORIN' instead of 'Caffè Florian')
- − The texture is very subtle, bordering on a flat digital fill
Verdict: OmniGen v2 performed significantly better by following the specific layout instructions, including the cloche and banner while providing legible text. DALL-E 2 failed to render the correct text or a coherent logo design, resulting in a disorganized jumble of letters and shapes.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
DALL-E 2
- + Matches the 'navy, white, muted red' color palette effectively.
- + Contains a large amount of complex vector-style line work.
- − Text is completely illegible and nonsensical ('ALLPOO APPLOO').
- − The composition is cluttered and does not follow the requested 6-step chronological structure.
- − Fails to include specific recognizable icons like the Saturn V or Earth orbit rings.
OmniGen v2
- + Follows the clean, modern infographic layout much more effectively.
- + Successfully groups icons and labels into a legible grid structure.
- + Adheres better to the requested color palette with clear white and navy sections.
- − Incorrect mission numbering ('Apolo 17' instead of Apollo 11) and spelling errors ('NSA' instead of NASA).
- − The icons do not perfectly match the 6 specific steps requested in the prompt.
- − Graphic elements like the planet icons are a bit generic.
Verdict: OmniGen v2 is the clear winner as it actually produces an infographic layout with a logical structure and recognizable icons, despite the mission number error and spelling mistakes. DALL-E 2 fails significantly on prompt adherence, producing garbled text and a chaotic composition that lacks the requested step-by-step logic.
Explore each model
Unified multimodal model for text-to-image generation, instruction-guided image editing, personalized generation, and virtual try-on