OpenAI's legacy image generation model supporting generations, edits with masks (inpainting), and variations
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
DALL-E 2
#59 of 62 in Text-to-Image
Vidu Q2
#42 of 62 in Text-to-Image
Where the votes landed
DALL-E 2
0%
win rate
Ties
0%
Vidu Q2
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
DALL-E 2
- + Features a glass cube and a wooden table.
- + Soft lighting follows the prompt's direction.
- − Failed almost all spatial instructions, including placing the sphere inside and the book on top.
- − The blue sphere is missing and replaced by a blue vase.
- − Low image resolution and lack of fine details.
Vidu Q2
- + Perfect adherence to all spatial instructions and color/object prompts.
- + High visual quality with realistic textures on the book, glass, and wood.
- + Excellent rendering of light and shadows, including the plant's shadow on the table.
- − None observed.
Verdict: Vidu Q2 followed every detail of the complex spatial prompt perfectly, creating a high-resolution, realistic scene with correct object placement. DALL-E 2 struggled significantly, failing to place the sphere inside or the book on top, and misinterpreted the blue sphere as a large blue pot.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
DALL-E 2
- + Successfully captures a shallow depth of field and soft reflections.
- + Achieves the 'imperfect framing' and 'candid' feel requested in the prompt.
- − The subject is so out of focus and obscured that the prompt details (elderly Japanese man, repairing) are indistinguishable.
- − Lacks clarity and overall visual quality compared to modern models.
Vidu Q2
- + Excellent adherence to all prompt details, including the specific action, the car, and the wet pavement.
- + High level of detail in skin texture and the mechanical components of the bicycle.
- + Effective use of reflections and lighting to convey the rainy atmosphere.
- − The 'motion blur' on the car is quite subtle, looking more static than a moving vehicle.
- − The anatomy of the man's hands is slightly distorted where they grip the bicycle.
Verdict: Vidu Q2 is the clear winner as it successfully renders all elements of the complex prompt, providing a clear, realistic, and detailed depiction of an elderly man repairing a bicycle. DALL-E 2 failed to keep the subject in focus, rendering a blurry and largely unidentifiable scene that misses the core narrative of the prompt.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
DALL-E 2
- + Captures a strong bokeh effect with torch-like coloring.
- − Extremely low resolution and clarity.
- − The prompt adherence is poor, failing to render the face, braids, or clear armor details.
- − Heavy digital artifacts make the subject almost unrecognizable.
Vidu Q2
- + Excellent adherence to all prompt details including the braided hair with beads, scars, and ornate engraving.
- + High visual quality with realistic textures on leather, cloth, and skin.
- + Strong lighting and atmosphere that matches the 'warm torchlight' and 'bokeh sparks' description.
- − The eyes, while detailed, are slightly asymmetrical.
Verdict: Vidu Q2 is the clear winner as it followed every specific detail of the prompt with high-fidelity textures and professional composition. DALL-E 2 produced an abstract, messy image that failed to render the character's face, hair, or the level of detail requested in the armor.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
DALL-E 2
- + Strong bold sans-serif typography on the right page.
- + Creative grid layout using geometric food photography.
- − Does not follow the layout requirements for specific sections like appetizers, pizza, or mains.
- − The food photos are heavily fragmented and unrecognizable as dishes.
- − Text is nonsensical and lacks the structure of a restaurant menu.
Vidu Q2
- + Excellent adherence to layout requirements including sections for appetizers, pizza, and mains.
- + High-quality, vibrant food photography that looks appetizing and professional.
- + Clean, modern minimalist design with well-organized price listings and descriptions.
- − Text contains several AI hallucinations and misspellings (e.g., 'Appeczyers', 'PitoaA').
- − Some menu items are repetitive or have layout overlapping issues in the top-right section.
Verdict: Vidu Q2 is the clear winner as it accurately interprets the user's request for a functional menu layout with specific sections and a professional aesthetic. While DALL-E 2 produces an artistic book-like layout, it fails to deliver a usable menu design and its food photography is overly abstract.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
DALL-E 2
- + Successfully creates a sense of glowing energy and embers.
- − Text is garbled and does not match the prompt (e.g., 'MARGIC BAGUEC').
- − The burger components are poorly defined and lack photorealistic detail.
- − Fails to include the price and starburst element.
Vidu Q2
- + Excellent text rendering, following all three text requirements perfectly.
- + High-quality photorealistic textures on the lettuce, patty, and seeds.
- + Dynamic and balanced composition with clear separation of exploded layers.
- − The currency symbol is a generic hash/double-bar rather than a clear Euro symbol.
Verdict: Vidu Q2 is the clear winner as it perfectly executes all text requirements and provides a crisp, professional advertisement layout. DALL-E 2 fails significantly on both text legibility and the quality of the food components, resulting in a blurry and incoherent image.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
DALL-E 2
- + Successfully captures a chalk-like texture.
- − The text is completely illegible and gibberish.
- − The composition is poor, zoomed in too far with no background context.
- − Fails to follow any of the specific textual instructions from the prompt.
Vidu Q2
- + Excellent adherence to the specific text requested, including the date and item names.
- + Authentic handwriting style with realistic chalk strokes and smear effects.
- + Clear, high-resolution composition with a warm café background.
- − Minor spelling errors in the menu items (e.g., 'Musshoom', 'Lemepun').
- − The transition of the price on the third item becomes messy.
Verdict: Vidu Q2 is the clear winner as it successfully rendered most of the complex prompt text with an authentic-looking chalkboard aesthetic. DALL-E 2 failed completely to produce legible words, resulting in gibberish that ignored the prompt's specific requirements.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
DALL-E 2
- + Successfully placed the horse on top of the astronaut as requested by the specific prompt instruction.
- + Maintains a gritty, cinematic space texture.
- − Lower resolution with significant noise and artifacting.
- − Anatomical details on the horse and astronaut are muddy and poorly defined.
Vidu Q2
- + Stunning visual quality with high resolution and vibrant colors.
- + Excellent texture work on the galaxy-patterned horse and the space suit.
- − Failed the negative constraint/specific positioning instruction entirely by placing the astronaut on top of the horse.
- − Leans more toward fantasy illustration than the requested surreal cinematic look.
Verdict: This comparison highlights the trade-off between prompt following and visual fidelity. DALL-E 2 followed the difficult spatial instruction to place the horse on top, but the image quality is low and distorted; conversely, Vidu Q2 produced a beautiful, high-quality image but completely ignored the core spatial requirement of the prompt.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
DALL-E 2
- + Attempts to use a wide-angle perspective.
- − Major anatomical distortions and disturbing facial artifacts on the human.
- − The capybara is unrecognizable and appears as a blurry yellow mass.
- − Fails to place the businessman in the back seat and lacks photorealism.
Vidu Q2
- + Excellent adherence to all prompt details including the driver's cap and professional expression.
- + High-quality lighting and realistic taxi interior textures.
- + Accurately depicts the bored businesswoman in the backseat as requested.
- − The capybara's hands have slightly too many fingers, though the pose is correct.
Verdict: Vidu Q2 is the clear winner as it followed every instruction in the prompt with high fidelity and photorealism. In contrast, DALL-E 2 produced a highly distorted image with severe artifacts and failed to correctly render either the capybara or the human subject.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
DALL-E 2
- + Captures a dark, moody 'vintage' atmosphere very well.
- + The border design is intricate and feels appropriately gothic.
- − Text is illegible and full of spelling errors.
- − Failed to include a clear glowing jack-o-lantern as requested.
- − Image is blurry with low visual clarity.
Vidu Q2
- + Clear, readable layout with distinct gothic font elements.
- + Excellent adherence to all visual prompts including the jack-o-lantern, webs, and thorns.
- + High resolution and polished graphical quality.
- − Contains minor spelling errors like 'Intovztion' and 'invieed'.
- − The date was rendered as '30.70.2025' instead of '30.10.2026'.
Verdict: Vidu Q2 is the clear winner as it successfully incorporated almost all prompt elements, including the specific text banners, the central jack-o-lantern, and the complex border. DALL-E 2 produced an atmospheric but unusable image with illegible text and poor clarity.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
DALL-E 2
- + Clean minimalist color palette.
- + Good 3D lighting and shadow rendering.
- − Failed significantly on text rendering, misspelling it as 'Sush'.
- − Missing the 'JAPAN' text and flag icon entirely.
- − Composition is unbalanced with most elements cut off or off-center.
Vidu Q2
- + High adherence to all prompt instructions including text and iconography.
- + Excellent 3D miniature style with smooth textures and appealing colors.
- + Correctly rendered the isometric diorama base and centered composition.
- − The flag icon is slightly simple in design compared to the rest of the 3D scene.
Verdict: Vidu Q2 is the clear winner as it followed every part of the complex prompt, including specific text elements and the diorama base. DALL-E 2 failed to include most of the requested text and produced a poorly composed image with misspelled words.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
DALL-E 2
- + Successfully included a golden retriever and a tabby kitten.
- − Low visual fidelity with significant artifacts and smudging.
- − The animals are anatomically distorted, particularly the cat and the unidentified brown mass at the bottom.
- − The background is painterly and blurry, failing the 'hyper-photorealistic' and '8K' requirements.
Vidu Q2
- + Excellent adherence to all prompt elements including the specific animal types, lighting, and environment.
- + High visual quality with sharp details in fur, flowers, and butterflies.
- + Beautifully rendered golden hour lighting with 'god rays' and bokeh effects.
- − Contains an extra golden retriever puppy not specifically requested in the count.
Verdict: Vidu Q2 is the clear winner as it produced a vibrant, high-quality image that perfectly captured the requested atmosphere and fine details. DALL-E 2 failed significantly on technical quality, producing distorted animals and a muddy composition that lacked the requested photorealism.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
DALL-E 2
- + Successfully captures a minimalist vector style
- + Adheres well to the requested monochromatic warm brown color palette
- − Text consists of completely nonsensical characters
- − Fails to include the requested banner element
- − Steam effect is poorly defined and looks like an artifact
Vidu Q2
- + Excellent visual quality with professional shading and highlight details
- + Includes all requested elements including the cloche, steam, and banner
- + Attempts the specific dates and names requested in the prompt
- − Contains spelling errors in the brand name and 'Est' prefix
- − Wait-style text repetition at the bottom creates visual clutter
Verdict: Vidu Q2 is the clear winner as it successfully incorporates every element of the prompt, including the banner and steam, into a cohesive and visually appealing logo. While both models struggled with text accuracy, Vidu Q2's rendering of 'Est. 1720' is nearly perfect, whereas DALL-E 2 produced entirely unintelligible gibberish and missed the requested banner component.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
DALL-E 2
- + Captures the requested navy, white, and red NASA-inspired color palette.
- + Achieves a complex, tech-heavy aesthetic that feels like a mission control dashboard.
- − Text is completely unintelligible and includes major spelling errors in the title ('ALLPOO').
- − Failed to follow the requested 6-step chronological structure, resulting in a chaotic layout.
- − Lacks the specific requested icons like the Saturn V rocket.
Vidu Q2
- + Follows the requested chronological infographic structure much more accurately.
- + Successfully uses many of the specified icons including planets, orbits, and lunar modules.
- + Maintains a clean, consistent vector style with high legibility.
- − Text contains several spelling errors, though still more readable than model A.
- − Missed the 'navy' color requirement, opting for a very light gray background.
- − The icons for the landing phase repeat the same image three times with minor variations.
Verdict: Vidu Q2 is the clear winner as it successfully interprets the request for a structured 6-step infographic, providing distinct icons for the launch, orbit, and landing phases. In contrast, DALL-E 2 produced a disorganized layout with severe text butchering and failed to include the specific mission steps requested.
Explore each model
ShengShu Technology's text-to-image and reference-to-image model with support for character consistency and multi-reference image processing