Head to head
Esc

Models · slot A

to navigate to pick

DALL-E 3 OpenAI Vidu Q2 ShengShu Technology

Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.

DALL-E 3

20.5 arena score

#38 of 62 in Text-to-Image

Skill signature · Text-to-Image

Vidu Q2

19.8 arena score

#42 of 62 in Text-to-Image

Vote tally

Where the votes landed

DALL-E 3

0.0%

win rate

Ties

0.0%

Vidu Q2

100.0%

win rate

0.0% 0.0% ties 100.0%
Shared challenges 13

Challenge by challenge

The strongest take from each model on every shared challenge, with the AI judge's read.

Geometric Composition

Text-to-Image

“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”

DALL-E 3
Vidu Q2

AI Judge Analysis

DALL-E 3

  • + Excellent photorealistic lighting and textures
  • + Highly creative interpretation with a detailed miniature landscape inside the sphere
  • Failed the spatial instructions by placing the book inside the cube and the sphere on top of the book
  • The 'glass cube' is more of a wooden frame with glass panels

Vidu Q2

  • + Perfect adherence to the spatial requirements of the prompt
  • + Accurate representation of all requested objects in their correct positions
  • + Realistic caustic reflections from the glass onto the table
  • The plant is slightly cut off at the top of the frame
  • The glass cube has a minor artifacts where the top edge meets the book

Verdict: While DALL-E 3 produces a more artistic and visually stunning image, it fails significantly on spatial prompt adherence by placing the red book inside the cube instead of on top. Vidu Q2 followed every instruction perfectly, correctly positioning the blue sphere inside the cube and the red book on top while maintaining high visual quality.

Candid Street Photography

Text-to-Image

“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”

DALL-E 3
Vidu Q2

AI Judge Analysis

DALL-E 3

  • + Excellent composition with a unique 'foreground frame' perspective.
  • + Atmospheric lighting and reflections that perfectly capture a rainy evening mood.
  • + Strong cinematic aesthetic with a well-balanced layout.
  • The subject appears somewhat emaciated and stylistically exaggerated rather than natural.
  • Anatomical issues with the feet and hands are visible.
  • Lacks the requested motion blur from passing cars; the vehicle is static.

Vidu Q2

  • + Captured realistic skin textures and fine details on the hands and face.
  • + Includes subtle motion blur on the passing car as requested.
  • + The bicycle shows realistic wear and tear, and the puddle reflections are convincing.
  • The composition is cluttered and lacks a clear focal point.
  • The framing feels accidental rather than 'imperfect' in an artistic sense.
  • The bicycle structure is physically impossible with two front sections merging.

Verdict: DALL-E 3 produced a far more visually appealing and cinematic image, though it struggled with the request for 'no stylization' and movement. Vidu Q2 followed the technical prompts better regarding skin texture and motion blur, but it failed significantly on the basic anatomy of the bicycle and overall composition. DALL-E 3 is the preferred choice for its artistic coherence, despite the morphological flaws in the human figure.

Fantasy Warrior

Text-to-Image

“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”

DALL-E 3
Vidu Q2

AI Judge Analysis

DALL-E 3

  • + Excellent depiction of warm torchlight reflecting off the metal and skin.
  • + Highly detailed texture on the fabric and leather elements in the foreground.
  • + Effective use of shallow depth of field and bokeh.
  • Failed to include the requested braided hair with beads.
  • The skin texture looks slightly over-processed and 'plastic' in high-contrast areas.

Vidu Q2

  • + Successfully followed all prompt instructions, including the braided hair with beads.
  • + Balanced composition that shows a wider range of the ornate plate armor.
  • + Clear and realistic skin texture with subtle scarring and dirt.
  • The lighting is a bit flat compared to the 'warm torchlight' requested.
  • Some armor details, such as the rivets, appear slightly warped or inconsistent.

Verdict: While DALL-E 3 produces a more dramatic and cinematic lighting effect, it failed to fulfill specific prompt requirements like the braided hair. Vidu Q2 followed every instruction in the prompt, resulting in a more accurate representation of the character, despite having slightly less impactful lighting.

Modern Clean Menu

Text-to-Image

“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”

DALL-E 3
Vidu Q2

AI Judge Analysis

DALL-E 3

  • + Provides a variety of layout options in a collage format.
  • + Excellent food photography that looks professional and appetizing.
  • + Strong grid-based composition that adheres to the request.
  • Text is largely nonsensical and garbled.
  • Shows four layouts rather than one focused menu design.

Vidu Q2

  • + Closer to a functional menu with clear price placeholders.
  • + Clean, bold sans-serif fonts that are mostly legible.
  • + Better categorization of sections like Pizza and Mains.
  • Food photos are repetitive and lack variety compared to the prompt.
  • Some spelling errors in section headers (e.g., 'Apecizen').
  • Layout feels slightly more generic than Model A.

Verdict: Model B is the winner as it creates a more functional and coherent menu design with legible sections and price points, despite some spelling errors. Model A provides better food photography and grid layouts but fails significantly on text legibility and actual menu utility.

Magic Burger Explosion: Fiery Photorealism Challenge

Text-to-Image

“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”

DALL-E 3
Vidu Q2

AI Judge Analysis

DALL-E 3

  • + Excellent photographic texture on the meat and fresh vegetables
  • + Dynamic composition with embers and flying debris creating a sense of impact
  • + Sophisticated lighting with a strong glow reflecting off the bottom of the bun
  • Multiple spelling errors in the text including 'AGIC BURGR' and 'Limiited'
  • Failed to follow the starburst requirement for the price tag

Vidu Q2

  • + Perfect text rendering for all requested phrases including the price
  • + Accurately followed the starburst requirement for the price tag
  • + Strong adherence to the 'fiery, glowing effect' for the typography
  • The burger ingredients look slightly more plastic and less photorealistic than Image A
  • The euro symbol is substituted with a generic currency symbol

Verdict: Vidu Q2 is the clear winner for an advertisement task because it rendered all text perfectly, whereas DALL-E 3 had multiple typos in the main title and secondary message. While DALL-E 3 had slightly better realistic textures on the food, Vidu Q2 followed specific layout instructions like the starburst and the fiery text effect much more accurately.

Chalkboard Menu

Text-to-Image

“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”

DALL-E 3
Vidu Q2

AI Judge Analysis

DALL-E 3

  • + Excellent chalk texture and artistic flourishes
  • + The composition feels like a professional menu layout
  • + Strong lighting and atmospheric café context
  • Numerous spelling errors including 'GRILILLED', 'OCCTUS', and 'TRUFLE'
  • The numbers are wildly inaccurate relative to the prompt (e.g., $234 instead of $24)
  • Text becomes unreadable gibberish in the smaller sections

Vidu Q2

  • + Successfully rendered the date 'APRIL 30, 2026' perfectly
  • + Captured the 'handwritten cursive' and 'slant' requested in the prompt
  • + Much better adherence to the specified pricing ($28 for octopus)
  • Several spelling errors like 'Musshoom', 'Lemepun', and 'Octopd'
  • The last item is a mess of repeated words ('Cookies Chiir cpokies')
  • Chalk texture is less realistic than Model A, looking a bit like a digital brush

Verdict: Both models struggled with the exact spelling of the menu items, following the common AI trend of text hallucination. Vidu Q2 is the winner because it followed the text prompts for the title and prices much more closely, whereas DALL-E 3 produced beautiful art that completely ignored the specific requested text and prices.

The Reversed Rodeo

Text-to-Image

“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”

DALL-E 3
Vidu Q2

AI Judge Analysis

DALL-E 3

  • + Excellent cinematic atmosphere with a vast starry background
  • + Smooth blending of the horse into the ethereal environment
  • Failed the negative constraint: the astronaut is riding the horse, not vice versa
  • Lower level of intricate detail on the astronaut's suit compared to the competitor

Vidu Q2

  • + High level of sharp, intricate detail on the spacesuit and horse tack
  • + Creative galaxy-textured coat on the horse
  • + Dynamic composition with vibrant colors
  • Failed the negative constraint: the astronaut is riding the horse, not vice versa
  • The horse's frontal anatomy looks slightly distorted where the neck meets the chest

Verdict: Both DALL-E 3 and Vidu Q2 failed the specific negative constraint to have the 'horse on top' of the astronaut, instead opting for the more common interpretation of an astronaut riding a horse. Vidu Q2 is the preferred image due to its superior sharpness, vibrant colors, and more detailed rendering of the subjects, whereas DALL-E 3 feels slightly more generic in its execution.

The Capybara Taxi Driver

Text-to-Image

“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”

DALL-E 3
Vidu Q2

AI Judge Analysis

DALL-E 3

  • + Excellent fur texture rendering and lighting on the capybara's face.
  • + High visual clarity and a cohesive cinematic art style.
  • + Intelligent detail in the background with a 'Capybara' sign on a building.
  • Fails to include the human businesswoman in the back seat entirely.
  • The taxi cap is black instead of the requested yellow.

Vidu Q2

  • + Successfully includes both the capybara driver and the businesswoman in the back seat.
  • + Accurately represents the businesswoman's bored expression as she looks at her phone.
  • + Followed more prompt constraints including the yellow color of the hat.
  • The transition between the capybara head and the human-like body is slightly awkward.
  • Some anatomical issues with the capybara's hands/paws on the steering wheel.

Verdict: While DALL-E 3 produced a higher quality, more aesthetically pleasing single-subject portrait, it failed the primary composition task by omitting the passenger. Vidu Q2 successfully followed the complex prompt by including all characters with the correct expressions and colors, making it the better adherence to the specific instructions.

The Halloween Invitation

Text-to-Image

“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”

DALL-E 3
Vidu Q2

AI Judge Analysis

DALL-E 3

  • + Exquisite gothic aesthetic with deeply layered 3D-like textures.
  • + High visual quality with moody, cinematic lighting and a professional layout.
  • + Captures the vintage 'scary' atmosphere perfectly with web and thorn details.
  • Fails significantly on text rendering, with most body text being illegible gibberish.
  • Missed the specific event details like 'The Arches' and the correct year in the text.

Vidu Q2

  • + Successfully included all requested event details including location and date.
  • + Stronger adherence to the specific text requests, including the banner text.
  • + Good compositional balance with the central jack-o-lantern.
  • Several spelling errors in the main title and banner (e.g., 'Intoviztion', 'invieed').
  • Visual style is more illustrative and lacks the 'cinematic' polish and depth of the competitor.
  • The date contains a typo (30.70.2025).

Verdict: DALL-E 3 produces a much more visually stunning and atmospheric image that perfectly captures the gothic, cinematic prompt, but it fails to render any legible specific text. Vidu Q2 follows the instructions for text content and layout much more closely, though it suffers from multiple spelling errors and a flatter artistic style. DALL-E 3 is preferred for its artistic quality, while Vidu Q2 is closer to a functional invitation despite its typos.

Isometric Miniature Diorama Scenes

Text-to-Image

“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”

DALL-E 3
Vidu Q2

AI Judge Analysis

DALL-E 3

  • + Excellent 3D isometric execution with clean textures.
  • + Great lighting and shadow work that enhances the miniature effect.
  • + High level of rendering detail on the sushi and diorama base.
  • Failed to place the text 'JAPAN' and 'SUSHI' at top-center as requested.
  • Text rendering is incomplete, missing the word 'SUSHI' entirely.

Vidu Q2

  • + Followed the text layout instructions perfectly with 'JAPAN' and 'SUSHI' at top-center.
  • + Strong variety in sushi types depicted while maintaining the cartoon aesthetic.
  • + Successfully included all elements including the flag icon and solid background.
  • The diorama base is less 'isometrically' accurate than Model A.
  • Minor artifacting on the textures of the shrimp and gunkan sushi.

Verdict: While DALL-E 3 produced a more polished 3D render with superior lighting, it failed significantly on the text placement and content instructions. Vidu Q2 followed the complex layout instructions perfectly, including the specific text positioning and flag icon, making it the better overall response to the prompt.

Adorable Baby Animals in Sunny Meadow

Text-to-Image

“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”

DALL-E 3
Vidu Q2
0% wins 0% ties 100% wins

AI Judge Analysis

DALL-E 3

  • + Excellent adherence to lighting requests with clear 'god rays' and dew sparkles.
  • + Very consistent stylistic rendering of fur and expressive eyes.
  • + Cohesive composition that feels intimate and magical.
  • Has a distinct digital illustration/CGI feel rather than 'hyper-photorealistic'.
  • Strange 'butterfly-animal' hybrids where butterfly bodies have been replaced with fuzzy heads.

Vidu Q2

  • + Achieves a much higher level of photorealism as requested.
  • + Successfully depicts the movement and 'playfully chasing' aspect of the prompt.
  • + All four requested animal types are present and identifiable.
  • Composition includes an extra puppy that was not requested.
  • The kitten's tail has a slightly awkward anatomical connection to its body.
  • Butterflies feel a bit 'pasted on' with some lacking natural drop shadows.

Verdict: DALL-E 3 creates a charming, storybook-style illustration with magical lighting, but it fails the 'photorealistic' requirement and produces strange butterfly hybrids. Vidu Q2 captures a much more realistic scene with believable fur textures and energetic poses that better match the 'chasing' and '8K' aspects of the prompt. While Vidu Q2 adds an extra puppy, its adherence to the requested aesthetic makes it the superior choice.

Vintage Cafe Logo

Text-to-Image

“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”

DALL-E 3
Vidu Q2

AI Judge Analysis

DALL-E 3

  • + Excellent visual quality with clear, professional vector-style graphics.
  • + Captures the vintage aesthetic perfectly with stippling textures and balanced composition.
  • + Includes the required cloche and date elements in a coherent design.
  • Failed to include the specific name 'Caffè Florian', substituting it with 'Coffee House'.

Vidu Q2

  • + Successfully incorporated a banner as requested.
  • + Closely followed the requested color palette of warm browns and creams.
  • Significant text hallucinations and spelling errors like 'Farmiin' and 'Esttt'.
  • The cloche handle is detached and floating awkwardly above the steam.
  • Lacks the professional polish and clean lines of a vector logo.

Verdict: While DALL-E 3 failed to use the specific name requested, it produced a high-quality, professional-grade logo that looks like a finished product. Vidu Q2 attempted the text but failed due to severe spelling errors and poor structural coherence, such as the floating cloche handle. DALL-E 3 is the preferred choice for its superior artistic execution and vector-like clarity.

Apollo 11: Journey to Tranquility

Text-to-Image

“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”

DALL-E 3
Vidu Q2

AI Judge Analysis

DALL-E 3

  • + Exhibits a professional and complex layout consistent with high-end infographics.
  • + Captures the requested navy, muted red, and light gray NASA palette perfectly.
  • + Includes stylistically consistent, textured circles that evoke lunar surfaces.
  • Fails to follow the specific 6-step sequence requested in the prompt.
  • Contains significant text gibberish and visual repetition across the three layout options.
  • The inclusion of Space Shuttle-style imagery is historically inaccurate for the Apollo 11 mission.

Vidu Q2

  • + Attempts to follow a sequential step-by-step layout as requested.
  • + The vector icons (lunar module, rocket, earth) are clean and have consistent line weights.
  • + Adheres reasonably well to the flat-vector style with subtle gradients.
  • Severe spelling errors throughout (e.g., 'Alfonch' for Apollo, 'Aeulth Orot').
  • Fails to meet the 'navy' color palette requirement, using mostly light gray and white.
  • Logic errors in icons, such as the lunar module being repeated for 'descent' and 'landing' without variation.

Verdict: Model B (Vidu Q2) is the preferred choice because it actually attempts to organize the infographic into the requested steps, whereas Model A (DALL-E 3) generates a decorative but functionally useless three-panel mockup with historically inaccurate spacecraft. While Model B suffers from poor text rendering, its iconography is more relevant to the specific prompt instructions regarding the mission phases.

Next steps

Explore each model