Head to head
Esc

Models · slot A

to navigate to pick

Imagen 4.0 Generate 001 Google Vidu Q2 ShengShu Technology

Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.

Imagen 4.0 Generate 001

17.1 arena score

#54 of 62 in Text-to-Image

Skill signature · Text-to-Image

Vidu Q2

19.8 arena score

#42 of 62 in Text-to-Image

Vote tally

Where the votes landed

Imagen 4.0 Generate 001

0%

win rate

Ties

0%

Vidu Q2

0%

win rate

Shared challenges 13

Challenge by challenge

The strongest take from each model on every shared challenge, with the AI judge's read.

Geometric Composition

Text-to-Image

“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”

Imagen 4.0 Generate 001
Vidu Q2

AI Judge Analysis

Imagen 4.0 Generate 001

  • + Very clean and photorealistic textures
  • + Excellent handling of soft, diffused lighting
  • + Minimalist and aesthetically pleasing composition
  • The cube appears more like solid acrylic/columns rather than a hollow glass box
  • The sphere appears to be levitating without explanation

Vidu Q2

  • + Perfect adherence to the 'hollow cube' spatial concept
  • + Detailed and realistic plant life visible through the glass
  • + Complex light interactions with shadows and reflections
  • The lighting is harsh rather than the requested 'soft' light
  • Slightly more cluttered composition compared to Model A

Verdict: Both models followed the prompt instructions well. Imagen 4.0 created a more artistic, soft, and high-quality image, but it interpreted the cube as a solid block of glass. Vidu Q2 captured the physical spatial relationship more accurately by making the cube a hollow container, though the lighting was more intense than requested.

Candid Street Photography

Text-to-Image

“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”

Imagen 4.0 Generate 001
Vidu Q2

AI Judge Analysis

Imagen 4.0 Generate 001

  • + Excellent shallow depth of field and bokeh from city lights.
  • + Realistic skin textures and fine details on the man's clothing and face.
  • + Strong cinematic composition while maintaining the requested candid feel.
  • The tool in his hands is slightly warped and not clearly defined as a specific bicycle tool.
  • Motion blur on the passing vehicle is present but feels a bit static.

Vidu Q2

  • + Successfully captured the 'imperfect framing' request with a more off-center, candid crop.
  • + The reflections on the wet pavement are very prominent and well-rendered.
  • + The bicycle shows realistic wear and tear consistent with an older man's daily commuter.
  • Significant anatomical issues with the hands, including an extra thumb/finger structure.
  • The car in the background lacks the requested motion blur, appearing relatively sharp.
  • The overall image quality is slightly softer with less defined skin texture compared to Model A.

Verdict: Imagen 4.0 is the clear winner due to its superior technical execution of skin textures and anatomical details, specifically the hands. While Vidu Q2 followed the 'imperfect framing' prompt more literally, it failed significantly on the human anatomy and missed the specific request for motion blur on the passing cars.

Fantasy Warrior

Text-to-Image

“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”

Imagen 4.0 Generate 001
Vidu Q2

AI Judge Analysis

Imagen 4.0 Generate 001

  • + Excellent rendition of the 'hair braided with small beads' prompt with numerous distinct beads.
  • + Superior engraving detail on the plate armor with very clean scrollwork.
  • + Strong adherence to the warm torchlight and bokeh sparks lighting request.
  • The facial features look slightly plastic and less lifelike than the competition.
  • The armor design feels a bit repetitive and flat despite the ornate texture.

Vidu Q2

  • + Highly realistic skin texture and lifelike eyes that better convey the 'battle-worn' character.
  • + Superior rendering of the cloth underlayer's texture, showing realistic weave.
  • + More natural integration of scars and dirt which look embedded in the skin rather than drawn on.
  • Failed to fully follow the instruction for beads in the hair, providing only a few on a single braid.
  • The engraving on the armor is less intricate and appears slightly more smudged compared to Model A.

Verdict: Imagen 4.0 excels at technical prompt adherence, particularly with the intricate armor engravings and the specific hair beads requested. However, Vidu Q2 produces a more emotionally resonant and lifelike character with superior skin and fabric textures, despite missing the specific count of beads requested. Imagen 4.0 is preferred if the goal is decorative detail, while Vidu Q2 is better for realistic character portraiture.

Modern Clean Menu

Text-to-Image

“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”

Imagen 4.0 Generate 001
Vidu Q2

AI Judge Analysis

Imagen 4.0 Generate 001

  • + Excellent adherence to the 'grid' prompt requirement.
  • + Clean, professional typography that is highly legible.
  • + High-quality, distinct food photography that matches the specified categories.
  • Text consists of nonsensical placeholder words.
  • The geometric accent shapes feel a bit repetitive.

Vidu Q2

  • + Uses a two-column layout traditional for menus.
  • + Vibrant food photography.
  • Text rendering is very poor with significant character distortion.
  • Does not follow the requested 'grid' structure for the overall design.
  • Messy visual layout with overlapping elements and inconsistent spacing.

Verdict: Imagen 4.0 follows the prompt's request for a minimalist grid design perfectly, delivering a clean and professional layout with high-quality visual elements. Vidu Q2 fails to maintain a clean grid structure and suffers from significant text distortion and a cluttered aesthetic that does not fit the 'minimalist' criteria.

Magic Burger Explosion: Fiery Photorealism Challenge

Text-to-Image

“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”

Imagen 4.0 Generate 001
Vidu Q2

AI Judge Analysis

Imagen 4.0 Generate 001

  • + Excellent typography rendering with consistent glow effects.
  • + Superior photorealistic quality on the burger textures, particularly the patties and bun.
  • + Clean starburst element that matches the specific price request perfectly.
  • The 'exploded' effect feels a bit linear and vertical rather than dynamic.
  • Missing some of the requested 'fiery background' intensity compared to the other model.

Vidu Q2

  • + Highly dynamic background with intense fire and glowing embers.
  • + Strong sense of motion with liquid sauce splashes and angled ingredients.
  • + Included the fiery glowing effect on the main title text as requested.
  • The currency symbol is incorrect, showing a stylized '#' or a malformed Euro sign.
  • The bun texture and starburst element look slightly more illustrative and less photorealistic than Image A.

Verdict: Imagen 4.0 Generate 001 provides a much more polished and professional ad layout with perfect text rendering and high-end food photography textures. While Vidu Q2 captures the explosive, fiery energy of the prompt more effectively, it fails on technical details like the currency symbol and overall image sharpness. Imagen 4.0 is the preferred choice for a commercial-ready outcome.

Chalkboard Menu

Text-to-Image

“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”

Imagen 4.0 Generate 001
Vidu Q2

AI Judge Analysis

Imagen 4.0 Generate 001

  • + Excellent text legibility and spelling for most of the required items.
  • + Clean and realistic wooden frame composition.
  • + Captured the request for 'Brown Butter Chocolate Chip Cookies' which was cut off in the prompt text.
  • The text style looks more like a digital marker or thin chalk pen than traditional textured chalk.
  • Included meta-text from the prompt instructions (e.g., 'Tittle', 'Menu', 'Footer') directly onto the board.
  • Spelling errors in the 'Berbs' for Herbs and several nonsense sentences in the footer.

Vidu Q2

  • + Beautiful, realistic chalk texture with authentic smudging and varying pressure.
  • + Captured the aesthetic of a 'cozy café' background much better than a flat studio shot.
  • + Handwriting style is more charming and artistic.
  • Significant spelling errors throughout almost every menu item (e.g., 'Musshoom', 'Octopd', 'Browd Botter').
  • Price of the first item was altered from $24 to $34.
  • The text becomes increasingly jumbled and illegible toward the bottom of the board.

Verdict: Imagen 4.0 followed the specific text instructions much more accurately, correctly predicting the completion of the truncated prompt and maintaining better spelling, despite the error of including meta-labels like 'Tittle'. Vidu Q2 has a much more beautiful and realistic chalk aesthetic, but it fails significantly on text rendering and spelling accuracy.

The Reversed Rodeo

Text-to-Image

“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”

Imagen 4.0 Generate 001
Vidu Q2

AI Judge Analysis

Imagen 4.0 Generate 001

  • + Excellent cinematic lighting with a realistic 3D feel.
  • + Coherent anatomy for both the astronaut and the horse.
  • + The composition feels grounded yet surreal with the planet curvature below.
  • Failed to follow the specific spatial instruction for the horse to be on top of the astronaut.

Vidu Q2

  • + Vibrant, dreamlike color palette that leans heavily into the surreal prompt.
  • + Detailed nebula textures integrated into the horse's coat.
  • + Dynamic and energetic composition.
  • Failed to follow the specific spatial instruction for the horse to be on top of the astronaut.
  • Minor anatomical distortions in the horse's front legs.

Verdict: Both Imagen 4.0 and Vidu Q2 failed the difficult spatial reasoning task of placing the horse on top of the astronaut, instead defaulting to the standard 'astronaut riding horse' trope. Imagen 4.0 is preferred for its superior cinematic lighting and cleaner lines, whereas Vidu Q2 feels slightly cluttered by comparison despite its creative colors.

The Capybara Taxi Driver

Text-to-Image

“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”

Imagen 4.0 Generate 001
Vidu Q2

AI Judge Analysis

Imagen 4.0 Generate 001

  • + Excellent adherence to the 'bored' expression for the passenger.
  • + Stronger composition with a more centered, symmetrical view of the taxi interior.
  • + Clean, legible text on the taxi sign and driver cap.
  • The capybara's paws look more like bird talons than capybara feet.
  • The perspective of the dashboard in the foreground is slightly flat.

Vidu Q2

  • + The capybara's 'hands' are rendered more realistically as mammalian paws.
  • + The lighting and reflections on the leather seats are very high quality.
  • + Great cinematic depth of field reflecting a busy Manhattan street.
  • The passenger's face is slightly blurry and lower detail compared to the driver.
  • The capybara's expression is slightly more alert than the 'calm' tone requested.

Verdict: Both models followed the complex prompt very well, but Imagen 4.0 provided a superior composition and better character expression for the businesswoman. While Vidu Q2 had better rendering of the capybara's paws, Imagen 4.0 felt more like a coherent, intentional scene with clearer details on the branding and text elements.

The Halloween Invitation

Text-to-Image

“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”

Imagen 4.0 Generate 001
Vidu Q2

AI Judge Analysis

Imagen 4.0 Generate 001

  • + Perfect text rendering for all requested fields including the invitation title, banner, and specific event details.
  • + Clean and polished composition with high-quality cinematic lighting.
  • + Excellent adherence to all prompt elements including the thorns, webs, and twisted trees.
  • The visual style is slightly more vector/illustration-like rather than looking like a vintage parchment poster.
  • The large vertical tube on the left is a bit abstract and doesn't clearly serve a thematic purpose.

Vidu Q2

  • + Captures the 'vintage parchment' texture more effectively than the competitor.
  • + Strong atmospheric framing with the thorny branches and webs.
  • Significant spelling errors throughout including 'Intovztion', 'invieed', and 'might of fiigts'.
  • Incorrect date '2025' and 'Tmm' instead of the requested details.
  • Text is poorly integrated and lacks the polished clarity of the other image.

Verdict: Imagen 4.0 is the clear winner due to its flawless text rendering and adherence to specific details like the date and location. While Vidu Q2 captures a more authentic vintage texture, it fails significantly on text legibility and accuracy, making the invitation unusable.

Isometric Miniature Diorama Scenes

Text-to-Image

“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”

Imagen 4.0 Generate 001
Vidu Q2

AI Judge Analysis

Imagen 4.0 Generate 001

  • + Excellent realistic PBR materials, especially on the tuna and roe.
  • + Soft, high-quality photographic lighting.
  • + High level of detail in food textures like the rice and ginger.
  • Completely failed to include requested text 'JAPAN' and 'SUSHI'.
  • Missing the flag icon and the light blue background.
  • Does not look like a 'cartoon scene' as requested.

Vidu Q2

  • + Perfect adherence to all text, flag, and color requirements.
  • + Strong '3D cartoon' aesthetic with smooth, refined textures.
  • + Correct isometric composition on a diorama base with a light blue background.
  • Some minor texture blurring on the sushi ingredients.
  • The rice grains look slightly oversized and chunky.

Verdict: While Imagen 4.0 produces a much more realistic and appetizing image with superior textures, it completely ignores the text, flag, and background color requirements. Vidu Q2 followed every instruction in the prompt, including complex text rendering and specific background colors, making it the better response to the prompt instructions.

Adorable Baby Animals in Sunny Meadow

Text-to-Image

“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”

Imagen 4.0 Generate 001
Vidu Q2

AI Judge Analysis

Imagen 4.0 Generate 001

  • + Excellent character interactions with the puppy hugging the kitten
  • + Extremely sharp and detailed textures on the fur and wildflowers
  • + Dynamic and clear composition with a central focal point
  • The style leans more toward digital illustration than 'hyper-photorealistic'
  • The paws on the kitten are somewhat repetitive in shape

Vidu Q2

  • + Successfully captures a more photorealistic lighting and depth of field
  • + Strong adherence to the movement aspect of the prompt with chasing/running poses
  • + Beautiful atmosphere with 'bokeh' effects and multiple butterflies
  • Anatomical issues including a dog with an extra tail/conjoined body part
  • Includes two golden retriever puppies instead of one
  • The rabbit has a somewhat dog-like snout and tail

Verdict: Imagen 4.0 produces a much cleaner and higher-quality image with charming character interactions, although it feels more like a 3D render than a photo. Vidu Q2 achieves a more realistic photographic style with beautiful lighting, but it suffers from significant anatomical errors and prompt counting inaccuracies. Imagen 4.0 is the preferred choice for its technical polish and adorable composition.

Vintage Cafe Logo

Text-to-Image

“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”

Imagen 4.0 Generate 001
Vidu Q2

AI Judge Analysis

Imagen 4.0 Generate 001

  • + Perfect text rendering for both the brand name and established date.
  • + Clean vector art aesthetic that matches the minimalist request.
  • + Balanced composition with excellent use of negative space.
  • The 'Est. 1720' is in a banner style but doesn't wrap around the cloche as classically as some vintage logos might.

Vidu Q2

  • + Good use of subtle background texture as requested.
  • + Captures a more 'vintage' illustrative style with the ribbon.
  • Major spelling errors in the brand name ('CAFFE FARMIIN' and 'CAFFEE FOPLI20').
  • Visual artifacts and clipping where the steam overlaps the cloche handle.
  • Typography is cluttered and inconsistent.

Verdict: Imagen 4.0 followed all instructions perfectly, producing a professional-grade vector logo with flawless text. In contrast, Vidu Q2 failed significantly on text rendering and produced a cluttered design with several spelling errors.

Apollo 11: Journey to Tranquility

Text-to-Image

“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”

Imagen 4.0 Generate 001
Vidu Q2

AI Judge Analysis

Imagen 4.0 Generate 001

  • + Excellent text rendering and legibility.
  • + Sophisticated layout that uses paths to connect mission phases logically.
  • + High-quality vector illustration style with a clean NASA-inspired palette.
  • Missed the final two requested steps (Descent and Landing).
  • The 'Launch' icon incorrectly shows Earth instead of the Saturn V rocket.
  • The trajectory lines are visually confusing and do not accurately represent real mission phases.

Vidu Q2

  • + Follows the requested content more closely by including the Lunar Module and landing surface.
  • + Captures the 'flat-vector' instruction with clean icon design.
  • + Uses a cohesive light-gray based color palette that fits the modern aesthetic.
  • Severe text corruption and gibberish rendering throughout the infographic.
  • Logical error showing five astronauts instead of the three on Apollo 11.
  • The 'Step' numbering and labels are disorganized and factually nonsensical.

Verdict: Imagen 4.0 Generate 001 is the superior image due to its professional composition and perfect text rendering, even though it failed to include the final two steps of the mission. Vidu Q2 followed the prompt's structural requirements better by including the lunar module and landing surface, but the image is unusable due to significant text glitches and gibberish. Imagen 4.0's output actually functions as a clean graphic, whereas Vidu Q2's output fails to communicate any clear information.

Next steps

Explore each model