Google's Imagen 3.0 text-to-image generation model, producing high-quality images with improved detail and lighting
Settled by community votes across 6 shared challenges, with an AI judge weighing in on each.
Imagen 3.0 Generate 002
#46 of 62 in Text-to-Image
Vidu Q2
#42 of 62 in Text-to-Image
Where the votes landed
Imagen 3.0 Generate 002
0%
win rate
Ties
0%
Vidu Q2
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
Imagen 3.0 Generate 002
- + Excellent grid layout with clear, logical sections
- + Food photography is high quality and consistent in style
- + Superior font choice and hierarchy for a professional menu look
- − Text is largely gibberish, though it maintains the appearance of real words
- − Includes repetitive header tags within the grid
Vidu Q2
- + Colorful and vibrant food presentation
- + Varied section layouts for visual interest
- − Poor typography with significant character distortion and inconsistencies
- − Composition feels cluttered and less professional
- − Grid layout requested in prompt is poorly executed
Verdict: Imagen 3.0 follows the prompt much more effectively, delivering a professional-grade grid layout that looks like a real menu design. While both models fail to produce legible text, Imagen 3.0's graphic design and photography are vastly superior to Vidu Q2, which suffers from significant font artifacts and a messy composition.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
Imagen 3.0 Generate 002
- + Excellent photographic texture on the meat and buns
- + Perfect text rendering for all main and secondary messages
- + Clean and professional aesthetic balance
- − The 'starburst' for the price is more of a jagged badge
- − Gibberish text above the price tag
Vidu Q2
- + Dynamic 'fiery' effect on the typography matches the prompt well
- + Vibrant and intense background flames
- + Good sense of motion with liquid splashes
- − The currency symbol is incorrect (looks like a double-crossed E/pound hybrid)
- − Lower overall photorealism compared to Model A
- − Text is somewhat compressed at the top
Verdict: Imagen 3.0 Generate 002 produces a much more professional and realistic food advertisement with superior texture and lighting. While Vidu Q2 followed the 'fiery text' instruction better, it failed on the currency symbol and had less convincing food textures.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
Imagen 3.0 Generate 002
- + Excellent text legibility for the majority of the prompt.
- + Superior chalk texture and realistic wood frame composition.
- + Follows the elegant cursive requirement for the header reasonably well.
- − Includes several spelling errors like 'Risottto' and 'Octpsuin'.
- − Duplicates menu lines in a nonsensical way at the bottom.
Vidu Q2
- + Features a very authentic, messy chalk texture with smudge effects.
- + Better background enviornment conveying a 'cozy café' atmosphere.
- + Captures the cursive slant and handwriting variation well.
- − Significant spelling failures and garbled text across all menu items.
- − Failed to render the correct pricing ($34 instead of $24).
- − Text at the bottom becomes illegible gibberish.
Verdict: Imagen 3.0 Generate 002 is the clear winner as it maintains much higher text coherence and follows the specific menu items requested, despite some minor spelling repetitions. Vidu Q2 produces a more stylistically authentic chalkboard texture, but the text is largely unreadable and fails to accurately reflect the prompt's content.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
Imagen 3.0 Generate 002
- + High visual quality with realistic lighting and cinematic atmosphere.
- + Excellent fur and fabric textures on the horse and spacesuit.
- − Fails the negative constraint, depicting the astronaut on top of the horse.
- − The front right leg of the horse has an anatomical distortion near the hoof.
Vidu Q2
- + Vibrant, surreal color palette that fits the 'cinematic' and 'surreal' tags.
- + Dynamic composition with glowing trails and nebula effects.
- − Fails the specific negative constraint, also placing the astronaut on top of the horse.
- − The horse's hind legs are poorly formed and blend into the gaseous background.
Verdict: Both Imagen 3.0 and Vidu Q2 failed the difficult spatial reasoning requirement of the prompt ('horse on top, not vice versa'), instead providing the standard trope of an astronaut riding a horse. Imagen 3.0 is the superior image due to its more grounded, high-fidelity rendering and realistic textures, whereas Vidu Q2 suffers from messy anatomical details in the horse's legs.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
Imagen 3.0 Generate 002
- + Excellent 3D miniature clay-like texture that matches the cartoon aesthetic perfectly.
- + Accurate text rendering and placement as requested in the prompt.
- + Perfectly executed isometric perspective and clean composition.
- − Text is slightly offset to the left rather than perfectly top-center.
- − The flag icon is a simple graphic rather than a 3D asset like the sushi.
Vidu Q2
- + Strong 3D glossy materials that give it a premium polished look.
- + Excellent text placement and creative integration of the flag icon.
- + Higher level of detail in the food textures and variety.
- − Perspective is slightly flatter than the requested 45 degree isometric angle.
- − The base diorama is smaller and less prominent than specified.
Verdict: Imagen 3.0 Generate 002 better captures the requested isometric miniature aesthetic with its consistent clay-like textures and clean wooden base. However, Vidu Q2 offers a more dynamic composition with very professional text placement and centered alignment, resulting in a more 'finished' graphic design look despite falling slightly short on the isometric angle.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
Imagen 3.0 Generate 002
- + Excellent fur texture rendering and realistic lighting
- + High anatomical accuracy for each animal specifically mentioned in the prompt
- + Clean composition with a clear focal point and shallow depth of field
- − Static posing rather than the requested 'chasing butterflies' action
Vidu Q2
- + Successfully captures the 'playfully chasing' action requested in the prompt
- + Includes all requested animal types plus extra to fill the scene
- + Lively energy that fits the 'joyful wholesome vibe'
- − Noticeable anatomical issues including a puppy with five legs
- − Lower resolution and less realistic fur texture compared to Model A
- − The butterflies appear flattened and less integrated into the 3D space
Verdict: Imagen 3.0 Generate 002 produces a far more polished and photorealistic image with superior textures and lighting, though the animals are sitting relatively still. Vidu Q2 captures the movement of the prompt better but suffers from significant anatomical errors, most notably a puppy with an extra leg and less coherent detail. Imagen 3.0 is preferred for its high technical quality and clear rendering.
Explore each model
ShengShu Technology's text-to-image and reference-to-image model with support for character consistency and multi-reference image processing