Google's Imagen 3.0 text-to-image generation model, producing high-quality images with improved detail and lighting
Settled by community votes across 6 shared challenges, with an AI judge weighing in on each.
Imagen 3.0 Generate 002
#36 of 62 in Text-to-Image
Wan 2.6
#28 of 62 in Text-to-Image
Where the votes landed
Imagen 3.0 Generate 002
0%
win rate
Ties
0%
Wan 2.6
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
Imagen 3.0 Generate 002
- + Stronger adherence to the 'grid' layout mention in the prompt
- + Clean and professional aesthetic using consistent block sizing
- + High-quality, distinct food photography for each tile
- − The text is largely gibberish symbols rather than legible letters
- − Redundant labeling and scattered placement of categories makes it harder to read as a functional menu
Wan 2.6
- + Includes vibrant color accents as requested in the prompt
- + Text structure is much more legible with clear pricing and item names
- + Excellent use of bold sans-serif fonts for hierarchy
- − The 'grid' layout is less balanced, with some photos being vertical and others horizontal
- − Some minor glitches in the price alignment and currency symbols
Verdict: Imagen 3.0 provides a more visually satisfying geometric grid that fits the 'modern minimalist' prompt well, but it fails significantly on text legibility. Wan 2.6 is the more functional design, offering vibrant accents, legible menu items, and realistic pricing, making it feel like a usable template despite a slightly less rigid grid structure.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
Imagen 3.0 Generate 002
- + Excellent photographic quality and lighting on the food items.
- + Clean and legible typography for the main title.
- + Precise execution of the 'exploded' concept with layers clearly separated.
- − The pricing starburst contains gibberish text.
- − The background fire feels a bit more static compared to the other model.
Wan 2.6
- + Strong sense of motion and dynamic energy with floating sauce and smoke.
- + Perfect adherence to the fiery typography request, including the title and price tag.
- + Excellent background ambiance with realistic glowing embers and flames.
- − The burger ingredients are less 'exploded' and more slumped together in the center.
- − The bun texture is slightly less sharp than in the competing image.
Verdict: While Imagen 3.0 Generate 002 produces a cleaner product shot with better food lighting, Wan 2.6 captures the 'Magic Burger' prompt's requested energy much more effectively. Wan 2.6 successfully integrates all text with the specific 'fiery' effect requested and creates a far more dynamic atmosphere with its use of smoke and flying ingredients.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
Imagen 3.0 Generate 002
- + Text is presented in a very clear, centered, and legible layout.
- + The chalkboard texture and wooden frame are clean and realistic.
- + Follows most numerical and date instructions correctly.
- − Several spelling errors and duplicated lines occur (e.g., 'Risottto', 'Octpsuin', 'Choouip').
- − The handwriting looks somewhat digital and uniform rather than having natural chalk-on-slate variability.
- − The title is in block letters rather than the 'elegant cursive' requested.
Wan 2.6
- + Excellent adherence to the 'elegant cursive' and handwritten chalk texture requirement.
- + Spelling is nearly perfect for all items including the 'Brown Butter Chocolate Chip Cookies'.
- + The cinematic perspective and chalk dust details create a cozy, realistic café atmosphere.
- − The perspective angle makes the bottom text slightly harder to read compared to a flat shot.
- − Small artifact on the '9' in the price for cookies.
Verdict: Wan 2.6 is the clear winner as it successfully rendered complex handwritten text with perfect spelling and the requested cursive style, whereas Imagen 3.0 made numerous spelling mistakes and duplicated lines. Wan 2.6 also better captured the requested 'chalk texture' and 'cozy café' atmosphere through realistic lighting and depth of field.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
Imagen 3.0 Generate 002
- + Excellent anatomical realism for both the horse and the astronaut's gear.
- + Clean, cinematic lighting with a high level of detail in the texture of the horse's coat.
- + Great depth and use of astronomical nebulae as a backdrop.
- − Failed to follow the specific spatial instruction 'horse on top, not vice versa'.
- − Standard composition that doesn't fully capture the 'surreal' aspect of the prompt.
Wan 2.6
- + Vibrant and dynamic color palette with striking lighting effects.
- + Highly detailed astronaut suit and intricate horse tack.
- − Failed to follow the specific spatial instruction 'horse on top, not vice versa'.
- − Presence of a ground/surface contradicts the 'in space' setting of the prompt.
Verdict: Both Imagen 3.0 and Wan 2.6 failed the negative constraint and specific positioning instruction, returning a standard 'astronaut riding a horse' image instead of the requested surreal inverted horse-on-human arrangement. Imagen 3.0 is slightly preferred for its more authentic 'deep space' feel, whereas Wan 2.6 includes ground and dust which detracts from the space setting.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
Imagen 3.0 Generate 002
- + Excellent soft 3D textures and lighting consistent with the cartoon prompt
- + Accurate isometric perspective with a clean, centered layout
- + Higher rendering quality of the rice and fish textures
- − Text is placed in the top-left rather than top-center as requested
- − Text and icons are slightly off-balance compared to the 3D scene
Wan 2.6
- + Perfectly followed layout instructions with text at top-center and small flag icon
- + Strong use of the diorama base concept
- + Bold, clear typography that aligns with the prompt's aesthetic
- − The sushi model is much smaller and has less relative detail than Model A
- − Text is somewhat oversized, occupying a large portion of the square frame
Verdict: While Imagen 3.0 has superior 3D modeling and texture quality, Wan 2.6 adhered much more strictly to the compositional layout requested, including the vertical placement of text and the flag icon. However, Imagen 3.0 provides a more visually appealing and detailed 'miniature' look that feels more like a finished 3D render.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
Imagen 3.0 Generate 002
- + Excellent clarity and sharpness in the animals' fur and eyes.
- + Beautiful, clean rendering of the golden hour lighting and dew drops in the grass.
- + Well-composed arrangement of all four requested animals with distinct features.
- − The posing is more static and ‘posed’ rather than ‘chasing and tumbling’.
- − The fox kit has a slightly unusual limb anatomy near the back.
Wan 2.6
- + Successfully captures dynamic movement with chasing and jumping poses.
- + Dramatic and beautiful god rays that perfectly match the sunrise request.
- + Whimsical atmosphere with floating seeds and more active butterflies.
- − Slightly lower texture detail on the animals' fur compared to Model A.
- − The kitten's facial expression is a bit distorted and less realistic.
Verdict: Imagen 3.0 Generate 002 produces a cleaner, higher-resolution look with better fine details on the fur, but the animals appear to be posing for a portrait. Wan 2.6 better captures the 'playfully chasing' action requested in the prompt and provides a more atmospheric scene with superior lighting effects, though it has slightly more artifacts in the kitten's face.
Explore each model
Alibaba's multimodal generation model from the Wan AI suite, supporting text-to-video, image-to-video, reference-to-video with audio, and text-to-image, in both Chinese and English