Google's Imagen 3.0 text-to-image generation model, producing high-quality images with improved detail and lighting
Settled by community votes across 6 shared challenges, with an AI judge weighing in on each.
Imagen 3.0 Generate 002
#36 of 62 in Text-to-Image
Wan 2.5 (Preview)
#28 of 62 in Text-to-Image
Where the votes landed
Imagen 3.0 Generate 002
0%
win rate
Ties
0%
Wan 2.5 (Preview)
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
Imagen 3.0 Generate 002
- + Excellent adherence to the 'grid' prompt requirement.
- + Clean, professional aesthetic suitable for a high-end casual restaurant.
- + Realistic food photography within the layout.
- − Sections for Appetizers/Pizza/Mains are repeated or fragmented across the grid incorrectly.
- − Text is largely illegible gibberish.
Wan 2.5 (Preview)
- + Clear and logical hierarchy for Appetizers, Pizza, and Mains sections.
- + Vibrant color accents through colored lines and headers that enhance the 'modern' feel.
- + Clean text rendering for headers.
- − Food photos are repetitive, with most containing the same basil garnish.
- − Contains several spelling errors like 'Menue' and 'Restormalit'.
- − Less professional photography style compared to the other model.
Verdict: Imagen 3.0 produces a more sophisticated visual design that looks like a real-world magazine-style menu, though its categorization is messy. Wan 2.5 (Preview) better follows the instructions for structured sections (Appetizers, Pizza, Mains) and vibrant accents, making it more functional as a layout template despite spelling errors and repetitive food imagery.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
Imagen 3.0 Generate 002
- + Excellent photorealistic rendering of food textures, especially the bun and patty.
- + Accurate rendering of the '€6.99' price in a starburst sticker.
- + High clarity and sharp focus on the central subject.
- − The text 'MAGIC BURGER' lacks the requested fiery, glowing effect, appearing more like metallic gold.
- − The starburst includes gibberish text above the price.
- − The explosion effect feels slightly stiff and vertical rather than dynamic.
Wan 2.5 (Preview)
- + Successfully applied the fiery, glowing effect to all text elements as requested.
- + Very dynamic and chaotic 'exploded' composition with ingredients drifting at various angles.
- + Strong atmosphere with a mix of smoke, embers, and dripping effects.
- − The burger ingredients look slightly less photorealistic and more illustrative compared to Image A.
- − The bottom bun is tilted at an awkward angle that feels disconnected from the rest of the stack.
Verdict: Both models followed the prompt well, but Wan 2.5 (Preview) captured the requested artistic style and fiery text effects much more effectively than Imagen 3.0. While Imagen 3.0 has superior photorealism in the food itself, Wan 2.5 (Preview) delivered a more cohesive ad concept with the dynamic movement and glowing typography requested.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
Imagen 3.0 Generate 002
- + Features a very high-quality wood grain and realistic blackboard surface texture.
- + Excellent chalk-like stroke details including powder residue and fading.
- − Significant spelling errors throughout the menu items such as 'Risottto', 'Octpsuin', and 'Choouip'.
- − Repeats the same lines of text with slightly different typos instead of continuing the prompt.
Wan 2.5 (Preview)
- + Excellent prompt adherence with near-perfect spelling for all menu items.
- + Realistic chalkboard aesthetic with smears and a dynamic, cozy café background.
- + The handwriting style is consistent and convincingly handmade across the entire board.
- − One minor text overlap where 'Cookies' repeats the price of '$9' on a separate line.
- − The title font is more of a script than a strictly 'elegant cursive' style as requested.
Verdict: Wan 2.5 (Preview) is the clear winner as it successfully rendered the complex menu items with correct spelling and a beautiful, consistent handwriting style. Imagen 3.0 struggled significantly with the text, producing numerous typos and repeating lines of text rather than finishing the menu items requested.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
Imagen 3.0 Generate 002
- + Excellent realistic lighting and shadows on the horse's muscles
- + High-quality nebulous background with a cinematic lighting feel
- − Failed the negative constraint; the astronaut is on top, not the horse
- − The reins are handled somewhat awkwardly into the astronaut's hand
Wan 2.5 (Preview)
- + Dynamic composition with the inclusion of Earth and a galaxy
- + Strong lighting on the space suit and horse's coat
- − Failed the negative constraint; the astronaut is riding the horse
- − The horse's front legs have anatomical blurring/speed artifacts that look like extra limbs
Verdict: Both models completely failed the negative constraint to have the 'horse on top' of the astronaut, instead providing a standard 'astronaut riding a horse' image. Imagen 3.0 (Model A) is slightly preferred for its superior anatomical rendering and more realistic, cinematic lighting, whereas Wan 2.5 (Model B) has messy motion blur on the horse's legs and slightly less coherent textures.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
Imagen 3.0 Generate 002
- + Excellent adherence to the 45° isometric perspective.
- + Includes a variety of sushi types as part of the scene.
- + The text and flag icon are neatly placed in the corner, keeping the focal point clear.
- − The text is placed in the top-left corner instead of the requested top-center.
- − The 3D clay style is a bit flat compared to the requested realistic PBR materials.
Wan 2.5 (Preview)
- + Perfectly centered text and flag icons as per the prompt.
- + Higher quality lighting and material rendering with a glossy finish.
- + Clean, minimalist composition that feels more professional.
- − Failed the isometric perspective, opting for a standard eye-level/front-facing view.
- − Rendered only a single piece of sushi rather than a scene.
Verdict: Imagen 3.0 followed the technical perspective and layout instructions much more accurately, creating a true isometric miniature scene. Wan 2.5 produced a more visually striking image with superior lighting and centered text, but failed significantly on the isometric camera angle requirement.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
Imagen 3.0 Generate 002
- + Excellent photographic realism in fur texture and lighting.
- + Natural, proportional anatomy and convincing poses for all four animals.
- + Beautifully soft, diffused golden hour lighting that feels authentic.
- − The fox kit has somewhat dark eyes that don't match the 'big expressive eyes' prompt as well as others.
- − Slightly less movement/action compared to the 'chasing' part of the prompt.
Wan 2.5 (Preview)
- + Captures the 'chasing and tumbling' action much more dynamically with running poses.
- + Strong execution of god rays and 'dew sparkles' using floating water droplets.
- + Very expressive eyes that stand out.
- − The fox's eyes look unnaturally blue and crystalline, veering away from photorealism.
- − General anatomy feels slightly more stylized and 'CG' compared to the puppy in Model A.
- − Minor artifact on the kitten's front paw.
Verdict: Imagen 3.0 Generate 002 produces a superior photorealistic masterpiece with incredible fur detail and natural lighting, feeling like a genuine photograph. While Wan 2.5 (Preview) does a better job of capturing the dynamic 'chasing' action and the specific 'dew sparkles' requested, its overall aesthetic is slightly more digital and less grounded in reality, particularly with the fox's eyes.
Explore each model
Alibaba's text-to-image and image-to-image generation model from the Wan AI suite, offering high-quality visual generation capabilities