Google's Imagen 4.0 Fast model optimized for speed and efficiency, suitable for high-volume image generation tasks
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
Imagen 4.0 Fast Generate 001
#52 of 62 in Text-to-Image
Wan 2.7
#39 of 62 in Text-to-Image
Where the votes landed
Imagen 4.0 Fast Generate 001
0%
win rate
Ties
0%
Wan 2.7
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
Imagen 4.0 Fast Generate 001
- + Excellent photorealism and lighting
- + Captures the interaction of light and glass perfectly
- + Clear and accurate rendering of the blue sphere and the plant through the glass
- − The plant is more besides the cube than behind it, though it is still partially visible through the corner
- − The 'cube' has a mirrored base which wasn't specifically requested but adds to the aesthetic
Wan 2.7
- + Highly accurate spatial arrangement per the prompt
- + Excellent rendering of the plant visible through multiple faces of the glass cube
- + Realistic textures on the wooden table and red book
- − The glass cube has internal vertical lines that look slightly like seams or reflections of frames that aren't there
- − The sphere appears more like a matte ball rather than a polished sphere
Verdict: Both models followed the prompt very well, but Wan 2.7 adhered more strictly to the spatial requirement of placing the plant directly behind the cube so it is heavily visible through the glass. While Imagen 4.0 has slightly more refined lighting and cleaner line-work, Wan 2.7 provides a more complex and accurate composition based on the specific arrangement requested.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
Imagen 4.0 Fast Generate 001
- + Excellent handling of wet pavement reflections and realistic rain depth
- + Very convincing skin texture and natural facial expression
- + Captures the 'imperfect framing' prompt with a creative POV shot
- − The bike anatomy is slightly warped near the handlebars
- − The foreground blur serves as a frame but might be too obscuring for some tastes
Wan 2.7
- + Authentic street environment that feels distinctly like a Japanese neighborhood
- + Good portrayal of wet clothing and rainy atmosphere
- + Accurate representation of a 50mm lens perspective
- − The bicycle rendering is physically impossible with missing spokes and disconnected frame parts
- − Lacks the requested motion blur from passing cars
- − The face and hands have a slightly waxy, less realistic texture
Verdict: Imagen 4.0 Fast Generate 001 provides a significantly more cinematic and high-quality image, successfully incorporating specific details like motion blur and highly realistic skin textures. While Wan 2.7 captures the environmental aesthetic well, it suffers from severe structural issues with the bicycle and a flatter overall rendering.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
Imagen 4.0 Fast Generate 001
- + High resolution image of an elderly man
- + Natural-looking forest environment
- − Completely failed to follow the prompt instructions regarding armor, braiding, and lighting
- − Incorrect composition (full body instead of close portrait)
- − Incorrect setting and attire for a paladin
Wan 2.7
- + Excellent adherence to all prompt details including braided hair with beads and engraved plate armor
- + Beautiful warm torchlight lighting and bokeh effects
- + Highly detailed textures on skin, scars, and leather straps
- − Slightly repetitive patterns in the armor engraving
- − Minor blurring on the very edges of the hair braids
Verdict: Imagen 4.0 Fast Generate 001 completely failed the prompt, delivering a modern man in a leather jacket in a garden instead of a paladin. Wan 2.7 followed every instruction perfectly, producing a high-quality, cinematic portrait that captures the battle-worn aesthetic and specific technical details requested.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
Imagen 4.0 Fast Generate 001
- + Strong adherence to the requested grid-based layout.
- + Includes vibrant colored accents as requested.
- + Features clear, bold sans-serif headings for sections.
- − Poor text legibility with many gibberish characters.
- − Repetitive food photos, showing similar pizzas multiple times.
- − Spelling error in the main heading 'APETIERS'.
Wan 2.7
- + Excellent text readability with coherent English words and realistic pricing.
- + Varied food photography that matches diverse menu categories.
- + High-quality graphic design elements like a logo, QR code, and social media icons.
- − The grid is slightly less distinct than Model A's interpretation.
- − Less 'minimalist' in style due to excessive decorative elements like the pen and rosemary outside the menu.
Verdict: Imagen 4.0 Fast Generate 001 provides a layout that strictly follows the requested grid structure and minimalist white aesthetic, but fails significantly on text accuracy and variety. Wan 2.7 produces a much more professional and usable result with legible text, high-quality diverse imagery, and a polished commercial feel, despite adding unnecessary background props. Wan 2.7 is preferred for its realistic execution of a functional menu.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
Imagen 4.0 Fast Generate 001
- + Excellent typography rendering with crisp, clean text lines.
- + Strong photorealistic textures on the bun and meat patties.
- + Clean composition that feels like a professional studio advertisement.
- − The 'exploded' effect is a bit static and vertically aligned rather than dynamic.
- − The background is more generic and lacks the intense fiery texture requested.
Wan 2.7
- + Highly dynamic and energetic composition with great sense of motion.
- + Excellent interpretation of the fiery theme with flames integrated into the text and background.
- + More creative 'exploded' view with flying ingredients like sauce splashes and seeds.
- − Text rendering is slightly less refined with some clumping in the 'MAGIC BURGER' font.
- − The image style leans more toward digital illustration than the requested 'photorealistic' detail.
Verdict: Imagen 4.0 Fast Generate 001 produces a very professional, clean, and photorealistic ad with superior text legibility. However, Wan 2.7 captures the 'dynamic' and 'fiery' spirit of the prompt much better, offering a more exciting visual arrangement of the ingredients. Wan 2.7 is the preferred choice for its superior creativity and adherence to the high-energy motion requested.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
Imagen 4.0 Fast Generate 001
- + Successfully rendered most requested text accurately.
- + Includes visible chalk-like smudges at the bottom of the board for realism.
- − Has spelling errors such as 'Octuphus' and 'Cookes'.
- − Text appears too uniform and clean, lacking the natural 'handwritten' slant and variation requested.
- − The 'Brown Butter' line is repeated and messy.
Wan 2.7
- + Excellent spelling accuracy across all complex menu items.
- + Captures the 'elegant cursive' style for the title much better.
- + The overall image composition including background elements creates a stronger 'cozy café' atmosphere.
- − The text has a drop-shadow effect that makes it look slightly digital rather than authentic chalk strokes.
- − One hanging lantern on the left is missing its bottom or base.
Verdict: Wan 2.7 is the clear winner as it accurately rendered all specific menu items with zero spelling errors and followed the stylistic request for an elegant cursive title. Imagen 4.0 Fast Generate struggled with spelling ('Octuphus') and layout, creating a redundant and cluttered list for the final item.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
Imagen 4.0 Fast Generate 001
- + Excellent cinematic lighting and gaseous nebula background
- + Highly detailed rendering of the spacesuit and horse's dapple texture
- + Dynamic pose with the horse rearing up
- − Prompt adherence failure: the astronaut is riding the horse, not the horse on top of the astronaut
Wan 2.7
- + Clear composition with the curvature of Earth in the background
- + Good anatomical rendering of the horse and detailed tack
- + Sharp focus and high resolution
- − Prompt adherence failure: the astronaut is riding the horse, ignoring the specific spatial instruction
Verdict: Both Imagen 4.0 Fast Generate 001 and Wan 2.7 failed the specific spatial logic test in the prompt ('horse on top, not vice versa'), instead providing the standard Interpretation of an astronaut riding a horse. Imagen 4.0 is slightly preferred for its more cinematic atmosphere and lighting, although both models were equally unsuccessful at the surreal logic challenge.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
Imagen 4.0 Fast Generate 001
- + Excellent fur texture and lighting on the capybara's face
- + Correct interior perspective which emphasizes the 'inside the taxi' request
- + Perfectly captures the 'bored' expression of the passenger
- − The capybara's paws look more like human-monkey hybrid hands with long sharp claws
- − The driver appears to be on the right side of the car, which is incorrect for a NY taxi
Wan 2.7
- + Successfully places the passenger and driver beside each other for a clear view of both faces
- + Accurately depicts the capybara's paws on the steering wheel
- − The passenger is sitting in the front passenger seat instead of the requested back seat
- − The capybara's head has a strange mask-like quality where it meets the neck
- − Composition feels less like being 'inside' the car and more like looking through a window
Verdict: Imagen 4.0 Fast Generate 001 provides a much more cinematic and high-quality image that correctly places the businesswoman in the back seat as requested. While Wan 2.7 has better paw anatomy, it fails the spatial prompt by putting the passenger in the front seat and has lower overall visual fidelity.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
Imagen 4.0 Fast Generate 001
- + Atmospheric moody lighting with a cinematic feel
- + Clean, modern parchment texture with torn edges
- − Major spelling errors in the title ('IINVIIATION') and footer ('NIGIT OF FNIGITS')
- − Formatting of the event details at the bottom is cramped and poorly aligned
Wan 2.7
- + Excellent text rendering with no spelling errors in any section
- + Highly detailed gothic border featuring thorns, webs, and skulls as requested
- + Better use of space and font choice for an 'elegant gothic' look
- − The illustration style is more comic-book/illustrative than 'cinematic' life-like lighting
Verdict: While Imagen 4.0 Fast Generate 001 creates a more atmospheric and moody background, it fails significantly on text accuracy with multiple glaring spelling errors ('IINVIIATION'). Wan 2.7 delivers a perfectly legible invitation with intricate gothic details that closely follow the prompt's request for thorns and elegant text.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
Imagen 4.0 Fast Generate 001
- + Excellent adherence to the 'minimal garnish' instruction.
- + Very clean and professional font rendering.
- + Good 3D isometric perspective on the diorama base.
- − The sushi models are a bit simplified, almost looking like plastic toys rather than refined food.
- − Composition feels slightly empty with only two pieces of sushi.
Wan 2.7
- + Beautiful Material rendering with realistic sub-surface scattering on the fish.
- + Dynamic and colorful variety of sushi pieces.
- + Stylized text is bold and fits the 'cartoon scene' aesthetic perfectly.
- − Included extra items like chopsticks and soy sauce not requested in the 'minimal garnish' prompt.
- − The 'JAPAN' text has a small artifact on the 'N'.
Verdict: Both models followed the prompt's layout and theme excellently. Imagen 4.0 Fast Generate 001 provides a cleaner, more minimal diorama, while Wan 2.7 delivers much higher quality textures and a more appealing variety of sushi pieces. Wan 2.7 is the preferred choice for its superior visual quality and more interesting interpretation of 'refined textures'.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
Imagen 4.0 Fast Generate 001
- + Excellent fur texture and lighting details on the animals
- + The soft sunset lighting is very natural and warm
- + High level of realism in the facial structure of the fox and puppy
- − Failed to include butterflies as requested in the prompt
- − The animals are sitting still rather than 'playfully chasing and tumbling'
- − The cat is a solid black kitten rather than the requested tabby kitten
Wan 2.7
- + Accurately includes all subjects including butterflies and a variety of wildflowers
- + Perfectly captures the 'playfully chasing' and dynamic movement requested
- + Better adherence to specific animal breeds like the golden retriever and tabby kitten
- − The fox's anatomy looks slightly stiff compared to the other animals
- − Some minor artifacting around the butterfly wings
- − The composition feels slightly more digital/rendered than Model A
Verdict: Wan 2.7 is the clear winner as it followed every part of the complex prompt, including the motion, the specific animal markings (tabby kitten), and the presence of butterflies. While Imagen 4.0 produced a very high-quality static portrait with beautiful lighting, it failed to capture the action and specific visual elements requested.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
Imagen 4.0 Fast Generate 001
- + Excellent typography accuracy including the accent on the 'E'.
- + Clean vector aesthetic that feels modern yet vintage.
- + Well-balanced composition with a clear central cloche icon.
- − Includes small, nonsensical filler text ('AFFD CARO') above the main name.
- − The steam effect is a bit faint and could be more pronounced.
Wan 2.7
- + Beautiful warm cream and brown tone palette with nice border detailing.
- + Clear steam illustration that matches the vintage style.
- + High-quality vector look with consistent line weights.
- − Misspelled the primary name as 'Florion' instead of 'Florian'.
- − The cloche dome appears transparent, which is slightly unusual for a minimalist logo.
Verdict: Imagen 4.0 Fast Generate 001 provides a much more professional and accurate logo, successfully following the textual requirements and spelling the brand name correctly. While Wan 2.7 has a pleasing color palette and nice decorative elements, the spelling error on the main subject makes the logo unusable for the specific prompt.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
Imagen 4.0 Fast Generate 001
- + Features a very clean, minimalist aesthetic that fits the vector style request.
- + Icons are well-balanced and professionally spaced.
- + Successfully followed the specific color palette requested.
- − Numerous spelling errors including 'APOLO', 'MOOR', and 'MOO+ON'.
- − The logical flow of the diagram is confusing and non-linear.
- − Included irrelevant 'Saturn Vicon' text and literal '+' symbols from the prompt description.
Wan 2.7
- + Perfectly followed the 6-step logical sequence in a clear, vertical hierarchy.
- + Text rendering is highly accurate with only minor typos like 'Descript' or 'Tranquiliry'.
- + Captures the NASA aesthetic perfectly with excellent detail-oriented iconography.
- − The white border around the dark navy poster is slightly inconsistent.
- − Some secondary text at the bottom is small and slightly blurry.
Verdict: Wan 2.7 significantly outperformed Imagen 4.0 in logic, spelling, and prompt adherence, providing a clear 6-step vertical infographic that makes sense as an educational poster. Imagen 4.0 failed to create a logical sequence and featured significant spelling errors in the largest headers, such as 'APOLO' and 'MOOR'.
Explore each model
Alibaba's Wan 2.7 image generation and editing model for text-to-image, reference-guided generation, and instruction-based image edits