Google's latest Imagen 4.0 text-to-image generation model with significantly better text rendering and overall image quality
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
Imagen 4.0 Generate 001
#55 of 62 in Text-to-Image
Stable Diffusion 3.5 Large
#29 of 62 in Text-to-Image
Where the votes landed
Imagen 4.0 Generate 001
50.0%
win rate
Ties
0.0%
Stable Diffusion 3.5 Large
50.0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
Imagen 4.0 Generate 001
- + Perfect adherence to object placement with the red book sitting on top of the cube.
- + High photorealism with convincing material textures on the book and table.
- + Excellent rendering of light and reflections within the glass cube.
- − The sphere appears to be floating mid-air, which might seem physically improbable without support.
Stable Diffusion 3.5 Large
- + Good realistic lighting and shadow work.
- + Includes the green plant and wooden table as requested.
- − Failed the spatial relationship prompt by placing the red book inside the cube instead of on top of it.
- − The glass cube has visible glue seams and sharp edges that look more like acrylic than solid glass.
Verdict: Imagen 4.0 followed all spatial instructions perfectly, placing the book on top of the glass cube and the sphere inside it. Stable Diffusion 3.5 Large failed the core spatial reasoning of the prompt by placing the book on the table inside the cube, resulting in a less accurate output despite its high visual quality.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
Imagen 4.0 Generate 001
- + Excellent skin texture and facial detail showing age and expression.
- + High-quality rendering of raindrops on clothing and metallic bicycle surfaces.
- + Strong bokeh effect that creates a cinematic street atmosphere.
- − The 'motion blur from passing cars' is minimal; the taxi looks almost stationary.
- − The bicycle mechanics are slightly nonsensical where the tool meets the derailleur.
Stable Diffusion 3.5 Large
- + Better capture of the 'motion blur' aspect with the car in the background.
- + Stronger depiction of visible rainfall in the air.
- + Candid, wide framing that captures more of the street environment.
- − The subject's hands and the bicycle handlebar area exhibit significant anatomical and structural distortion.
- − The skin texture and facial features are less refined and clear compared to the other model.
- − The bicycle's physical structure, especially the basket and frame junction, is messy.
Verdict: Imagen 4.0 provides a much higher level of detail and realism in the subject's face and clothing, effectively capturing the 'natural skin texture' requested in the prompt. While Stable Diffusion 3.5 Large does a better job of conveying motion in the background and the atmosphere of rain, it fails on technical details, particularly the distortion of the man's hands and the bicycle's geometry. Imagen 4.0 is preferred for its superior clarity and realistic cinematic quality.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
Imagen 4.0 Generate 001
- + Excellent intricate engraving on the plate armor with logical reflections
- + Includes detailed leather straps and buckles as requested
- + Strong dramatic lighting from the torch source with clear bokeh sparks
- − The 'beads' in the hair look more like metal tech-connectors or tubes rather than small decorative beads
- − The skin texture across the forehead appears slightly artificial/waxy despite the scarring
Stable Diffusion 3.5 Large
- + Exceptional realism in the facial features and lifelike eyes
- + Very natural-looking braids and hair texture
- + Perfect depiction of facial dirt and blood for a 'battle-worn' appearance
- − Missed the 'small beads' in the hair prompt entirely
- − The texture on the cloth underlayer is less defined compared to the armor detail
Verdict: Imagen 4.0 provides a more fantastical and decorative interpretation with highly detailed armor and leatherwork, though some elements like the hair beads look a bit robotic. Stable Diffusion 3.5 Large achieves a much higher level of facial realism and gritty atmosphere, but it fails to include the requested beads in the hair. Stable Diffusion is the winner for its superior skin textures and convincing 'battle-worn' aesthetic.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
Imagen 4.0 Generate 001
- + Excellent grid layout that integrates text and images cohesively.
- + Clean, high-quality photography with vibrant colors and professional lighting.
- + Accurate categorization of sections including Appetizers, Pizza, and Mains.
- − Text is largely gibberish despite looking visually clean.
- − Some layout graphics like secondary boxes are a bit distracting.
Stable Diffusion 3.5 Large
- + Strong 'Menu' header with excellent bold sans-serif typography.
- + High density of images and a professional narrow-column layout.
- + Interesting use of dividers and sub-headers.
- − The grid structure is less intentional, feeling more like a vertical scroll than a single-page design.
- − Noticeable spelling errors in simple category names, such as 'MAIMAES' for Mains.
- − The food photos on the side are crowded and lack the breathing room found in Model A.
Verdict: Imagen 4.0 Generate 001 provides a much cleaner and more professional restaurant menu layout, perfectly balancing high-quality photography with readable text sections. While Stable Diffusion 3.5 Large has a striking header and a unique vertical layout, it fails to organize the grid effectively and has significant spelling errors in the primary category headers.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
Imagen 4.0 Generate 001
- + Perfect text rendering for all requested copy.
- + Excellent 'exploded' view with clear separation of ingredients.
- + Very clean, high-resolution aesthetic suitable for actual advertising.
- − The fiery effect on the text is subtle rather than prominent.
Stable Diffusion 3.5 Large
- + Great visual energy with actual flames and embers.
- + Good photorealism on the texture of the meat and bun.
- + Dynamic lighting that matches the fiery environment.
- − Completely failed to include any of the requested text.
- − The burger is floating but not 'exploded' or separated as requested.
- − The starburst element is missing entirely.
Verdict: Imagen 4.0 followed the prompt instructions near-perfectly, delivering a professional-grade advertisement with an exploded layout and all requested text rendered accurately. Stable Diffusion 3.5 Large produced a high-quality artistic image of a burger on fire, but failed on almost every specific technical instruction including the explosive deconstruction and all text elements. Imagen 4.0 is the clear winner for its superior prompt adherence.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
Imagen 4.0 Generate 001
- + Excellent text legibility and spelling for most of the primary requested items.
- + Captures the requested handwriting style with natural variations as described.
- + Strong adherence to the specific formatting and year (2026) requested.
Stable Diffusion 3.5 Large
- + Creates a beautiful aesthetic scene including the café interior.
- + Chalk texture looks very realistic with smudging and varied pressure.
- + Good architectural and environmental lighting.
- − Significant spelling errors throughout, including the title 'TODAAY'.
- − Failed to render the correct year (2024 instead of 2026).
- − The font style is more blocky and digital-feeling rather than the requested elegant cursive.
Verdict: Imagen 4.0 significantly outperformed Stable Diffusion 3.5 Large by correctly following the complex text instructions and date requirements while maintaining high legibility. While Stable Diffusion 3.5 Large produced a more visually context-rich café scene, it failed on almost every textual detail and spelling check within the prompt.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
Imagen 4.0 Generate 001
- + Excellent visual clarity and high-resolution details on the space suit
- + Vibrant and cinematic color palette with rainbow light trails
- + Effective use of surrealism with the galaxy patterns on the horse's flank
- − The composition is a bit static compared to the sense of motion in the other image
- − The horse's mane and tail look slightly stylized/digital rather than natural
Stable Diffusion 3.5 Large
- + Dynamic sense of movement with the galloping pose and cosmic dust clouds
- + Impressive atmospheric integration of the horse and rider within the planetary background
- + Realistic textures on the astronaut's gear and the horse's coat
- − The horse's face has some anatomical distortion specifically around the muzzle and eyes
- − Slightly less 'clean' than Imagen 4.0, with some graininess in the star fields
Verdict: Both models followed the prompt successfully, including the specific instruction for the rider's orientation. Imagen 4.0 produced a cleaner, more vibrant, and balanced composition with a striking 'surreal' aesthetic, while Stable Diffusion 3.5 Large excelled at creating a sense of scale and dynamic action, though it suffered from minor anatomical artifacts on the horse.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
Imagen 4.0 Generate 001
- + Excellent adherence to the complex composition with both capybara and passenger visible.
- + Captures the bored business expression and phone usage perfectly.
- + Includes the requested 'TAXI' cap and clear night city lights through the window.
- − The capybara's paws look more like bird talons than capybara feet.
- − The taxi's roof sign is oddly duplicated or appearing inside the car frame.
Stable Diffusion 3.5 Large
- + Features a highly detailed capybara with realistic fur texture.
- + The lighting on the capybara's face is dramatic and visually appealing.
- + Good rendering of the dark jacket and yellow cap.
- − Completely missed the secondary subject/businesswoman in the back seat.
- − The capybara only has one hand near the wheel instead of both front paws as requested.
- − The capybara's ear is incorrectly positioned on top of the hat.
Verdict: Imagen 4.0 followed the complex prompt instructions much more accurately, successfully depicting the interaction between the capybara driver and the bored businesswoman passenger. Stable Diffusion 3.5 Large produced a high-quality close-up of the capybara but failed to include the passenger and didn't place the paws on the steering wheel as specified.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
Imagen 4.0 Generate 001
- + Perfect text rendering for all lines including the specific event details.
- + Excellent gothic visual style with a polished, professional finish.
- + Clear adherence to the central glowing jack-o-lantern requirement.
- − The parchment roll on the left side is a bit chunky and disrupts the framing symmetry.
Stable Diffusion 3.5 Large
- + Strong 'vintage' aesthetic with distressed parchment edges.
- + Includes all thematic elements like bats, webs, and twisted trees.
- + Dynamic lighting with the large moon.
- − Failed to include the specific event details (Date, Time, Location) at the bottom.
- − The banner text is significantly distorted and contains garbled characters.
- − The jack-o-lantern is not the central focus as requested.
Verdict: Imagen 4.0 followed the prompt instructions perfectly, rendering all the specific text details accurately within an elegant gothic layout. Stable Diffusion 3.5 Large captured the vintage aesthetic well but failed on technical prompt adherence, missing the event details and producing garbled text on the banner.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
Imagen 4.0 Generate 001
- + Excellent photorealistic textures and PBR materials, especially on the tuna and fish roe.
- + Clean and soft lighting that matches the 'gentle lighting' prompt perfectly.
- + Simple and elegant composition on a cylindrical base.
- − Completely failed to include any of the requested text ('JAPAN', 'SUSHI').
- − Background color is a neutral grey/white rather than the requested solid light blue.
- − Missing the flag icon.
Stable Diffusion 3.5 Large
- + Followed all text prompts accurately including the bold words 'JAPAN' and 'SUSHI'.
- + Accurately rendered the solid light blue background and the small flag icon.
- + Strong isometric 3d miniature aesthetic with a clear diorama base.
- − The text is placed on a sign rather than at 'top-center' of the image frame as requested.
- − The scene is a bit cluttered compared to the request for 'minimal garnish'.
- − Minor visual artifacts in the chopsticks and some sushi pieces.
Verdict: While Imagen 4.0 produces a much more visually pleasing and realistic-looking sushi miniature, it failed to follow nearly half of the specific prompt instructions regarding text, background color, and icons. Stable Diffusion 3.5 Large adhered to every part of the prompt, including the complex text and icon requirements, making it the more successful model for this specific challenge despite slightly less refined textures.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
Imagen 4.0 Generate 001
- + Perfect adherence to all requested animal types
- + Excellent clarity and sharp focus throughout the frame
- + Detailed depiction of dew drops and diverse wildflower species
- − The composition feels a bit crowded and illustrative rather than 'photorealistic'
- − The lighting effects such as god rays appear a bit artificial
Stable Diffusion 3.5 Large
- + Achieves a more convincing 'photorealistic' depth of field and soft lighting
- + Captures a more dynamic sense of movement and 'joyful' expression
- + Beautiful bokeh and natural integration of sun glares
- − Failed to include a 'tabby' kitten, showing a solid ginger kitten instead
- − The back legs of the fox and puppy are somewhat muddled or missing in the grass
Verdict: Imagen 4.0 Generate 001 followed the prompt's specific subject list more accurately by correctly including a tabby kitten, whereas Stable Diffusion 3.5 Large opted for a ginger one. However, Stable Diffusion 3.5 Large produced a much more photorealistic and emotionally resonant image with better lighting and depth, while Imagen 4.0 leaned toward a saturated, illustrative aesthetic.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
Imagen 4.0 Generate 001
- + Excellent text rendering with accurate spelling and accent mark.
- + Clean vector execution that perfectly captures the minimalist aesthetic.
- + High contrast and well-balanced composition.
- − The 'Est. 1720' banner is quite small and lacks detail.
- − Misses the 'subtle texture' request on the background compared to Model B.
Stable Diffusion 3.5 Large
- + Beautiful background texture and vintage corners that match the prompt's mood.
- + Creative use of decorative elements and various steam shapes.
- + Rich warm brown and cream color palette.
- − Major spelling error in the main text ('Cafféé').
- − The cloche dome icon is poorly formed with odd gaps and floating elements.
- − The banner element is busy and overlaps awkwardly with bottom flourishes.
Verdict: Imagen 4.0 delivers a much more professional and usable logo emblem with perfect typography and clean vector lines, whereas Stable Diffusion 3.5 Large struggles with core text spelling and the integrity of the central icon. While Stable Diffusion captures the requested vintage texture better, the structural failures in the design and spelling make it inferior to the clean execution of Imagen 4.0.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
Imagen 4.0 Generate 001
- + Excellent text rendering with clear, legible typography.
- + Followed the specific 6-step logical flow and NASA color palette.
- + Clean, modern flat-vector aesthetic that looks professional.
- − The logic of the diagram icons is a bit repetitive (Earth and Moon used multiple times incorrectly in some steps).
- − The Saturn V rocket misses some proportion accuracy.
Stable Diffusion 3.5 Large
- + High visual complexity and texture in the lunar surface.
- + Intricate layout that feels like a dense technical schematic.
- − Failed significantly on text rendering, resulting in 'gibberish' characters.
- − Depicted a Space Shuttle-style orbiter instead of the requested Saturn V for the Apollo mission.
- − Did not follow the specific 6-step sequence or provide clear iconography.
Verdict: Imagen 4.0 significantly outperformed Stable Diffusion 3.5 Large by correctly following the structured 6-step infographic request and rendering legible, accurate text. Stable Diffusion 3.5 Large suffered from typical AI text issues and hallucinated a Space Shuttle, which is historically inaccurate and contrary to the prompt's request for a Saturn V.
Explore each model
Stability AI's 8.1-billion parameter Multimodal Diffusion Transformer (MMDiT) text-to-image model featuring improved image quality, typography, complex prompt understanding, and resource-efficiency