Google's latest Imagen 4.0 text-to-image generation model with significantly better text rendering and overall image quality
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
Imagen 4.0 Generate 001
#55 of 62 in Text-to-Image
Stable Diffusion 3.5 Medium
#58 of 62 in Text-to-Image
Where the votes landed
Imagen 4.0 Generate 001
0%
win rate
Ties
0%
Stable Diffusion 3.5 Medium
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
Imagen 4.0 Generate 001
- + Excellent photorealistic texture on the red book and wooden table.
- + Sophisticated handling of reflections and light within the glass cube.
- + High resolution with sharp, clean edges.
- − The glass cube appears more like a solid block of glass rather than a hollow container, making the sphere look embedded.
- − The plant is very blurred in the background rather than being distinctly visible 'through' the glass as requested.
Stable Diffusion 3.5 Medium
- + Successfully shows the plant clearly through the glass walls of the cube.
- + Accurately depicts the cube as a hollow structure.
- + Follows all prompt elements including color and spatial relationships.
- − Lower visual quality with noticeable noise and soft textures.
- − The red book looks more like a flat red slab and lacks realistic book details like pages or binding.
- − Lighting is somewhat flat compared to the other model.
Verdict: Imagen 4.0 produces a much higher quality, photorealistic image with beautiful lighting, but it struggles with the physics of the scene, making the sphere look suspended in solid resin. Stable Diffusion 3.5 Medium follows the specific spatial instructions better by showing the plant through the hollow glass, but the overall image quality and texture rendering are significantly weaker. Imagen 4.0 is preferred for its superior aesthetic and professional finish.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
Imagen 4.0 Generate 001
- + Excellent natural skin texture and facial details.
- + Realistic rendering of wet surfaces and raindrops on clothing.
- + Effective use of shallow depth of field and bokeh.
- − The hands and fingers are somewhat merged and anatomically unclear.
- − Lacks the requested motion blur on the passing car.
Stable Diffusion 3.5 Medium
- + Captures a very authentic 'candid' street photography aesthetic.
- + Good atmosphere with believable reflections and lighting.
- + Better adheres to the imperfect framing request.
- − Anatomical failure where the man's hands melt into the bicycle frame.
- − The bicycle's structure is nonsensical with the pedal floating away.
- − Lower resolution and visible AI artifacts on the face.
Verdict: Imagen 4.0 provides a much higher quality image with superior textures and lighting, though it feels a bit more posed than 'candid'. Stable Diffusion 3.5 Medium captures the specific aesthetic mood of a 50mm street snap very well, but it suffers from severe anatomical and structural hallucinations, particularly in how the man interacts with the bicycle.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
Imagen 4.0 Generate 001
- + Excellent structural clarity and sharpness in the engraved plate armor.
- + Superior technical rendering of the leather straps and metal buckles.
- + Clean, professional composition with clear bokeh and lighting sources.
- − The 'braided' hair with beads looks more like modern mechanical hair clips or tubes than fantasy braids.
- − The skin looks a bit too airbrushed and clean for a 'battle-worn' description.
Stable Diffusion 3.5 Medium
- + Outstanding lifelike eyes and intense facial expression that fits the character's role.
- + Highly realistic application of dirt and scars on the skin skin texture.
- + Authentic braided hair texture that follows the prompt literally.
- − The 'beads' in the hair are largely absent or unnoticeable compared to the prompt requirements.
- − The lighting is a bit overexposed on the forehead, washing out some fine detail.
- − Some minor grain/noise in the background bokeh.
Verdict: Imagen 4.0 excels in rendering the hardware and materials, such as the intricate engravings and leather textures, though it fails to capture the 'battle-worn' grit accurately. Stable Diffusion 3.5 Medium creates a much more compelling and lifelike character with superior facial detail and realistic weathering, making it the better interpretation of the prompt's mood despite slightly less sharp armor details.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
Imagen 4.0 Generate 001
- + Excellent typography with clear, readable sections for Appetizers, Pizza, and Mains.
- + High-quality, vibrant food photography that looks professional and appealing.
- + Clean, modern grid layout with effective use of white space and geometric accents.
- − The placeholder text for ingredient descriptions is nonsensical gibberish.
Stable Diffusion 3.5 Medium
- + Captures the concept of a long-form casual dining menu layout.
- + Includes a large variety of food items to fill the space.
- − Text is highly distorted, blurry, and illegible throughout.
- − The layout feels cluttered and lacks the 'modern minimalist' aesthetic requested.
- − Low visual fidelity with muddy colors and messy alignment.
Verdict: Imagen 4.0 significantly outperforms Stable Diffusion 3.5 Medium by delivering a high-resolution, professional-grade graphic design that adheres perfectly to the grid layout and section prompts. While both models struggle with generating actual words in the body text, Imagen 4.0 uses sharp, bold headers that look intentional, whereas Stable Diffusion 3.5 Medium produces blurry, distorted artifacts for all text and images.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
Imagen 4.0 Generate 001
- + Excellent adherence to the 'exploded' concept with clear separation of ingredients.
- + Very high quality, legible text rendering with the requested glowing effect.
- + Dynamic and professional composition suitable for a real advertisement.
- − The '€6,99' uses a comma instead of a decimal point.
- − The cheese texture appears slightly more plastic than photorealistic.
Stable Diffusion 3.5 Medium
- + Strong fiery atmosphere with impressive lighting on the burger.
- + Correct numerical value with a decimal point.
- + Good rendering of textures on the bun and patty.
- − Failed to create an 'exploded' burger, showing it mostly assembled instead.
- − Text is small, lacks the requested glowing effect, and has a slight spelling error in 'LIMITED'.
- − Comparison of the scale between the burger and the starburst is awkward.
Verdict: Imagen 4.0 significantly outperformed Stable Diffusion 3.5 Medium by correctly interpreting the 'exploded' burger layout and rendering all requested text prominently with the correct effects. While Stable Diffusion 3.5 had nice lighting, it ignored the core compositional instruction of suspended individual components and struggled with clear text integration.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
Imagen 4.0 Generate 001
- + Excellent text legibility and spelling for all menu items and the date.
- + Successfully rendered the specific menu prices and full descriptions requested.
- + Captures the look of consistent chalk handwriting with natural variations.
- − Includes instructional text from the prompt (like 'Tittle', 'Menu', and 'Footer') as actual elements on the board.
- − The layout is a bit cluttered with redundant meta-text describing the handwriting.
Stable Diffusion 3.5 Medium
- + Very realistic chalk texture and dust effects on the blackboard.
- + Artistic chalk illustration style for the titles.
- − Poor text rendering with significant spelling errors and gibberish (e.g., 'Oocclutels', 'helButer').
- − Failed to follow the specific date and price details in the prompt.
- − The layout is confusing with overlapping and truncated words.
Verdict: Imagen 4.0 significantly outperformed Stable Diffusion 3.5 Medium in terms of prompt adherence and text accuracy, correctly spelling the complex menu items and dates. While Imagen 4.0 mistakenly included some of the descriptive instructions as text on the board, Stable Diffusion 3.5 Medium produced mostly illegible text and failed to follow the specific content requirements.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
Imagen 4.0 Generate 001
- + Excellent visual quality with a cinematic feel and sharp textures.
- + Vibrant lighting and creative use of space-themed patterns on the horse's coat.
- + Strong composition with a sense of scale and movement.
- − Failed the negative constraint; the astronaut is riding the horse, not the requested 'horse on top'.
Stable Diffusion 3.5 Medium
- + Natural-looking lighting on the planet surface reflection.
- + Good adherence to the astronaut's gear details.
- − Failed the negative constraint; the astronaut is riding the horse instead of the inverse.
- − Noticeable anatomy errors, including extra horse legs and distorted hooves.
- − Flat composition compared to the other model.
Verdict: Both models failed the specific spatial instruction to place the horse on top of the astronaut. However, Imagen 4.0 is the clear winner due to its superior visual quality, artistic composition, and lack of the severe anatomical glitches present in Stable Diffusion 3.5 Medium, which generated additional legs for the horse.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
Imagen 4.0 Generate 001
- + Excellent adherence to the prompt including the woman looking at her phone.
- + High resolution and clear details on the capybara's fur and clothing.
- + Well-balanced composition that clearly shows the interior and exterior context.
- − The capybara's paws look slightly like sharp claws rather than natural anatomical paws.
- − Includes a taxi sign on the inside roof which is a bit nonsensical.
Stable Diffusion 3.5 Medium
- + Natural-looking capybara fur and facial features.
- + Good cinematic lighting that matches a city at night.
- − Failed to show the woman looking at her phone as requested; she is just looking forward.
- − The capybara's paws are poorly rendered and overlap awkwardly.
- − The perspective is cramped, making it harder to see the taxi interior.
Verdict: Imagen 4.0 followed the complex prompt instructions much more accurately, specifically including the detail about the businesswoman looking at her phone. Stable Diffusion 3.5 Medium failed on that narrative beat and had more significant anatomical issues with the capybara's paws, although it captured a nice mood.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
Imagen 4.0 Generate 001
- + Excellent typography with perfect spelling of all complex phrases.
- + Clean, professional composition with high-quality cinematic lighting.
- + Well-integrated border of thorns and webs as requested.
- − The parchment is represented more as a side element than the main poster background.
- − Layout feels slightly digital/vector rather than a vintage physical poster.
Stable Diffusion 3.5 Medium
- + Stronger 'vintage parchment' feel for the main invitation area.
- + Includes all requested thematic elements likes bats and twisted trees.
- − Numerous spelling errors in the title and body text.
- − Jack-o-lanterns are not 'central' as requested by the prompt.
- − Graphics and lines are rougher and less polished compared to Image A.
Verdict: Imagen 4.0 significantly outperformed Stable Diffusion 3.5 Medium by rendering all text perfectly and following the layout instructions precisely. While Stable Diffusion 3.5 captured a more rustic parchment aesthetic, its failure to spell basic words correctly and the off-center placement of the pumpkins makes it inferior for a specific design task.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
Imagen 4.0 Generate 001
- + Excellent PBR textures and material realism
- + Perfectly captures the miniature 3D diorama aesthetic
- + High visual clarity and balance
- − Completely ignored the text requirements ('JAPAN', 'SUSHI')
- − Missed the light blue background requirement
Stable Diffusion 3.5 Medium
- + Successfully rendered the requested text 'JAPAN' and 'SUSHI'
- + Accurately followed the light blue background and isometric camera angles
- + Clean, high-clarity output
- − Text rendering has minor artifacts on the letter 'S'
- − Lack of detail variety in the sushi types compared to Model A
- − Missed the flag icon requirement
Verdict: Stable Diffusion 3.5 Medium is the winner because it adhered to the complex prompt instructions including text placement, background color, and isometric layout, whereas Imagen 4.0 significantly failed by ignoring all text and background requirements. While Imagen 4.0 produced a more realistic and visually appealing 3D model of sushi, it failed the fundamental task of following the user's compositional instructions.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
Imagen 4.0 Generate 001
- + Successfully includes all four requested animals: puppy, kitten, bunny, and fox kit.
- + Excellent composition showing the animals dynamicly playing and tumbling as requested.
- + High level of detail in the foreground foliage, dew drops, and fur textures.
- − The style leans more toward a digital illustration than the requested 'hyper-photorealistic' look.
- − The puppy has a slight anatomical oddity where its leg blends into the kitten.
Stable Diffusion 3.5 Medium
- + Rich, warm lighting that perfectly captures the 'golden sunrise' and 'god rays' requested.
- + Very cute character designs with expressive eyes.
- + Higher contrast and more vibrant color palette.
- − Failed the prompt by omitting the baby bunny entirely.
- − The kitten lacks tabby markings, appearing as a solid orange/ginger cat.
- − The animals are mostly stationary rather than 'playfully chasing butterflies and tumbling'.
Verdict: Imagen 4.0 is the clear winner for prompt adherence, accurately depicting all four specific baby animals and the requested playful tumbling action. While Stable Diffusion 3.5 Medium produced a warm and visually striking image, it failed to include the bunny and ignored the core 'tumbling' interaction requested in the prompt.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
Imagen 4.0 Generate 001
- + Perfect text rendering for all requested strings including accents.
- + Clean minimalist vector aesthetic that fits the brand identity.
- + Excellent adherence to all prompt elements including color palette and steam.
- − Simple composition might be seen as safe or generic.
Stable Diffusion 3.5 Medium
- + Intricate vintage engraving style with high artistic detail.
- + Strong atmospheric character and good use of classic typography.
- − Frequent spelling errors including 'Florrian' and 'Est 170'.
- − Layout is cluttered and doesn't follow the 'minimalist' requirement.
- − The cloche is poorly defined and resembles a jar or teapot lid.
Verdict: Imagen 4.0 delivers a professional, production-ready logo that perfectly replicates the requested text and minimalist style. In contrast, Stable Diffusion 3.5 Medium produces an overly complex illustration with several core spelling errors and misses the 'minimalist' part of the prompt.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
Imagen 4.0 Generate 001
- + Excellent text rendering with no spelling errors.
- + Clean, professional flat-vector aesthetic that matches the 'modern infographic' prompt.
- + Clear, logical visual flow for the mission steps.
- − Missed the final two steps (Descent and Landing) in the sequence.
- − Scientific accuracy of the diagram is low, confusing 'Translunar' with a lunar orbit.
Stable Diffusion 3.5 Medium
- + Successfully attempted all text labels for the specific steps requested.
- + Creative celestial composition with a unique artistic style.
- − Major text rendering issues with significant 'gibberish' across almost all labels.
- − Poor layout for an infographic where the pathing lines are cluttered and messy.
- − Fails the 'clean, crisp lines' requirement with grainy textures and artifacting.
Verdict: Imagen 4.0 is the clear winner despite missing the final two steps of the sequence, as it produced a professional, legible, and aesthetically pleasing infographic. Stable Diffusion 3.5 Medium struggled significantly with the text rendering and the 'clean' vector style requested, resulting in a cluttered image with nonsensical labels.
Explore each model
Stability AI's 2.5-billion parameter Multimodal Diffusion Transformer with improvements (MMDiT-X) text-to-image model optimized for consumer hardware, featuring improved image quality, typography, and complex prompt understanding