Head to head
Esc

Models · slot A

to navigate to pick

Imagen 4.0 Fast Generate 001 Google Qwen Image 2.0 Alibaba

Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.

Imagen 4.0 Fast Generate 001

17.7 arena score

#52 of 62 in Text-to-Image

Skill signature · Text-to-Image

Qwen Image 2.0

21.7 arena score

#34 of 62 in Text-to-Image

Vote tally

Where the votes landed

Imagen 4.0 Fast Generate 001

0%

win rate

Ties

0%

Qwen Image 2.0

0%

win rate

Shared challenges 13

Challenge by challenge

The strongest take from each model on every shared challenge, with the AI judge's read.

Geometric Composition

Text-to-Image

“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”

Imagen 4.0 Fast Generate 001
Qwen Image 2.0

AI Judge Analysis

Imagen 4.0 Fast Generate 001

  • + Excellent photorealism with realistic glass reflections and refractions.
  • + The red book features convincing texture and embossed details.
  • + High quality lighting that accurately simulates window light on the table surface.
  • The glass cube has a mirrored base which wasn't specifically requested.
  • The plant is positioned more to the side than directly behind the cube.

Qwen Image 2.0

  • + Successfully positions the plant directly behind the cube.
  • + Accurately places the window on the left side of the frame.
  • + The composition feels more balanced and central.
  • The blue sphere appears to be floating mid-air inside the cube without support.
  • The cube's geometry is slightly inconsistent, with a vertical line through the middle that doesn't align with a standard cube.

Verdict: Imagen 4.0 Fast Generate 001 produces a much more realistic image with superior textures and lighting, although it adds a mirrored tray to the bottom of the cube. Qwen Image 2.0 follows the spatial instructions well but suffers from physics issues, specifically the floating sphere and awkward glass paneling.

Candid Street Photography

Text-to-Image

“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”

Imagen 4.0 Fast Generate 001
Qwen Image 2.0

AI Judge Analysis

Imagen 4.0 Fast Generate 001

  • + Excellent handling of wet pavement reflections
  • + Follows the 'imperfect framing' prompt with a creative voyeuristic perspective
  • + Strong adherence to the requested cinematic color palette
  • The subject's face is partially cropped out by the frame
  • The bicycle geometry is slightly warped near the handlebars

Qwen Image 2.0

  • + Superb natural skin texture and facial details
  • + Accurately depicts the mechanical action of repairing the bicycle chain
  • + Consistent lighting and realistic wet street atmosphere
  • The motion blur on the passing car is very subtle compared to the request
  • Composition is a bit standard rather than 'imperfect' as requested

Verdict: Both models followed the prompt well, but they excelled in different areas. Imagen 4.0 Fast Generate captured the 'imperfect framing' and cinematic mood more effectively, while Qwen Image 2.0 provided significantly better anatomical and skin detail on the man's face and hands. Qwen Image 2.0 is the overall winner for its superior technical execution and realistic human textures.

Fantasy Warrior

Text-to-Image

“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”

Imagen 4.0 Fast Generate 001
Qwen Image 2.0

AI Judge Analysis

Imagen 4.0 Fast Generate 001

  • + Features a realistic, high-quality image of an older man.
  • Completely failed to follow the prompt's subject and setting instructions.
  • Displays a man in a modern leather jacket and jeans in a garden instead of a paladin in armor.
  • Lacks all requested technical elements like bokeh sparks and engraved plate.

Qwen Image 2.0

  • + Excellent adherence to all prompt details including ornate engraved armor, braided hair with beads, and scars.
  • + Strong cinematic lighting with warm highlights and bokeh sparks.
  • + High texture detail on the skin, fabric, and metal surfaces.
  • Slight anatomical distortion in the way the hand rests on the sword hilt.

Verdict: Imagen 4.0 Fast Generate 001 completely failed the prompt, producing a modern man in a garden rather than a fantasy knight. Qwen Image 2.0 followed every specific detail of the prompt, creating a rich, character-driven portrait with excellent lighting and texture.

Modern Clean Menu

Text-to-Image

“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”

Imagen 4.0 Fast Generate 001
Qwen Image 2.0

AI Judge Analysis

Imagen 4.0 Fast Generate 001

  • + Excellent utilization of the requested sections (Appetizers, Pizza, Mains) in a realistic layout.
  • + High-quality, consistent food photography that fits the professional menu context.
  • + Great use of vibrant color accents and bold sans-serif fonts to differentiate sections.
  • Several spelling errors in headings, such as 'Apetiers'.
  • Text within the menu items is mostly placeholder gibberish.

Qwen Image 2.0

  • + Clean, modern grid layout that feels very minimalist and visual.
  • + Varied food photography that matches the category headers above it.
  • + Good use of rounded corners and consistent typography for a casual dining feel.
  • The 'Mains' column shows pizzas in two of the photos, which is inconsistent with the header.
  • The text at the bottom and the pricing labels have significant rendering artifacts and garbled characters.

Verdict: Imagen 4.0 Fast Generate 001 provides a much more convincing and professional menu layout that feels like a real graphic design piece, despite some typos in the headings. Qwen Image 2.0 has a nice minimalist aesthetic, but the logic of the food images under the headers is flawed and the text rendering is quite messy at the bottom.

Magic Burger Explosion: Fiery Photorealism Challenge

Text-to-Image

“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”

Imagen 4.0 Fast Generate 001
Qwen Image 2.0

AI Judge Analysis

Imagen 4.0 Fast Generate 001

  • + Excellent typography with a consistent glowing neon-ember effect.
  • + Very clean 'exploded' view where all components are clearly separated as requested.
  • + Photorealistic texture on the bun and patties.
  • The 'starburst' for the price is a bit simplistic compared to the rest of the professional ad layout.
  • The bottom half of the burger is less 'exploded' than the top half.

Qwen Image 2.0

  • + Dynamic use of fire and smoke to enhance the 'fiery' theme from the prompt.
  • + Rich textures and lighting on the cheese and dripping sauce.
  • + Strong sense of motion with flying crumbs and embers.
  • The components are not well 'exploded'; the patties, cheese, and tomatoes are largely touching or merged together.
  • The text 'LIMITED TIME ONLY' is small and lacks the fiery effect seen in the main title.

Verdict: Imagen 4.0 Fast Generate 001 provides a much better 'exploded' view, which was a core instruction of the prompt, and maintains consistent text effects across all elements. While Qwen Image 2.0 has more dramatic fire effects and lighting, it fails to separate the burger components effectively and misses the styling cues for the secondary text.

Chalkboard Menu

Text-to-Image

“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”

Imagen 4.0 Fast Generate 001
Qwen Image 2.0

AI Judge Analysis

Imagen 4.0 Fast Generate 001

  • + Very clean and legible text rendering.
  • + Accurately follows the specific pricing and menu items mentioned in the prompt.
  • The font looks too digital and uniform, failing the 'no digital fonts' requirement.
  • Spelling errors present like 'Octuphus' and 'Cookes'.
  • Composition is flat and lacks the 'cozy café' atmosphere requested.

Qwen Image 2.0

  • + Authentic chalk texture with realistic smudges and varying pressure.
  • + Excellent 'cozy café' background and lighting that adds to the atmosphere.
  • + Better adherence to the cursive title and handwritten style request.
  • The text layout is slightly cluttered due to line breaks.
  • Small spelling artifact in the title character rendering.

Verdict: Qwen Image 2.0 is the clear winner as it successfully captured the requested 'cozy café' atmosphere and provided an authentic chalk texture that looks truly handwritten. Imagen 4.0 Fast Generate 001 produced text that looked like a digital font and failed to provide any environmental context, appearing as a flat graphic rather than a photo.

The Reversed Rodeo

Text-to-Image

“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”

Imagen 4.0 Fast Generate 001
Qwen Image 2.0

AI Judge Analysis

Imagen 4.0 Fast Generate 001

  • + Excellent anatomical rendering of the horse.
  • + High cinematic quality with dramatic lighting and detailed nebula background.
  • + Clear facial details through the astronaut helmet.
  • Fails the specific logic constraint of the prompt (horse on top of astronaut).
  • Standard interpretation of an astronaut riding a horse.

Qwen Image 2.0

  • + Dynamic composition with the inclusion of Earth and floating water droplets.
  • + Good level of detail on the space suit and horse's mane.
  • + Vibrant and clear contrast.
  • Fails the specific logic constraint of the prompt (horse on top of astronaut).
  • Anatomical issues with the horse's legs and hooves.

Verdict: Both models failed the negative constraint/spatial reasoning test to place the horse on top of the astronaut, instead providing the typical image of an astronaut riding a horse. Imagen 4.0 Fast Generate 001 is the winner due to superior anatomical correctness and a more polished, cinematic aesthetic compared to the distorted limbs and generic styling in Qwen Image 2.0.

The Capybara Taxi Driver

Text-to-Image

“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”

Imagen 4.0 Fast Generate 001
Qwen Image 2.0

AI Judge Analysis

Imagen 4.0 Fast Generate 001

  • + Excellent photorealistic lighting and skin textures for both the capybara and the human.
  • + The human appears to be in the passenger seat rather than the back seat, allowing for a better view of her expression.
  • + High level of detail on the capybara's fur and the taxi driver cap.
  • The woman is in the front passenger seat instead of the back seat as requested.
  • The capybara's hands/paws look like long, thin human fingers which is slightly uncanny.

Qwen Image 2.0

  • + Successfully placed the woman in the back seat as requested.
  • + The capybara's paws look more anatomically correct for the species.
  • + Captures the gritty, over-saturated look of a New York street scene at night effectively.
  • Lower overall image resolution and clarity compared to the other model.
  • The background lighting is a bit messy and distracting.
  • The textures on the car interior and clothing are less defined.

Verdict: Qwen Image 2.0 followed the spatial instructions more accurately by placing the passenger in the back seat, though it suffered from lower fidelity. Imagen 4.0 produced a much more realistic and detailed image with superior lighting, but failed the prompt by putting the woman in the front seat next to the driver.

The Halloween Invitation

Text-to-Image

“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”

Imagen 4.0 Fast Generate 001
Qwen Image 2.0

AI Judge Analysis

Imagen 4.0 Fast Generate 001

  • + Clean layout with clear separation of elements
  • + Cinematic lighting on the central pumpkin
  • + High contrast and vibrant colors
  • Spelling error in the title: 'IINVIITATION'
  • Typo in the scroll text: 'NIGIT OF FRIGITS'
  • Poorly formatted event details at the bottom

Qwen Image 2.0

  • + Perfect text rendering for all requested strings
  • + Superior gothic aesthetic with thorns and webs detail
  • + Excellent source-appropriate font choices
  • Slightly less 'cinematic' lighting compared to Model A
  • Composition feels a bit more cluttered

Verdict: Qwen Image 2.0 followed every instruction perfectly, including complex text rendering with zero spelling errors. While Imagen 4.0 Fast Generate 001 featured more vibrant lighting, it failed significantly on text accuracy with multiple typos and awkward spacing in the invitations details.

Isometric Miniature Diorama Scenes

Text-to-Image

“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”

Imagen 4.0 Fast Generate 001
Qwen Image 2.0

AI Judge Analysis

Imagen 4.0 Fast Generate 001

  • + Perfectly executes the isometric 3D cartoon style with soft, refined textures.
  • + Text layout is elegantly integrated into the composition.
  • + Excellent adherence to the 'miniature diorama' aesthetic with clean geometric shapes.
  • The sushi pieces are slightly more abstract/stylized than 'realistic PBR' might suggest.
  • The flag icon is placed to the right instead of following a strictly top-center vertical stack.

Qwen Image 2.0

  • + High realism in the food textures and materials.
  • + Follows the text hierarchy instructions literally with large bold 'JAPAN' and 'SUSHI' below.
  • + Good use of a wooden diorama base.
  • Fails the 'isometric' perspective request, using a more standard perspective.
  • Lacks the '3D cartoon' and 'miniature' aesthetic, looking more like a standard food photograph.
  • The text and flag overlay look like a flat 2D graphic rather than integrated into the scene's space.

Verdict: Imagen 4.0 Fast Generate 001 succeeded in capturing the specific isometric 3D cartoon style and miniature aesthetic requested, resulting in a cohesive and professional-looking graphic. Qwen Image 2.0 produced more realistic food but failed on the perspective and stylistic requirements, making it look like a composite photo rather than a designed diorama.

Adorable Baby Animals in Sunny Meadow

Text-to-Image

“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”

Imagen 4.0 Fast Generate 001
Qwen Image 2.0

AI Judge Analysis

Imagen 4.0 Fast Generate 001

  • + Excellent fur texture rendering and soft lighting.
  • + Coherent and unified composition with clear subject focus.
  • Failed to include the requested butterflies.
  • The animals are sitting still rather than 'playfully chasing' as prompted.
  • Incorrect animal breeds: generated a brown/white spaniel-like dog and a black kitten instead of a Golden Retriever and a tabby.

Qwen Image 2.0

  • + Successfully followed all prompt instructions, including the specific breeds and chasing butterflies.
  • + Highly dynamic 'tumbling' action that perfectly captures the playful vibe.
  • + Beautiful light rays and dew sparkling on the wildflowers.
  • Minor anatomical distortion on the kittens paws/limbs during the action.
  • The fox's face is slightly obscured by the tumbling pose.

Verdict: Qwen Image 2.0 is the clear winner as it fulfilled every part of the complex prompt, including specific animal breeds and the action of chasing butterflies. While Imagen 4.0 produced a high-quality static portrait, it missed the 'tabby' and 'golden retriever' requirements and failed to depict the requested activity.

Vintage Cafe Logo

Text-to-Image

“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”

Imagen 4.0 Fast Generate 001
Qwen Image 2.0

AI Judge Analysis

Imagen 4.0 Fast Generate 001

  • + Excellent typography rendering with the correct accent on 'Caffè'
  • + Superior minimalist composition perfect for a vector logo
  • + Accurate representation of the requested banner and cloche elements
  • Includes small, nonsensical placeholder text above the main title
  • The steam effect is a bit faint and simplistic

Qwen Image 2.0

  • + Strong 'vintage' illustrative style with nice shading
  • + Good text rendering for both the name and the date
  • + Dynamic interpretation of the steam and cloche
  • The composition feels cramped as the dome is very large and overlaps the text area
  • Less 'minimalist' than requested, leaning more into a detailed illustration

Verdict: Imagen 4.0 Fast Generate 001 provides a much better execution of a 'minimalist vector logo' with a clean, balanced layout that is typical of professional branding. While Qwen Image 2.0 has an attractive illustrative style, it fails the minimalist constraint and the composition feels a bit crowded.

Apollo 11: Journey to Tranquility

Text-to-Image

“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”

Imagen 4.0 Fast Generate 001
Qwen Image 2.0

AI Judge Analysis

Imagen 4.0 Fast Generate 001

  • + Clean vector aesthetic with excellent use of the requested color palette.
  • + Good iconography for the Lunar Module and crew silhouettes.
  • Numerous spelling errors including 'APOLO', 'MOOR', and 'MOO + ON'.
  • The flow of the infographic is confusing and does not logically follow the numbered steps.

Qwen Image 2.0

  • + Perfect text rendering for all mission steps and astronaut names.
  • + Logical vertical layout that clearly illustrates the progression from launch to landing.
  • + Strong adherence to the specific icon requests for each step.
  • Spelling error in 'Translunjar' (contains an extra 'j').
  • Slightly less 'flat' than Model A, with some heavier glow effects around the text.

Verdict: Qwen Image 2.0 is the clear winner for its superior instructional flow and high-quality text rendering, despite one minor typo. While Imagen 4.0 Fast Generate 001 captures a more professional graphic design style, its severe spelling issues and nonsensical layout make it fail as an infographic.

Next steps

Explore each model