Head to head
Esc

Models · slot A

to navigate to pick

Imagen 3.0 Generate 002 Google Z-Image Turbo Alibaba

Settled by community votes across 6 shared challenges, with an AI judge weighing in on each.

Imagen 3.0 Generate 002

21.0 arena score

#36 of 62 in Text-to-Image

Skill signature · Text-to-Image

Z-Image Turbo

25.3 arena score

#12 of 62 in Text-to-Image

Vote tally

Where the votes landed

Imagen 3.0 Generate 002

0%

win rate

Ties

0%

Z-Image Turbo

0%

win rate

Shared challenges 6

Challenge by challenge

The strongest take from each model on every shared challenge, with the AI judge's read.

Modern Clean Menu

Text-to-Image

“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”

Imagen 3.0 Generate 002
Z-Image Turbo

AI Judge Analysis

Imagen 3.0 Generate 002

  • + Strict adherence to the 4x4 grid layout requested.
  • + Includes pizza as a specific recurring food motif which matches the headers.
  • + Excellent use of bold sans-serif typography that remains legible.
  • The grid is slightly repetitive with mostly pizza images.
  • Smaller body text is heavily garbled at high resolution.

Z-Image Turbo

  • + Strong 'vibrant accents' through the use of orange bars, improving visual hierarchy.
  • + Higher diversity in food photography, showing pasta and meats alongside pizza.
  • + Text rendering for headers like 'APPETIZERS' is very clean.
  • The layout is less integrated, with text blocks breaking the grid consistency.
  • Large 'PIZZA MANS' header is awkward in its placement and phrasing.
  • Includes a typo in the word 'SECTION' (rendered as 'SE TIIION').

Verdict: Imagen 3.0 provides a more cohesive and professional layout that strictly adheres to the 'grid' requirement, creating a balanced and realistic menu structure. While Z-Image Turbo has more vibrant color accents and diverse food photos, its layout feels slightly disjointed and contains more noticeable text Errors. Imagen 3.0 is preferred for its superior composition and cleaner design aesthetic.

Magic Burger Explosion: Fiery Photorealism Challenge

Text-to-Image

“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”

Imagen 3.0 Generate 002
Z-Image Turbo

AI Judge Analysis

Imagen 3.0 Generate 002

  • + Excellent 'exploded' layout with clearly separated floating layers.
  • + High textural detail on the bun and patty.
  • + Professional graphic design layout with the logo at the base.
  • Nonsense gibberish text inside the starburst above the price.
  • The 'fiery' effect on the text is subtle rather than glowing.

Z-Image Turbo

  • + Perfect adherence to text rendering with the requested glowing/fiery effect.
  • + Highly appetizing and photorealistic rendering of the meat and cheese.
  • + Successful starburst and secondary message placement.
  • Failed the 'exploded' instruction; the burger is mostly assembled rather than suspended in mid-air.
  • The '€' symbol and some edges show minor generation artifacts.

Verdict: Imagen 3.0 successfully captured the 'exploded' burger concept with distinct floating layers, but struggled with clean text rendering in the starburst. Z-Image Turbo delivered superior text effects and appetite appeal, but failed to separate the burger components into the requested deconstructed/exploded view. Ultimately, Imagen 3.0 is preferred for more accurately following the core structural prompt of an exploded burger.

Chalkboard Menu

Text-to-Image

“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”

Imagen 3.0 Generate 002
Z-Image Turbo

AI Judge Analysis

Imagen 3.0 Generate 002

  • + Excellent chalk-like texture with smudges and natural variations
  • + Detailed wood frame and realistic board background
  • + Included additional thematic text at the bottom
  • Multiple spelling errors across almost every line (e.g., 'Risottoo', 'Octpsuin', 'Choouip')
  • Repeated entire lines of text with different spellings, creating a cluttered look
  • Did not use elegant cursive for the title as requested

Z-Image Turbo

  • + Highly accurate spelling for almost all items
  • + Clear, legible layout with consistent letter spacing
  • + Successfully completed the 'Brown Butter' item into a logical 'Chocolate Cookies'
  • Minor typo in 'Mustroom' instead of Mushroom
  • Text lacks the 'slight slant' requested in the prompt
  • Lettering feels slightly more like a digital chalk font than authentic variable handwriting

Verdict: Z-Image Turbo is the clear winner because it maintains high legibility and spelling accuracy, whereas Imagen 3.0 suffers from severe hallucinations and repetitive text blocks. While Imagen 3.0 has a more authentic chalk texture on the board, Z-Image Turbo's ability to follow the complex menu list without significant errors makes it much more useful.

The Reversed Rodeo

Text-to-Image

“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”

Imagen 3.0 Generate 002
Z-Image Turbo

AI Judge Analysis

Imagen 3.0 Generate 002

  • + Excellent cinematic lighting with a vibrant nebula background.
  • + High-quality textures on the space suit and horse's mane.
  • + Balanced and dynamic composition.
  • Failed to follow the specific spatial instruction (horse on top of astronaut).

Z-Image Turbo

  • + Detailed representation of a modern astronaut's suit.
  • + Clear, sharp focus on the central subjects.
  • Failed to follow the specific spatial instruction (horse on top of astronaut).
  • The background is quite plain and lacks the cinematic depth of the other model.
  • The horse's back legs have anatomical issues near the hooves.

Verdict: Both Imagen 3.0 and Z-Image Turbo reached a 'tie' in prompt adherence by failing to follow the complex spatial instruction where the horse was requested to be on top of the astronaut. However, Imagen 3.0 is the superior image due to its stunning cinematic lighting, detailed textures, and cohesive background, whereas Z-Image Turbo has anatomical errors and a flat aesthetic.

Isometric Miniature Diorama Scenes

Text-to-Image

“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”

Imagen 3.0 Generate 002
Z-Image Turbo

AI Judge Analysis

Imagen 3.0 Generate 002

  • + Excellent adherence to technical specs like the 45° isometric angle.
  • + Accurate rendering of the Japanese flag.
  • + High-quality 3D clay-like textures and clean typography.
  • Text is placed in the top-left corner instead of the requested top-center.

Z-Image Turbo

  • + Perfectly centered text as requested in the prompt.
  • + Soft, appealing lighting and realistic subsurface scattering on the fish texture.
  • + Strong composition with a clear focal point.
  • Incorrectly featured the flag of China instead of the Japanese flag.
  • The text 'SUSHI' is slightly misaligned with the text 'JAPAN'.

Verdict: Imagen 3.0 Generate 002 is the clear winner for its accuracy, correctly identifying the Japanese flag and providing a more detailed 3D scene. While Z-Image Turbo captures the centered text layout better, its inclusion of the Chinese flag for a Japanese theme is a significant factual error that breaks prompt adherence.

Adorable Baby Animals in Sunny Meadow

Text-to-Image

“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”

Imagen 3.0 Generate 002
Z-Image Turbo

AI Judge Analysis

Imagen 3.0 Generate 002

  • + Includes all four requested animals clearly visible.
  • + Exceptional fur texture and lighting on the animals.
  • + Captures the 'tumbling' aspect of the prompt with the bunny on its back.
  • The fox kit has slightly strange anatomy where its paw meets the grass.
  • The butterflies are quite small relative to the scene.

Z-Image Turbo

  • + Very expressive faces with open mouths conveying a 'joyful' vibe.
  • + Good interpretation of 'chasing butterflies' with the animals looking upwards.
  • + Bright, vibrant color palette.
  • Anatomical issues including the kitten having an extra paw-like growth on its chest and the fox having two tails merging.
  • The kitten's facial structure is slightly distorted.

Verdict: Imagen 3.0 Generate 002 is the superior image because it delivers much higher anatomical accuracy and more photorealistic fur textures. While Z-Image Turbo captures the joyful energy well, it suffers from significant AI artifacts like merged tails and extra limbs that detract from the '8K masterpiece' requirement.

Next steps

Explore each model