Head to head
Esc

Models · slot A

to navigate to pick

GPT Image 2 OpenAI Imagen 4.0 Fast Generate 001 Google

Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.

GPT Image 2

28.1 arena score

#3 of 62 in Text-to-Image

Top 3 in Text-to-Image
Skill signature · Text-to-Image

Imagen 4.0 Fast Generate 001

17.6 arena score

#53 of 62 in Text-to-Image

Vote tally

Where the votes landed

GPT Image 2

0%

win rate

Ties

0%

Imagen 4.0 Fast Generate 001

0%

win rate

Shared challenges 13

Challenge by challenge

The strongest take from each model on every shared challenge, with the AI judge's read.

Geometric Composition

Text-to-Image

“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”

GPT Image 2
Imagen 4.0 Fast Generate 001

AI Judge Analysis

GPT Image 2

  • + Excellent photographic realism and texture on the book and table.
  • + Accurate glass thickness and internal reflections.
  • + Lighting feels natural and integrates well with the outdoor window background.
  • The plant is more on the side than directly behind the cube as requested.

Imagen 4.0 Fast Generate 001

  • + Perfect adherence to the spatial 'behind the cube' instruction for the plant.
  • + Translucent blue material for the sphere adds visual interest.
  • + Clear lighting direction from the left as specified.
  • The bottom of the cube appears to be a mirror rather than clear glass.
  • The glass walls are unnaturally thin and lack realistic refraction.

Verdict: GPT Image 2 is the superior image due to its high level of photorealism and accurate material properties, such as the thickness of the glass and the wood grain texture. While Imagen 4.0 followed the spatial instruction for the plant slightly better, its use of a mirrored base inside the cube and less convincing glass rendering makes it feel more like a 3D render than a photo.

Candid Street Photography

Text-to-Image

“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”

GPT Image 2
Imagen 4.0 Fast Generate 001

AI Judge Analysis

GPT Image 2

  • + Excellent skin texture and hyper-realistic facial details
  • + Accurately represents the 50mm lens and shallow depth of field request
  • + Includes the motion blur on passing cars as specifically requested
  • The 'imperfect framing' is missing as it feels quite centered and deliberate
  • Visible AI artifact with the bicycle rack clipping into the man's hands

Imagen 4.0 Fast Generate 001

  • + Strong creative interpretation of 'imperfect framing' using a frame-within-a-frame composition
  • + Very realistic reflections on the wet pavement
  • + Good sense of light rain with visible raindrops hitting the ground
  • The man's Face and head are partially cut off at the top
  • Less detail in natural skin texture compared to Model A
  • The 'motion blur' on the background car is less pronounced than requested

Verdict: GPT Image 2 provides a more standard high-quality photography look with superior detail on the subject's face and excellent adherence to the 'motion blur' requirement. Imagen 4.0 Fast Generate 001 takes a more artistic approach to the 'imperfect framing' prompt but suffers from poor subject placement and slightly lower detail in the human features.

Fantasy Warrior

Text-to-Image

“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”

GPT Image 2
Imagen 4.0 Fast Generate 001

AI Judge Analysis

GPT Image 2

  • + Excellent adherence to all prompt details including braided hair with beads, engraved armor, and torchlight.
  • + High-quality skin texture showing realistic dirt and faint scars.
  • + Great composition with effective use of shallow depth of field and bokeh sparks.
  • None identified based on the provided prompt.

Imagen 4.0 Fast Generate 001

  • + Realistic texture on the leather jacket.
  • Complete failure to follow the prompt regarding subject, setting, and equipment.
  • Generated a modern man in a garden instead of a paladin in armor.
  • Lack of specific details like braids, beads, torchlight, or plate armor.

Verdict: GPT Image 2 perfectly captured every detail of the complex prompt, from the specific textures of the battle-worn armor to the lighting and character styling. In contrast, Imagen 4.0 Fast Generate 001 failed entirely, producing a modern portrait of an older man in a garden that bears no resemblance to the requested paladin character. GPT Image 2 is the clear winner for both technical execution and creative adherence.

Modern Clean Menu

Text-to-Image

“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”

GPT Image 2
Imagen 4.0 Fast Generate 001

AI Judge Analysis

GPT Image 2

  • + Exceptional text rendering with perfect spelling and logical price points.
  • + Highly professional layout with sophisticated graphic design elements like icons and accent lines.
  • + High-quality, distinct food photography that accurately represents the described dishes.
  • The layout is slightly crowded compared to the minimalist request, though it remains clean.

Imagen 4.0 Fast Generate 001

  • + Strong adherence to the minimalist, white-space focused layout.
  • + Logical grid structure that is easy to follow at a glance.
  • + Good use of vibrant color blocks for section accents.
  • Text is largely illegible gibberish with numerous spelling errors.
  • The food photos lack variety, showing pizzas in sections that are not for pizza.
  • Prices and layout alignment are inconsistent and messy.

Verdict: GPT Image 2 is far superior because it generates a fully functional, professional menu with perfect English text and appropriate high-quality food imagery for each section. Imagen 4.0 Fast Generate 001 provides a good structural layout but fails significantly on text legibility and image relevance, placing pizzas in the 'Apetiers' and 'Main Courses' sections.

Magic Burger Explosion: Fiery Photorealism Challenge

Text-to-Image

“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”

GPT Image 2
Imagen 4.0 Fast Generate 001

AI Judge Analysis

GPT Image 2

  • + Excellent photorealistic texture on the meat patty and fresh vegetables.
  • + Dynamic and cohesive fiery aesthetic applied across all text and the background.
  • + Perfect layout that feels like a professional, high-energy advertisement.
  • The text takes up a very large portion of the frame, slightly crowding the burger.

Imagen 4.0 Fast Generate 001

  • + Clean, symmetrical explosion of ingredients that is easy to read.
  • + Accurate rendering of the starburst and price tag.
  • + The burger components are very distinct from one another.
  • Text error in 'LIMITED TIME ON ONLY' including an extra word.
  • The burger textures appear somewhat plastic or artificial compared to the other model.
  • The background is relatively flat and lacks the requested 'fiery' intensity.

Verdict: GPT Image 2 (Model A) is the clear winner as it captures the 'fiery' and 'dynamic' spirit of the prompt with much higher photorealism and professional composition. While Model B followed the layout instructions well, it suffered from a typo in the secondary text and a lack of visual intensity in the background.

Chalkboard Menu

Text-to-Image

“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”

GPT Image 2
Imagen 4.0 Fast Generate 001

AI Judge Analysis

GPT Image 2

  • + Excellent text rendering with no spelling errors.
  • + The chalk texture is highly realistic with authentic dusty strokes.
  • + The background environment and lighting create a convincing 'cozy café' atmosphere.
  • The slant of the text is slightly inconsistent between lines.

Imagen 4.0 Fast Generate 001

  • + Successfully included all requested text and menu items.
  • + Clear, legible layout with centered alignment.
  • Contains spelling errors such as 'Octuphus' and 'Cookes'.
  • The text looks more like a digital font or marker than realistic grainy chalk.
  • The composition is repetitive, including the 'Brown Butter' item twice.

Verdict: GPT Image 2 is the superior model as it perfectly followed all text requirements without any spelling errors and captured a much more realistic chalk texture. In contrast, Imagen 4.0 Fast committed multiple spelling mistakes and rendered text that looked too clean and digital, failing the 'handwritten-style' and 'chalk texture' prompt requirements.

The Reversed Rodeo

Text-to-Image

“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”

GPT Image 2
Imagen 4.0 Fast Generate 001

AI Judge Analysis

GPT Image 2

  • + Excellent adherence to the specific 'horse on top' spatial instruction
  • + High level of texture detail on the spacesuit and lunar surface
  • The horse's Anatomy and posture are quite stiff and anatomically awkward
  • The lighting on the horse feels a bit flat compared to the background

Imagen 4.0 Fast Generate 001

  • + Dynamic composition with a classic cinematic aesthetic
  • + Good lighting and color palette in the nebula background
  • Completely failed the negative constraint to put the horse on top
  • Anatomy issues with the horse's front legs and back hoof orientation

Verdict: GPT Image 2 is the clear winner because it followed the difficult and specific spatial instruction for the horse to be riding the astronaut. Imagen 4.0 Fast Generate 001 produced a generic 'astronaut riding a horse' image, failing the primary challenge of the prompt.

The Capybara Taxi Driver

Text-to-Image

“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”

GPT Image 2
Imagen 4.0 Fast Generate 001

AI Judge Analysis

GPT Image 2

  • + Excellent photorealism in the capybara's fur and textures.
  • + Cinematic lighting that perfectly captures the New York taxi night atmosphere.
  • + The businesswoman has a naturally bored, blurred expression that perfectly matches the 'normalcy' of the prompt.
  • The passenger is holding the phone with three hands/multiple fingers in a distorted way.
  • The capybara's right paw is merged oddly with the steering wheel.

Imagen 4.0 Fast Generate 001

  • + Successfully shows two paws clearly on the steering wheel.
  • + The capybara's outfit is sharp and includes a professional-looking tie.
  • + The composition is clear and well-lit with all main elements visible.
  • The capybara's paws look more like long-fingered primate hands than capybara paws.
  • The lighting on the capybara's head is a bit inconsistent with the dark taxi interior.
  • The businesswoman looks more like a studio model than a passenger in a moving taxi.

Verdict: GPT Image 2 (Model A) wins on photorealism and atmosphere, capturing the gritty, cinematic feel of a New York night taxi ride. While Imagen 4.0 Fast Generate 001 (Model B) followed the paw requirement more literally, it suffered from unrealistic hand-like paws and a less convincing integration of the characters into the environment.

The Halloween Invitation

Text-to-Image

“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”

GPT Image 2
Imagen 4.0 Fast Generate 001

AI Judge Analysis

GPT Image 2

  • + Excellent typography with perfect spelling in all three text areas.
  • + Rich, detailed gothic border featuring thorns, webs, and skulls which matches the vintage theme.
  • + Sophisticated lighting and atmospheric background including the requested NYC Arches bridge and skyline.
  • The parchment texture is very dark, making some of the thorny border details slightly muddy.

Imagen 4.0 Fast Generate 001

  • + Strong composition with a clear 'parchment' aesthetic.
  • + Included the requested twisted trees and glowing jack-o-lantern clearly.
  • Significant text errors including 'IINVIIATION' and 'NIGIT OF FNIGHTS'.
  • The art style is more of a modern digital illustration rather than the requested 'vintage gothic' or 'cinematic' look.
  • The text layout at the bottom is cluttered with the location and time merged into one line.

Verdict: GPT Image 2 followed the prompt with much higher fidelity, producing a professional-grade invitation with flawless text and a complex, atmospheric gothic style. In contrast, Imagen 4.0 Fast Generate 001 struggled with the spelling of the requested phrases and opted for a simpler, cartoonish illustration style that lacked the requested vintage cinematic quality.

Isometric Miniature Diorama Scenes

Text-to-Image

“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”

GPT Image 2
Imagen 4.0 Fast Generate 001

AI Judge Analysis

GPT Image 2

  • + Excellent PBR textures and lighting that create a premium 3D look.
  • + High attention to detail in the sushi variety and miniature garden elements.
  • + Perfect execution of the text layout and flag icon.
  • Slightly more complex than 'minimal' garnish requested.
  • Chopsticks are slightly thicker than natural scale.

Imagen 4.0 Fast Generate 001

  • + Strictly adheres to the 'minimal' garnish and diorama base instruction.
  • + Clean isometric 45-degree angle.
  • + Accurate text rendering.
  • Low level of detail compared to GPT Image 2.
  • Materials look more like plastic/clay than refined food textures.
  • The composition feels a bit empty for a main subject.

Verdict: GPT Image 2 is the superior image due to its exceptional level of detail, realistic PBR material rendering, and vibrant colors that make the dish look appealing while maintaining the isometric miniature style. Imagen 4.0 Fast Generate 001 is technically accurate to the 'minimal' part of the prompt but lacks the visual polish and 'refined textures' present in the first model.

Adorable Baby Animals in Sunny Meadow

Text-to-Image

“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”

GPT Image 2
Imagen 4.0 Fast Generate 001

AI Judge Analysis

GPT Image 2

  • + Excellent adherence to the 'chasing butterflies' and 'tumbling together' action prompts.
  • + Perfectly captures the specific animal types including a golden retriever and a tabby kitten.
  • + Beautiful lighting effects including god rays and dew sparkles that enhance the magical atmosphere.
  • The fox has a slightly surreal, oversized eye sparkle that looks a bit digital.
  • The kitten's paw reaching out has a slightly soft focus compared to the rest of its body.

Imagen 4.0 Fast Generate 001

  • + Very clean, high-resolution rendering with a classic portrait feel.
  • + Beautifully soft, realistic fur texture on all animals.
  • + Pleasant bokeh effect in the background flowers.
  • Missed the 'tabby' kitten (it is black), 'golden retriever' (looks like a spaniel mix), and completely missed the butterflies.
  • The animals are sitting still rather than 'playfully chasing' or 'tumbling' as requested.
  • The composition is a static lineup rather than a dynamic scene.

Verdict: GPT Image 2 is much more successful at capturing the dynamic elements of the prompt, including all four specific animals and their playful interactions with butterflies. While Imagen 4.0 has high visual quality, it fails on several key prompt instructions, resulting in a static portrait that misses the butterflies and incorrect animal breeds.

Vintage Cafe Logo

Text-to-Image

“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”

GPT Image 2
Imagen 4.0 Fast Generate 001

AI Judge Analysis

GPT Image 2

  • + Excellent typography with correct accents and elegant spacing.
  • + Intricate vintage detailing and texture that fits the Est. 1720 era.
  • + Cohesive composition with a beautiful border and banner.
  • The design is more ornate than 'minimalist' as requested in the prompt.

Imagen 4.0 Fast Generate 001

  • + Successfully captures a more minimalist vector style.
  • + Accurate text rendering for the main logo name and date.
  • Nonsensical text 'AFFD CARO' included above the main name.
  • Composition feels slightly disjointed with thin lines compared to the heavy cloche.
  • The background is cold blue-grey rather than the requested warm cream tones.

Verdict: GPT Image 2 perfectly captures the 'Vintage' and 'Classic Typography' aspects of the prompt with high-quality artistic execution, even if it leans more ornate than minimalist. Imagen 4.0 Fast Generate 001 provides a cleaner minimalist look but fails on color palette requirements and includes hallucinated text characters. GPT Image 2 is the superior choice for its professional branding aesthetic and superior detail.

Apollo 11: Journey to Tranquility

Text-to-Image

“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”

GPT Image 2
Imagen 4.0 Fast Generate 001

AI Judge Analysis

GPT Image 2

  • + Perfectly follows all 6 requested steps in a logical sequence.
  • + Excellent typography and icon accuracy, including the NASA logo and mission patch.
  • + High-quality vector aesthetic with a professional layout and coherent details.
  • The 'Saturn V' illustration includes some unnecessary smoke effects that detract slightly from a purely flat vector style.

Imagen 4.0 Fast Generate 001

  • + Adheres well to the requested flat-vector minimalism.
  • + Uses the requested color palette accurately.
  • Multiple spelling errors including 'APOLO', 'MOOR', and 'VICON' (taking the prompt too literally).
  • The sequence of the infographic is confusing and non-linear.
  • Fails to include all 6 distinct steps as separate, navigable icons.

Verdict: GPT Image 2 (Model A) is the clear winner as it provides a professional-grade, accurate infographic that follows the chronological steps requested in the prompt. In contrast, Imagen 4.0 (Model B) suffers from poor layout logic, significant spelling errors, and a failure to clearly represent the mission progression.

Next steps

Explore each model