Head to head
Esc

Models · slot A

to navigate to pick

GPT Image 1.5 OpenAI Imagen 4.0 Generate 001 Google

Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.

GPT Image 1.5

27.1 arena score

#7 of 62 in Text-to-Image

Top 3 in Image Editing
Skill signature · Text-to-Image

Imagen 4.0 Generate 001

17.0 arena score

#55 of 62 in Text-to-Image

Vote tally

Where the votes landed

GPT Image 1.5

100.0%

win rate

Ties

0.0%

Imagen 4.0 Generate 001

0.0%

win rate

100.0% 0.0% ties 0.0%
Shared challenges 13

Challenge by challenge

The strongest take from each model on every shared challenge, with the AI judge's read.

Geometric Composition

Text-to-Image

“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”

GPT Image 1.5
Imagen 4.0 Generate 001

AI Judge Analysis

GPT Image 1.5

  • + Excellent adherence to lighting instructions with a soft, natural window glow.
  • + Realistic physical interaction between the sphere resting on the base and the book on top.
  • + High-quality rendering of the glass texture and transparency.
  • The cube looks more like an open glass box or container than a solid cube.

Imagen 4.0 Generate 001

  • + Perfect geometric representation of a solid glass cube.
  • + Clean, modern aesthetic with vibrant colors.
  • + Great texture detail on the book cover.
  • The blue sphere is levitating, which defies natural physics without explanation.
  • The plant is almost entirely obscured and not clearly visible through the glass as requested.
  • The reflection on the right side of the cube is confusing and lacks the transparency of the front face.

Verdict: GPT Image 1.5 is the winner because it correctly interprets the spatial relationship between the objects, placing the plant clearly behind and visible through the glass, while maintaining realistic physics. Imagen 4.0 Generate 001 produces a cleaner cube, but the sphere levitates unnaturally and the transparency of the glass is visually inconsistent.

Candid Street Photography

Text-to-Image

“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”

GPT Image 1.5
Imagen 4.0 Generate 001

AI Judge Analysis

GPT Image 1.5

  • + Excellent depiction of light rain with realistic water droplets on the jacket and bicycle.
  • + The composition feels authentic and 'candid' with appropriate focal depth for a 50mm lens.
  • + High level of mechanical detail on the bicycle parts and the man's tools.
  • The car on the left has slightly warped geometry in the rear wheel.
  • The man's hands while working are a bit muddled in the shadows.

Imagen 4.0 Generate 001

  • + Captures the 'Tokyo' street atmosphere well with the iconic yellow taxi in the background.
  • + Strong facial texture and character expression that feels natural.
  • + Clear implementation of the shallow depth of field and bokeh.
  • The bicycle frame is physically impossible, with a frame bar appearing to pass directly through the front tire.
  • The tool in the man's hand is floating and not properly gripped.
  • The man seems relatively dry despite the visible rain in the environment.

Verdict: GPT Image 1.5 produced a much more coherent and realistic image, particularly in how the rain interacts with the subjects and the logical assembly of the bicycle. While Imagen 4.0 captures a nice cinematic atmosphere, it suffers from significant structural errors in the bicycle and tool interaction.

Fantasy Warrior

Text-to-Image

“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”

GPT Image 1.5
Imagen 4.0 Generate 001

AI Judge Analysis

GPT Image 1.5

  • + Excellent photorealism with lifelike eyes and skin texture.
  • + Authentic battle-worn appearance with convincing dirt and scratches.
  • + Highly detailed rendering of both the armor and the fabric underlayer.
  • The bokeh sparks are a bit large and slightly distracting.

Imagen 4.0 Generate 001

  • + Ornate engraving on the armor is very sharp and intricate.
  • + Unique interpretation of hair braids with many metallic beads.
  • + Good use of warm torchlight on one side of the face.
  • The skin texture and hair look more like a digital painting or 3D render than a lifelike photo.
  • The fire effect on the left is somewhat stylized and lacks realism.

Verdict: GPT Image 1.5 produces a much more realistic and emotive portrait that perfectly captures the 'battle-worn' aesthetic with grit and photographic texture. While Imagen 4.0 Generate 001 features very clean and intricate armor engravings, its overall rendering feels more like a video game character and lacks the 'lifelike eyes' and 'dirt' realism requested by the prompt.

Modern Clean Menu

Text-to-Image

“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”

GPT Image 1.5
Imagen 4.0 Generate 001

AI Judge Analysis

GPT Image 1.5

  • + Excellent text rendering with no spelling errors
  • + Logical alignment between food items and price points
  • + Professional layout that feels like a real usable menu
  • The food photos are slightly less high-end in presentation compared to Image B
  • Layout is more traditional than minimalist

Imagen 4.0 Generate 001

  • + Modern and artistic grid layout with vibrant accents
  • + High-quality, aesthetically pleasing food photography
  • + Strong adherence to the 'minimalist' and 'grid' descriptors
  • Text is nonsensical and includes placeholder gibberish
  • Missing several items listed for specific sections like Appetizers and Pizza
  • The grid structure makes it difficult to read as a functional menu

Verdict: GPT Image 1.5 produced a fully functional, professional menu with perfectly legible text and logical sections that perfectly match the prompt requirements. While Imagen 4.0 Generate 001 provides a more visually striking and artistic layout, its failure to generate readable text or complete the requested item sections makes it less useful as a design prototype.

Magic Burger Explosion: Fiery Photorealism Challenge

Text-to-Image

“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”

GPT Image 1.5
Imagen 4.0 Generate 001

AI Judge Analysis

GPT Image 1.5

  • + Excellent photorealistic texture on the meat patty and toasted bun.
  • + Vibrant and high-energy background with realistic glowing embers and sparks.
  • + Strong adherence to the 'fiery' text effect requested in the prompt.
  • The text is slightly crowded at the top and bottom edges.
  • The €6.99 price tag has a slight visual artifact on the decimal point.

Imagen 4.0 Generate 001

  • + Clean, professional layout with good legibility for the main title.
  • + Precise rendering of the price starburst and currency symbol.
  • + Good suspension and spacing of the individual ingredients.
  • The lighting on the ingredients is a bit flat and plastic-looking compared to Model A.
  • The background is less 'fiery' and dynamic, feeling more like a static studio shot with a few sparks.
  • The price text uses a comma instead of a decimal point.

Verdict: GPT Image 1.5 is the winner because it better captures the 'fiery' and 'dynamic' mood requested in the prompt, with superior photorealistic textures on the food itself. While Imagen 4.0 Generate 001 offers a very clean layout, the food looks less appetizing and the overall atmosphere is less impactful than the high-intensity glow and smoke shown in GPT Image 1.5.

Chalkboard Menu

Text-to-Image

“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”

GPT Image 1.5
Imagen 4.0 Generate 001

AI Judge Analysis

GPT Image 1.5

  • + Excellent text accuracy with no spelling errors
  • + Highly realistic chalk texture with dusty smudges and authentic strokes
  • + Accurately completed the partial prompt for 'Brown Butter Chocolate Chip Cookies'
  • The 'cursive' requirement for the title is subtle rather than elegant

Imagen 4.0 Generate 001

  • + Natural wood frame adds to the cafe aesthetic
  • + Legible handwriting style
  • Included meta-text and labels like 'Tittle', 'Menu', and 'Footer' within the image
  • Multiple spelling errors such as 'Berbs' instead of 'Herbs' and 'veriations'
  • Text looks like a digital font or vector stroke rather than realistic chalk

Verdict: GPT Image 1.5 followed the prompt instructions perfectly, including inferring the end of the truncated product name correctly and rendering the text with a very convincing chalk texture. In contrast, Imagen 4.0 Generate 001 failed by including the prompt's structural labels (like 'Footer') as literal text on the board and suffered from several spelling mistakes.

The Reversed Rodeo

Text-to-Image

“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”

GPT Image 1.5
Imagen 4.0 Generate 001

AI Judge Analysis

GPT Image 1.5

  • + Excellent level of cinematic detail in the lunar surface and equipment.
  • + Highly realistic textures on the space suit and horse hair.
  • Failed the negative constraint entirely by placing the astronaut on top of the horse.
  • The horse lacks any life-support equipment for a space environment.

Imagen 4.0 Generate 001

  • + Creative interpretation of a space-horse with nebulae patterns on its coat.
  • + Beautiful lighting and ethereal atmosphere.
  • Failed the negative constraint by placing the astronaut on top of the horse.
  • The composition is a bit more generic compared to the gritty detail of the other image.

Verdict: Both GPT Image 1.5 and Imagen 4.0 failed to follow the specific spatial instruction to have the 'horse on top' of the astronaut, both defaulting to the common trope of an astronaut riding a horse. GPT Image 1.5 is preferred for its superior cinematic quality, grit, and detailed rendering of the lunar environment, which better matches the 'highly detailed' part of the prompt.

The Capybara Taxi Driver

Text-to-Image

“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”

GPT Image 1.5
Imagen 4.0 Generate 001

AI Judge Analysis

GPT Image 1.5

  • + Excellent photorealistic texture on the capybara's fur and the taxi interior.
  • + Dramatic, moody cinematic lighting that feels very authentic to a New York night.
  • + Superior integration of the capybara's paws on the steering wheel.
  • The passenger in the back is slightly out of focus, though this helps with depth of field.
  • The interior feels a bit cramped compared to the wide view of Model B.

Imagen 4.0 Generate 001

  • + Crystal clear composition showing both the driver and the passenger perfectly.
  • + Accurately represents the 'bored' expression of the businesswoman as requested.
  • + Very clean UI-like clarity for all elements in the frame.
  • The image has a smooth, digital 'AI look' that lacks the grittiness and photorealism of the first.
  • The capybara's claws looking like sharp talons is slightly unnatural for the species.

Verdict: GPT Image 1.5 is the preferred choice because it achieves a high level of cinematic photorealism with convincing textures and lighting that perfectly match the 'New York at night' atmosphere. While Imagen 4.0 provides a clear and well-composed scene with a great character expression for the passenger, it looks more like a high-end digital illustration than a real photograph.

The Halloween Invitation

Text-to-Image

“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”

GPT Image 1.5
Imagen 4.0 Generate 001

AI Judge Analysis

GPT Image 1.5

  • + Excellent typography across all required text elements.
  • + Highly cohesive vintage gothic aesthetic with a dark parchment texture.
  • + Perfect adherence to all requested details including the thorns and webs border.
  • The parchment texture is quite noisy, which might affect readability at small sizes.

Imagen 4.0 Generate 001

  • + Clean, readable layout with a modern-spooky illustrative style.
  • + Good use of color contrast between the blue sky and orange pumpkin.
  • The 'parchment' element is strangely rendered as a vertical roll on the left side rather than the background.
  • Failed to correctly render the requested thorn border, showing smooth curved lines instead.
  • The composition feels more like a modern cartoon than 'vintage gothic'.

Verdict: GPT Image 1.5 is the clear winner as it perfectly captures the 'vintage gothic' aesthetic and correctly interprets all prompt elements, including the parchment background and thorny border. Imagen 4.0 Generate 001 fails to execute the border correctly and interprets 'dark parchment poster' as a weird vertical roll, resulting in a less professional invitation design.

Isometric Miniature Diorama Scenes

Text-to-Image

“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”

GPT Image 1.5
Imagen 4.0 Generate 001

AI Judge Analysis

GPT Image 1.5

  • + Excellent adherence to the text prompt requirements including 'JAPAN', 'SUSHI', and the flag icon.
  • + Strong isometric 3d diorama composition that fits the requested style perfectly.
  • + Great material rendering on the wood and ceramic elements.
  • One of the chopsticks appears to be bent or misaligned near the tip.
  • Slightly more objects than 'minimal garnish' might suggest, though it fits the theme well.

Imagen 4.0 Generate 001

  • + High realism in the food textures, especially the fish and roe.
  • + Clean, minimalist composition on the diorama base.
  • Completely failed to include the requested text ('JAPAN', 'SUSHI') and flag icon.
  • The background is gray/off-white instead of the requested solid light blue.
  • Lacks the 45° isometric 3D cartoon style requested, opting for a photographic look instead.

Verdict: GPT Image 1.5 followed every instruction in the prompt, including the specific text and icon requirements, as well as the stylistic choice of a 3D isometric cartoon scene. Imagen 4.0 Generate 001 produced a high-quality photograph of sushi but failed nearly all specific formatting constraints including the text, background color, and overall style. GPT Image 1.5 is the clear winner for its superior prompt adherence.

Adorable Baby Animals in Sunny Meadow

Text-to-Image

“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”

GPT Image 1.5
Imagen 4.0 Generate 001
100% wins 0% ties 0% wins

AI Judge Analysis

GPT Image 1.5

  • + Excellent photorealistic textures on the fur and individual blades of grass.
  • + Dynamic, playful composition with plausible physical interaction between the animals.
  • + Beautiful lighting with realistic god rays and backlighting that creates a warm, cohesive atmosphere.
  • The kitten has an extra paw visible in the tumble.
  • The butterfly Scale is slightly large compared to the animals.

Imagen 4.0 Generate 001

  • + Very clear and distinct rendering of all four requested animals.
  • + Vibrant colors and a wide variety of detailed wildflower species.
  • + Adheres well to the 'dew sparkles' requirement with visible droplets on the grass.
  • Lacks photorealism, leaning heavily into a digital illustration or greeting card style.
  • The fox appears to be floating awkwardly rather than jumping or tumbling naturally.
  • The lighting feels flat and artificial compared to the requested 'god rays' effect.

Verdict: GPT Image 1.5 successfully captures the 'hyper-photorealistic' and 'wholesome' vibe of the prompt with stunning lighting and natural fur textures, despite a minor anatomical error in the kitten's paw count. Imagen 4.0 Generate 001 feels much more like a 2D illustration, failing the photorealism requirement and lacking the atmospheric depth found in GPT's version. GPT Image 1.5 is the clear winner for its superior blend of composition, lighting, and texture realism.

Vintage Cafe Logo

Text-to-Image

“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”

GPT Image 1.5
Imagen 4.0 Generate 001

AI Judge Analysis

GPT Image 1.5

  • + Excellent classic typography with a high-end vintage feel
  • + Beautiful use of texture and shading on the cloche dome
  • + Perfect execution of the 'Est. 1720' banner
  • Failed the requirement for a light background by using a solid black one
  • The logo elements feel a bit heavy for a 'minimalist' prompt

Imagen 4.0 Generate 001

  • + Successfully followed the requirement for a light background with subtle texture
  • + Strong minimalist vector aesthetic that is clean and professional
  • + Accurate text rendering for both the name and date
  • The 'Est. 1720' ribbon is a bit too simple compared to a 'classic banner'
  • Typography is more modern than 'classic' or 'vintage'

Verdict: GPT Image 1.5 produced a much more visually impressive and stylistically accurate vintage logo, but completely failed to follow the instruction for a light background. Imagen 4.0 Generate 001 strictly adhered to all prompt constraints, including the background and minimalist style, making it the more reliable choice despite the simpler font selection.

Apollo 11: Journey to Tranquility

Text-to-Image

“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”

GPT Image 1.5
Imagen 4.0 Generate 001

AI Judge Analysis

GPT Image 1.5

  • + Successfully included all 6 requested steps with matching icons and labels.
  • + Perfect adherence to the flat-vector infographic style.
  • + Excellent text rendering for both labels and names.
  • The layout is a bit crowded with large boxes.
  • Includes names of astronauts which were not explicitly in the prompt, though they fit the theme.

Imagen 4.0 Generate 001

  • + Graceful and modern layout with a clear narrative flow.
  • + Very clean vector aesthetic and consistent iconography.
  • + Strict adherence to the requested color palette.
  • Failed to include steps 5 (Descent) and 6 (Landing).
  • The icons do not logically match the labels (e.g., Earth icon for Launch, small moon for Translunar).

Verdict: GPT Image 1.5 is the clear winner as it followed the complex multi-step instructions perfectly, including specific icons for all six mission stages. Imagen 4.0 Generate 001 produced a visually pleasing layout but missed the final two steps of the prompt and failed to align the iconography with the textual labels.

Next steps

Explore each model