Head to head
Esc

Models · slot A

to navigate to pick

GPT Image 1.5 OpenAI Imagen 4.0 Ultra Generate 001 Google

Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.

GPT Image 1.5

27.1 arena score

#5 of 62 in Text-to-Image

Top 3 in Image Editing
Skill signature · Text-to-Image

Imagen 4.0 Ultra Generate 001

22.1 arena score

#33 of 62 in Text-to-Image

Vote tally

Where the votes landed

GPT Image 1.5

76.9%

win rate

Ties

7.7%

Imagen 4.0 Ultra Generate 001

15.4%

win rate

76.9% 7.7% ties 15.4%
Shared challenges 13

Challenge by challenge

The strongest take from each model on every shared challenge, with the AI judge's read.

Geometric Composition

Text-to-Image

“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”

GPT Image 1.5
Imagen 4.0 Ultra Generate 001
50% wins 25% ties 25% wins

AI Judge Analysis

GPT Image 1.5

  • + Excellent adherence to the 'small' scale of the sphere relative to the cube.
  • + Very realistic rendering of thick glass edges and reflections.
  • + Natural lighting and high-quality textures on the table and book cover.
  • The green plant in the background is quite cluttered compared to the clean aesthetic of the foreground.

Imagen 4.0 Ultra Generate 001

  • + Beautiful bokeh effect and professional photographic composition.
  • + Impressive text rendering on the book spine ('ANCIENT TALES').
  • + Sophisticated handling of light and shadow across the wooden table.
  • The blue sphere appears to be floating rather than sitting inside the cube.
  • The cube looks more like a solid glass block than a hollow container.

Verdict: GPT Image 1.5 followed the prompt instructions more accurately, particularly regarding the placement and scale of the sphere inside the glass cube. While Imagen 4.0 Ultra Generate 001 produced a more artistic and visually stunning image with impressive text rendering, the central subject looks like a solid block with a floating sphere, failing the physics of the requested scene.

Candid Street Photography

Text-to-Image

“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”

GPT Image 1.5
Imagen 4.0 Ultra Generate 001
100% wins 0% ties 0% wins

AI Judge Analysis

GPT Image 1.5

  • + Captures motion blur in the passing car perfectly as requested.
  • + The 50mm shallow depth of field is well-executed with professional-looking bokeh.
  • + Excellent environmental storytelling with the toolkit and damp clothing textures.
  • The bicycle mechanics are slightly physics-defying with the rear rack and chain alignment.
  • The red reflection on the pavement is a bit oversaturated compared to the ambient light.

Imagen 4.0 Ultra Generate 001

  • + Incredible skin texture and facial detail, looking very realistic and non-stylized.
  • + Highly detailed rain droplets on the man's jacket.
  • + Excellent composition using the brick wall as a leading line.
  • Failed to include motion blur on the passing car, which appears static.
  • The bicycle frame and pedals have some structural inconsistencies where they meet the chainring.

Verdict: GPT Image 1.5 adhered better to the technical requirements of the prompt, specifically the motion blur and the overall 'candid street photo' feel. While Imagen 4.0 Ultra provided superior skin textures and sharp details on the clothing, it missed the motion blur instruction and felt slightly more posed.

Fantasy Warrior

Text-to-Image

“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”

GPT Image 1.5
Imagen 4.0 Ultra Generate 001

AI Judge Analysis

GPT Image 1.5

  • + Exceptional photographic realism in skin texture and eyes.
  • + Beautiful lighting integration with realistic reflections on the engraved armor.
  • + Effective use of shallow depth of field reflecting the prompt.
  • The 'paladin' character leans slightly into a generic 'pretty warrior' aesthetic compared to the battle-worn request.

Imagen 4.0 Ultra Generate 001

  • + Strong adherence to the 'battle-worn' descriptor with visible aging and scars.
  • + Very clear and detailed bead-work throughout the braids.
  • + Good visual balance with the inclusion of the physical torch as a light source.
  • The skin texture looks slightly more digital/rendered compared to the realism of Model A.
  • The bokeh sparks appear a bit more repetitive and artificial.

Verdict: GPT Image 1.5 wins due to its superior photographic quality and the way it handles light reflecting off the ornate armor textures. While Imagen 4.0 Ultra captures the character's 'battle-worn' age and braids more distinctly, the overall composition and lifelike eyes of GPT Image 1.5 create a more compelling and professionally rendered image.

Modern Clean Menu

Text-to-Image

“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”

GPT Image 1.5
Imagen 4.0 Ultra Generate 001
100% wins 0% ties 0% wins

AI Judge Analysis

GPT Image 1.5

  • + Perfect text rendering with zero spelling errors
  • + Excellent logical organization of categories and pricing
  • + High-quality, appetizing food photography that matches the text descriptions
  • Simple layout is highly functional but follows a very standard template style

Imagen 4.0 Ultra Generate 001

  • + Stronger adherence to the 'grid' request in the prompt
  • + Clean, modern minimalist aesthetic with plenty of white space
  • Garbled, nonsensical text in titles and descriptions
  • Confusing grouping where pizza items appear under 'Appetizers'
  • Inconsistent image sizes and alignment within the grid

Verdict: GPT Image 1.5 is the clear winner as it produces a fully functional, professional menu with perfect text and logical content mapping. While Imagen 4.0 Ultra attempted a more sophisticated grid layout, it failed significantly on text legibility and categorized the food items incorrectly.

Magic Burger Explosion: Fiery Photorealism Challenge

Text-to-Image

“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”

GPT Image 1.5
Imagen 4.0 Ultra Generate 001

AI Judge Analysis

GPT Image 1.5

  • + Excellent adherence to text rendering with vibrant glowing effects.
  • + Dynamic and high-energy composition with realistic textures on the patty and bun.
  • + The 'exploded' effect feels more visceral with sauce drips and small flying fragments.
  • The background is slightly busy, which can distract from the main subject.
  • The '€' symbol and numbers are slightly crowded in the starburst.

Imagen 4.0 Ultra Generate 001

  • + Very clean and professional composition with a focused central subject.
  • + Clear, legible typography that follows the 'limited time only' placement instruction perfectly.
  • + Photorealistic textures on the lettuce and tomatoes are particularly sharp.
  • The 'exploded' effect feels a bit static and less 'dynamic' than the other model.
  • The lighting on the burger components is a bit flat compared to the fiery background.

Verdict: GPT Image 1.5 captures the 'dynamic' and 'fiery' energy requested much better, using small particles and dripping sauces to enhance the feeling of motion. Imagen 4.0 Ultra is remarkably clean and displays superior text legibility, but it feels more like a standard product shot than an 'exploded' magic burger. GPT Image 1.5 is the winner for its superior atmospheric detail and more exciting interpretation of the prompt.

Chalkboard Menu

Text-to-Image

“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”

GPT Image 1.5
Imagen 4.0 Ultra Generate 001

AI Judge Analysis

GPT Image 1.5

  • + Excellent chalk texture with realistic dust and smudging effects.
  • + Authentic handwritten aesthetic with natural variations in letter size and spacing.
  • + Perfect spelling and adherence to the cut-off prompt item name.
  • The 'cursive' requirement for the title is only partially met, leaning more toward print.

Imagen 4.0 Ultra Generate 001

  • + Strong composition with a clear frame and environmental context.
  • + Very clean and readable text rendering.
  • + Good contrast between the chalkboard and the letters.
  • The text looks more like a digital font or a paint marker rather than chalk.
  • Lacks the specific 'elegant cursive' style requested for the title.
  • The letter forms are overly uniform, missing the natural handwriting variation requested.

Verdict: GPT Image 1.5 is the clear winner for its superior rendering of chalk texture and authentic handwriting, which perfectly captures the 'realistic chalk' requirement. While Imagen 4.0 Ultra has a nice layout, its text appears too clean and digital, failing to mimic the dusty, textured quality of a real chalkboard as effectively as GPT Image 1.5.

The Reversed Rodeo

Text-to-Image

“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”

GPT Image 1.5
Imagen 4.0 Ultra Generate 001

AI Judge Analysis

GPT Image 1.5

  • + Excellent texture and cinematic lighting on both the horse and the space suit.
  • + Dynamic composition with a sense of motion in the dust and debris.
  • + High level of detail in the background, including nebulae and planetary bodies.
  • Failed the negative constraint: the prompt asked for the horse to be on top of the astronaut, but the astronaut is riding the horse.
  • The horse lacks a helmet or breathing apparatus, which reduces the 'surreal' or 'logical' consistency within the celestial setting.

Imagen 4.0 Ultra Generate 001

  • + Creative interpretation with futuristic details like light-up horseshoes and a custom horse visor.
  • + Clean, polished visual style with a dreamy, surreal atmosphere.
  • + The horse has a specialized space helmet, which aligns better with the space environment.
  • Failed the negative constraint: like the other model, it depicts an astronaut riding a horse instead of the horse on top.
  • The composition feels slightly more static compared to the action-oriented style of the competitor.

Verdict: Both models failed the specific prompt instruction to place the horse on top of the astronaut, with both defaults resulting in the traditional 'astronaut riding horse' trope. GPT Image 1.5 is the preferred image because of its superior cinematic textures, gritty detail, and more complex background composition, whereas Imagen 4.0 Ultra feels more like a generic digital illustration.

The Capybara Taxi Driver

Text-to-Image

“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”

GPT Image 1.5
Imagen 4.0 Ultra Generate 001

AI Judge Analysis

GPT Image 1.5

  • + Excellent photorealism with gritty cinematic lighting
  • + Perfect adherence to the 'bored expression' prompt for the passenger
  • + Texture of the capybara fur and clothing looks very realistic
  • The capybara's paws look slightly more like human hands/fingers than rodent paws

Imagen 4.0 Ultra Generate 001

  • + Includes more of the taxi interior and the dashboard
  • + Clear text on the taxi cap
  • + Vibrant colors and clean composition
  • The passenger looks distressed or anxious rather than bored
  • The image has a smooth, digital 'AI' sheen that lacks the requested photorealism
  • The capybara's claws are overly sharp and look like talons

Verdict: GPT Image 1.5 is the clear winner as it achieves a truly photorealistic cinematic look that feels like a real movie still. It perfectly captured the subtle 'bored' expression of the businesswoman, whereas Imagen 4.0 Ultra gave her an expression of concern and had a much more synthetic, digital texture.

The Halloween Invitation

Text-to-Image

“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”

GPT Image 1.5
Imagen 4.0 Ultra Generate 001

AI Judge Analysis

GPT Image 1.5

  • + Excellent 'vintage parchment' texture and aged aesthetic.
  • + Sophisticated, integrated typography that fits the gothic theme perfectly.
  • + Fantastic use of atmospheric, cinematic lighting and a moody background.
  • The thorn border is a bit messy and overlaps the text elements slightly.

Imagen 4.0 Ultra Generate 001

  • + Very clean and legible typography throughout the image.
  • + Strong adherence to all layout elements including the scroll banner and border thorns.
  • + Good use of color contrast between the cool blue sky and warm pumpkin glow.
  • Lacks the requested 'vintage parchment' feel, appearing more like a modern digital illustration.
  • The border thorns look somewhat plastic and lack the organic texture of Model A.

Verdict: GPT Image 1.5 is the winner because it captures the 'vintage' and 'dark parchment' elements of the prompt with much more artistic depth and atmosphere. While Imagen 4.0 Ultra is more legible and clean, it feels like a modern cartoon or clip-art illustration, whereas GPT Image 1.5 looks like a high-quality physical movie poster or authentic invitation.

Isometric Miniature Diorama Scenes

Text-to-Image

“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”

GPT Image 1.5
Imagen 4.0 Ultra Generate 001

AI Judge Analysis

GPT Image 1.5

  • + Excellent PBR textures and lighting that create a high-quality 3D render feel.
  • + Perfect centered typography and flag icon alignment.
  • + Rich and appealing detailed composition with varied elements like the teapot and soy sauce.
  • The diorama base contains a lot of extra elements contrary to the 'minimal garnish' request.
  • One chopstick is slightly misaligned with its base.

Imagen 4.0 Ultra Generate 001

  • + Strong adherence to the 'minimal' request with a clean, simple white diorama base.
  • + Highly accurate 45° isometric perspective.
  • + Text is sharp and legible.
  • The sushi and ginger textures look a bit plastic and less realistic than requested.
  • Typography and flag placement are slightly off-center compared to the image subject.
  • The plate appears to be floating unnaturally above the base.

Verdict: GPT Image 1.5 produced a much more visually stunning result with superior PBR materials and professional lighting, giving it a high-end 3D render quality. While Imagen 4.0 followed the 'minimal' part of the prompt more closely, its textures appeared flat and the composition felt less polished, making GPT Image 1.5 the overall winner for its sheer visual quality and detail.

Adorable Baby Animals in Sunny Meadow

Text-to-Image

“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”

GPT Image 1.5
Imagen 4.0 Ultra Generate 001
100% wins 0% ties 0% wins

AI Judge Analysis

GPT Image 1.5

  • + Excellent texture work on the fur and paws, feeling very tactile and soft.
  • + Beautiful lighting with visible god rays and realistic dew sparkles on the flowers.
  • + Captures the 'tumbling together' part of the prompt with a natural, intertwined composition.
  • The fox kit's anatomy looks a bit distorted, especially the mouth and eye alignment.

Imagen 4.0 Ultra Generate 001

  • + Very clear and vibrant colors with distinct butterfly designs.
  • + Accurate representation of all four requested animals with clear silhouettes.
  • + Good usage of dew drops on the grass and strong lighting rays.
  • Has a notably 'digital illustration' look rather than the requested 'hyper-photorealistic' style.
  • The animals feel more like they are floating or standing in place rather than 'tumbling together'.

Verdict: GPT Image 1.5 followed the stylistic requirements much better, delivering a truly photorealistic scene with soft, detailed textures and a natural, chaotic sense of play. Imagen 4.0 Ultra Generate 001 produced a very clean and cute image, but it leans heavily into a CGI/illustrative aesthetic rather than the realism requested by the prompt.

Vintage Cafe Logo

Text-to-Image

“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”

GPT Image 1.5
Imagen 4.0 Ultra Generate 001
100% wins 0% ties 0% wins

AI Judge Analysis

GPT Image 1.5

  • + Excellent typography with correct accent on 'Caffè'
  • + Strong vector illustration style with attractive shading
  • + Clear and legible ribbon banner
  • Ignored the request for a light background with subtle texture
  • Steam effect is a bit chunky compared to the overall logo style

Imagen 4.0 Ultra Generate 001

  • + Perfectly follows the light background with subtle texture requirement
  • + Captures a more minimalist, clean aesthetic
  • + Accurate text rendering and placement
  • The 'E' in 'Caffè' has a slightly awkward tail
  • Steam lines are very thin and faint compared to other elements

Verdict: Imagen 4.0 Ultra Generate 001 is the winner because it followed all prompt instructions, including the specific request for a textured light background. While GPT Image 1.5 produced a high-quality graphic with superior illustration details, it completely ignored the background color requirement, resulting in a black background.

Apollo 11: Journey to Tranquility

Text-to-Image

“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”

GPT Image 1.5
Imagen 4.0 Ultra Generate 001
50% wins 0% ties 50% wins

AI Judge Analysis

GPT Image 1.5

  • + Strictly followed the requested 6-step chronological sequence with accurate icons.
  • + Excellent text rendering for main labels and supporting details like astronaut names.
  • + Perfect adherence to the flat-vector style and specified NASA color palette.
  • Includes a redundant Earth icon in the first 'Launch' step which wasn't specifically requested.
  • The Saturn V rocket is slightly cut off at the top of the frame.

Imagen 4.0 Ultra Generate 001

  • + Captures a professional poster layout with a centered focal point.
  • + Used the requested color palette effectively across the design.
  • Failed to follow the requested 6-step chronological structure, opting for a circular layout with nonsensical steps.
  • Text is mostly gibberish or repetitive, failing to provide the requested information labels.
  • Icons do not clearly correspond to the specific mission phases requested (e.g., Saturn V, lunar orbit).

Verdict: GPT Image 1.5 is the clear winner as it followed every instruction, including the specific 6-step sequence, icon types, and text labels. Imagen 4.0 Ultra generated a visually appealing layout but failed significantly on prompt adherence, producing garbled text and ignoring the specific chronological steps requested.

Next steps

Explore each model