OpenAI's state-of-the-art image generation model with better instruction following and adherence to prompts
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
GPT Image 1.5
#7 of 62 in Text-to-Image
Imagen 4.0 Fast Generate 001
#53 of 62 in Text-to-Image
Where the votes landed
GPT Image 1.5
100.0%
win rate
Ties
0.0%
Imagen 4.0 Fast Generate 001
0.0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
GPT Image 1.5
- + Perfectly follows the spatial placement of the plant behind the cube.
- + Excellent refraction and reflection of the plant through the glass panels.
- + Very realistic thickness to the glass and texture on the red book.
- − The blue sphere feels slightly large for a 'small' sphere.
- − The lighting is a bit flat compared to the dramatic shadows in the competitor.
Imagen 4.0 Fast Generate 001
- + Beautiful lighting and shadow play from the left window.
- + The blue sphere has a nice glassy, marble-like texture.
- + Clear, sharp focus on all objects.
- − Fails the spatial requirement: the plant is sitting next to the cube rather than behind it.
- − The base of the cube appears to be a mirror rather than transparent glass.
- − Perspective distortion on the top edge of the cube makes the book appear to be floating slightly.
Verdict: GPT Image 1.5 is the clear winner because it accurately follows the complex spatial instructions, successfully showing the green plant behind the cube and visible through the glass. While Imagen 4.0 Fast Generate 001 provides a more aesthetically pleasing lighting setup, it fails the prompt by placing the plant to the side and creating a mirrored base for the cube instead of a transparent one.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
GPT Image 1.5
- + Excellent depiction of rain with visible droplets on the man's jacket and hat.
- + Realistic skin texture and an authentic, natural-looking pose for the subject.
- + Strong cinematic composition with beautiful reflections and light bokeh in the background.
- − The motion blur of the passing car is somewhat static and lacks a sense of true speed.
Imagen 4.0 Fast Generate 001
- + Successfully incorporates the 'imperfect framing' prompt with a window-like border.
- + Clear, high-quality reflection on the wet pavement.
- + Good natural skin texture and facial detail.
- − The 'motion blur from passing cars' is completely absent; the cars in the background are sharp.
- − Anatomical issues with the hands, specifically the right hand which looks mangled.
- − The composition feels slightly more staged than 'candid'.
Verdict: GPT Image 1.5 is the clear winner as it successfully balances all technical requirements, particularly the atmospheric effects of the light rain and the candidate feel requested. Imagen 4.0 Fast Generate 001 fails to produce the requested motion blur and has significant artifacts in the subject's hands.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
GPT Image 1.5
- + Perfect adherence to all prompt elements including armor, beads, scars, and lighting.
- + Exceptional detail in skin texture, armor engravings, and leather straps.
- + Dynamic composition with realistic bokeh and cinematic torchlight reflections.
- − The hair strands over the eye are slightly blurred compared to the sharp facial features.
Imagen 4.0 Fast Generate 001
- + High resolution and natural-looking outdoor lighting.
- − Failed to follow the prompt entirely, showing a modern man in a garden instead of a paladin.
- − Missing all requested elements: armor, scars, torchlight, and beads.
- − Distance is a wide shot rather than the requested close portrait.
Verdict: GPT Image 1.5 followed the prompt with extreme precision, delivering a cinematic and highly detailed portrait of a battle-worn warrior that matched every specific requirement. Imagen 4.0 Fast Generate 001 failed the task completely, producing an unrelated image of an elderly man in a garden which suggests a total failure in prompt interpretation or a generation error.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
GPT Image 1.5
- + Excellent text readability with no spelling errors
- + High-quality, appetizing food photography that relates to the menu items
- + Clean, professional layout that matches the 'casual dining' instruction
- − Layout is a split column rather than a strict 'grid' of photos
Imagen 4.0 Fast Generate 001
- + Successfully creates a grid-based layout for the photos
- + Uses vibrant color blocks as accents
- + Includes the requested white background
- − All text is gibberish and 'Appetizers' is misspelled
- − Food photos are repetitive and only show pizza, regardless of the section label
- − Overall resolution and clarity of the text and images are poor
Verdict: GPT Image 1.5 produced a functional and professionally designed menu with perfectly rendered text and high-quality food photography. In contrast, Imagen 4.0 Fast Generate 001 failed significantly on text generation, producing gibberish and repetitive imagery that did not correlate with the menu sections.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
GPT Image 1.5
- + Excellent photorealistic textures on the beef patty and fresh vegetables.
- + Dynamic composition with embers and flying pieces that create a strong sense of motion.
- + Perfect text rendering and integration with the overall visual theme.
- − The lighting is very high contrast, which might make some details feel slightly cluttered.
Imagen 4.0 Fast Generate 001
- + Clean, professional studio-style layout.
- + Excellent literal interpretation of the fire-outlined text.
- + Clear separation of ingredients for an 'exploded' view.
- − The textures look more like 3D renders than photorealistic photography.
- − Typo in the secondary message: 'LIMITED TIME ON ONLY'.
- − The composition feels a bit more static and less punchy than Model A.
Verdict: GPT Image 1.5 is the clear winner due to its superior photorealistic textures and successful execution of all text elements without errors. While Imagen 4.0 Fast Generate 001 provides a clean layout, its failure on the text accuracy and the more synthetic-looking lighting makes it less effective as a dynamic advertisement.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
GPT Image 1.5
- + Excellent text rendering with perfect spelling and realistic chalk texture.
- + Authentic handwriting style with natural variations as requested.
- + Realistic chalkboard background with smudges and atmospheric lighting.
- − The lighting creates some glare at the top which slightly reduces contrast.
Imagen 4.0 Fast Generate 001
- + Clear, legible text and a well-defined wooden frame.
- + Followed the prompt to include a wooden border for the chalkboard.
- − Multiple spelling errors including 'Octuphus' and 'Cookes'.
- − Redundant text repetition of the cookie item.
- − The text looks more like a digital font rather than natural chalk handwriting.
Verdict: GPT Image 1.5 followed the prompt perfectly, producing realistic, well-textured handwriting with no spelling errors. In contrast, Imagen 4.0 Fast Generate 001 suffered from several spelling mistakes, repetitive items, and a font style that appeared too digital for a chalk-based prompt.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
GPT Image 1.5
- + Excellent cinematic lighting and texture on the horse's coat.
- + Detailed background featuring planets, a lunar lander, and realistic space debris.
- + Highly surreal concept of a horse gallop kicking up space dust on a lunar surface.
- − Failed the specific spatial instruction 'horse on top, not vice versa'.
- − The astronaut's left leg disappears into the horse's body.
Imagen 4.0 Fast Generate 001
- + Features a clear, well-rendered face inside the helmet.
- + Clean, ethereal composition with a smooth nebula background.
- + Good anatomical consistency for the horse and rider legs.
- − Failed the specific spatial instruction 'horse on top, not vice versa'.
- − The pose is somewhat static compared to the cinematic request.
- − The lack of riding gear (saddle/reins) makes the astronaut appear to be floating through the horse.
Verdict: Both GPT Image 1.5 and Imagen 4.0 failed the key 'surreal' negative constraint to have the horse on top of the astronaut, instead opting for the common 'astronaut riding a horse' trope. GPT Image 1.5 is the preferred image because its cinematic quality, detailed lunar background, and dynamic action better capture the requested 'surreal and highly detailed' aesthetic, whereas Imagen 4.0 is more simplistic.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
GPT Image 1.5
- + Excellent photorealism with gritty, cinematic lighting that matches a NYC night atmosphere.
- + Perfect adherence to the 'both front paws on the steering wheel' instruction.
- + The capybaras expression and fur texture are highly realistic.
- − The passenger's hands and phone interaction are slightly blurry.
- − The framing is very tight, showing less of the car interior.
Imagen 4.0 Fast Generate 001
- + Captures the 'bored' expression of the businesswoman very effectively.
- + Clean composition that shows more of the taxi's driver-side interior.
- + High level of detail in the capybara's chauffeur-style uniform.
- − The capybara's hands/paws are not realistically placed on the steering wheel, appearing to hover or merge awkwardly.
- − The lighting feels a bit more like a studio set than a genuine night-time NYC street.
- − The capybara's head is slightly disconnected from its neck/shoulders.
Verdict: GPT Image 1.5 is the clear winner due to its superior photorealistic quality and faithful adherence to the physical interaction requested (paws on the steering wheel). While Imagen 4.0 Fast Generate 001 captures the passenger's expression well, it fails on the anatomical logic of the capybara driving, whereas GPT Image 1.5 creates a convincing, gritty cinematic scene that feels like a real photograph.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
GPT Image 1.5
- + Excellent typographic accuracy for all requested text segments.
- + Highly cohesive gothic aesthetic with detailed thorn and web borders.
- + Atmospheric lighting that creates a genuine vintage cinematic feel.
- − The dark color palette slightly reduces the contrast of the bottom event details.
Imagen 4.0 Fast Generate 001
- + Clean, readable layout with a distinct parchment border.
- + Good interpretation of the 'twisted trees' prompt element.
- − Significant spelling errors including 'IINVITATION' and 'FNIGHTS'.
- − Visual style appears more like digital clip-art than a 'vintage gothic' poster.
- − Text layout at the bottom is bunched together awkwardly.
Verdict: GPT Image 1.5 follows all prompt instructions perfectly, delivering an atmospheric and highly detailed vintage invitation with flawless text. In contrast, Imagen 4.0 Fast Generate 001 suffers from multiple spelling errors and a more simplified, cartoonish art style that lacks the cinematic polish requested.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
GPT Image 1.5
- + Excellent material rendering with realistic wood, ceramic, and food textures.
- + Dynamic and high-quality 3D diorama composition with rich details like a teapot and soy sauce bottle.
- + Perfect text rendering and alignment with the requested flag icon.
- − Includes more elements than the 'minimal garnish' requested in the prompt.
- − The 45° angle is slightly less rigid than a strict isometric projection.
Imagen 4.0 Fast Generate 001
- + Strict adherence to 'minimal' aesthetic and isometric perspective.
- + Clean, soft-cartoon style that feels cohesive and simple.
- + Accurate placement of text and flag icon.
- − Rice texture appears as uniform white bumps, lacking the PBR realism requested.
- − The composition feels slightly empty compared to the rich detail of the other model.
- − The white border around the square canvas is unnecessary.
Verdict: GPT Image 1.5 is the clear winner due to its superior texture rendering and high-quality PBR materials, making the scene feel like a professional 3D render. While Imagen 4.0 Fast Generate 001 captures the 'minimal' and 'isometric' aspects well, it lacks the visual depth and material sophistication that the prompt explicitly requested.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
GPT Image 1.5
- + Perfect adherence to all requested animal types (Golden Retriever, Tabby, Bunny, Fox).
- + Excellent dynamic composition with the 'tumbling' and 'chasing' actions clearly depicted.
- + Strong lighting effects with visible god rays and sparkling dew consistent with a sunrise.
- − The cat's anatomy is slightly distorted, particularly the leg with too many toe pads.
- − The fox's paw in the bottom right looks somewhat muddy and poorly defined.
Imagen 4.0 Fast Generate 001
- + High level of fur detail and realistic textures on all animals.
- + Soft, pleasing bokeh in the background with consistent lighting.
- − Failed multiple prompt requirements: no butterflies, no 'tabby' kitten, and no 'golden retriever' (dog is brown/white).
- − The animals are sitting still rather than 'playfully chasing' or 'tumbling' as requested.
Verdict: GPT Image 1.5 followed the complex prompt much more accurately, including all four specific animal breeds and the action of chasing butterflies. While Imagen 4.0 has high textural quality, it missed several key descriptors including the butterfly element and the specific dog and cat breeds requested. GPT Image 1.5 better captured the 'joyful wholesome vibe' through dynamic posing.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
GPT Image 1.5
- + Excellent typography with a hand-lettered feel
- + Detailed vintage texture and shading on the cloche
- + Perfect text accuracy including the accented letter
- − Ignored the 'light background' instruction, resulting in a black background
- − The cloche handle is slightly off-center from the steam
Imagen 4.0 Fast Generate 001
- + Successfully followed the 'light background' and 'minimalist' prompt instructions
- + Clean, balanced vector-style layout
- + Accurate spelling and good use of the banner element
- − Included minor gibberish text ('AFFD CARO') above the main name
- − The steam effect is very faint and hard to see
Verdict: GPT Image 1.5 produced a more stylistically impressive logo with beautiful texture, but it completely failed the instruction for a light background. Imagen 4.0 followed all instructions better, including the color palette and background, although it included some unnecessary filler text.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
GPT Image 1.5
- + Perfectly adhered to the 6-step sequence with accurate labels.
- + Legible, correctly spelled text including astronaut names and 'Tranquillity'.
- + Clean, consistent flat-vector style across all panels.
Imagen 4.0 Fast Generate 001
- + Very clean and minimalist layout with a modern infographic feel.
- + Used the requested color palette effectively.
- + Accurate icons for the lunar module descent and landing.
- − Failed to follow the 6-step sequence, repeating 'Launch' and missing 'Earth Orbit'.
- − Spelling errors such as 'APOLO', 'SATURN VICON', and 'MOOR'.
- − Confusing diagram logic with non-sequential arrows.
Verdict: GPT Image 1.5 performed significantly better by following the specific 6-step sequence outlined in the prompt and maintaining perfect spelling. While Imagen 4.0 Fast Generate 001 had a nice minimalist aesthetic, its failure to organize the information logically and the presence of numerous typos made it an inferior infographic.
Explore each model
Google's Imagen 4.0 Fast model optimized for speed and efficiency, suitable for high-volume image generation tasks