OpenAI's previous image generation model that accepts both text and image inputs and produces image outputs
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
GPT Image 1
#29 of 62 in Text-to-Image
Imagen 4.0 Ultra Generate 001
#33 of 62 in Text-to-Image
Where the votes landed
GPT Image 1
0%
win rate
Ties
0%
Imagen 4.0 Ultra Generate 001
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
GPT Image 1
- + Excellent adherence to the glass cube geometry with clearly defined edges
- + Natural warm lighting and realistic textures on the wooden table and plant
- + Accurate arrangement of all spatial elements requested in the prompt
- − The blue sphere is quite large and matte, whereas the prompt suggested a 'small' sphere
- − The glass cube appears to have an open frame or very thick seams rather than being a solid glass object
Imagen 4.0 Ultra Generate 001
- + Beautiful rendering of light and caustic reflections through the glass
- + Highly detailed red book with legible, elegant text
- + Realistic scale of the small blue sphere relative to the cube
- − The plant is more 'beside' or 'flanking' the cube rather than being clearly 'behind' and visible through the glass
- − The blue sphere appears to be floating unnaturally rather than resting inside the cube
Verdict: Both models followed the prompt instructions very well, but GPT Image 1 feels like a more accurate spatial arrangement of the plant behind the glass. However, Imagen 4.0 Ultra Generate 001 produced a more aesthetically pleasing image with superior lighting, better scale for the 'small' sphere, and impressive detail on the book spine.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
GPT Image 1
- + Excellent skin and hair texture that feels authentic for a 50mm candid portrait.
- + Strong cinematic atmosphere with realistic light rain droplets on the bike seat.
- + The color grading is cohesive and muted, fitting a 'no stylization' realistic request.
- − The bike anatomy is slightly jumbled, particularly the chain area and how the hand interacts with it.
- − The background cars look a bit static and lack the 'motion blur' requested.
Imagen 4.0 Ultra Generate 001
- + Successfully captures the requested motion blur in the passing white car.
- + The lighting on the pavement and brick wall is very realistic and shows great reflection work.
- + The framing feels more 'imperfect' and candid as requested.
- − The hands have anatomical issues, with fingers merging into the bike chain.
- − The bicycle's orientation is physically awkward, appearing to lean against the wall at a sharp angle without support.
- − The face has a slightly 'over-sharpened' digital look compared to the natural texture of the other model.
Verdict: GPT Image 1 captures a more natural and intimate portrait with superior skin textures and a beautifully captured 'light rain' feel, though it misses the motion blur request. Imagen 4.0 Ultra Generate 001 follows the prompt more literally regarding motion blur and framing, but suffers from significant anatomical errors in the hands and bicycle structure. GPT Image 1 is the winner for its overall photographic coherence and believable cinematic quality.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
GPT Image 1
- + Excellent photorealistic facial texture and lifelike eyes
- + Beautifully intricate engraving on the armor with subtle wear
- + Natural and atmospheric lighting that integrates perfectly with the bokeh sparks
- − The hair beads are very subtle and few in number
- − Fabric underlayer details are mostly obscured by shadow
Imagen 4.0 Ultra Generate 001
- + Strong adherence to the bead requirement with many distinct silver beads
- + Clear visibility of the leather straps and metal buckles mentioned in the prompt
- + Visible torch source adds context to the lighting
- − The scars look a bit like digital brush strokes rather than realistic tissue
- − Overall image has a slightly more artificial, CG look compared to the photorealism of the competitor
Verdict: GPT Image 1 excels in photorealism, creating a hauntingly human portrait with atmospheric lighting and incredible armor detail. Imagen 4.0 Ultra follows the prompt's technical requirements more literally, specifically with the leather straps and numerous hair beads, but has a more synthetic, game-cinematic aesthetic. GPT Image 1 is the winner for its superior visual quality and more natural interpretation of 'battle-worn'.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
GPT Image 1
- + Excellent high-resolution food photography that looks appetizing and professional.
- + Clean, readable layout with appropriate use of white space and bold sans-serif fonts.
- + Sections for Appetizers and Pizza are clearly labeled and logically organized.
- − Missing the 'mains' header section despite having images of main dishes.
- − Text contains significant spelling errors in the descriptions (e.g., 'Apperoiation descrigion').
Imagen 4.0 Ultra Generate 001
- + Successfully includes all three requested sections: appetizers, pizza, and mains.
- + Layout follows a strict grid pattern as requested.
- + Good use of vibrant color accents on the sides of images.
- − The food photos are very small and some contain visual artifacts.
- − Text rendering is messy with many garbled characters and gibberish names.
- − The composition feels slightly crowded with 12 small images.
Verdict: GPT Image 1 produces a much more professional and appetizing result with superior food photography and a cleaner aesthetic suitable for a modern restaurant. While Imagen 4.0 Ultra included the 'Mains' header that GPT missed, its text is unreadable and the small grid size makes the food look less appealing.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
GPT Image 1
- + Excellent typography rendering with consistent fiery glow effect.
- + Rich, photorealistic textures on the burger bun and lettuce.
- + Strong use of depth and lighting on the floating ingredients.
- − Failed to render the price accurately, showing '€.99' instead of '€6.99'.
- − Lacks the sauce layer explicitly mentioned in the prompt.
Imagen 4.0 Ultra Generate 001
- + Accurately rendered all requested text including the correct price of '€6.99'.
- + Dynamic swirling background enhances the sense of motion and 'magic'.
- + Includes all requested components including multiple layers of sauce.
- − The center of the starburst is a bit cluttered compared to the clean layout of the other model.
- − The top bun's sesame seed distribution looks slightly repetitive.
Verdict: While both models provided high-quality photorealistic images, Imagen 4.0 Ultra is the superior choice because it adhered to all text requirements, whereas GPT Image 1 omitted the most critical part of the price. Imagen 4.0 Ultra also provided a more dynamic background that better fit the 'magic' theme of the burger.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
GPT Image 1
- + Excellent chalk texture showing individual particulates and pressure variations
- + Perfect text accuracy including the final truncated item
- + Consistent handwriting style across all lines of text
- − The 'cursive' request for the title was not fully met, as it remains blocky
- − Lacks the atmospheric background context of a cafe
Imagen 4.0 Ultra Generate 001
- + Successfully applied a more flowing, cursive-influenced style for the title
- + Cozy atmosphere achieved via the brick wall and lighting
- + Clear layout with consistent bullet points
- − The text looks more like a digital marker or vector font than real chalk
- − Missing the chalky texture and grit requested in the prompt
- − Failed to properly finish the third item based on the truncated prompt text
Verdict: GPT Image 1 is the clear winner for its superior rendering of chalk textures and perfect adherence to the prompt's specific text. While Imagen 4.0 Ultra creates a more complete scene with better 'cursive' interpretation, its text looks like a clean digital font rather than the physical chalk requested.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
GPT Image 1
- + Excellent cinematic lighting and texture on the horse's coat.
- + Strongly captures a 'surreal' and dramatic atmosphere through minimalist composition.
- − Failed the negative constraint; the astronaut is riding the horse.
- − The horse has no breathing apparatus or gear for the space environment, reducing the 'highly detailed' sci-fi aspect.
Imagen 4.0 Ultra Generate 001
- + Included creative sci-fi details like a horse helmet and glowing hoof-boots.
- + Higher color vibrance and complex background elements like the nebula and asteroid.
- − Failed the negative constraint; the astronaut is riding the horse instead of the horse being on top.
- − The composition feels a bit cluttered compared to the cinematic request.
Verdict: Both GPT Image 1 and Imagen 4.0 Ultra failed the specific logic constraint of the prompt to have the horse on top of the astronaut. However, GPT Image 1 is the superior image due to its much more convincing cinematic lighting and realistic textures, whereas Imagen 4.0 Ultra feels more like a digital illustration for a children's book.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
GPT Image 1
- + Excellent photorealistic texture on the capybara fur and the jacket
- + Perfect composition with the woman correctly placed in the back seat
- + Cinematic lighting that captures the authentic feel of a New York night taxi ride
- − The capybara's front paws look a bit like primate hands
- − Lacks specific 'TAXI' text on the cap as implied by common tropes
Imagen 4.0 Ultra Generate 001
- + Features the requested cap with 'TAXI' text clearly rendered
- + Clean and sharp car interior details
- + Good adherence to the pose for both characters
- − The passenger is sitting in the front passenger seat instead of the back seat
- − The capybara's claws are overly long and sharp, looking more like a sloth's
- − The overall lighting and texture look more like 3D CGI than a photorealistic scene
Verdict: GPT Image 1 is much more successful at achieving the requested 'photorealistic' style and correctly placing the woman in the back seat, whereas Imagen 4.0 Ultra placed her in the front next to the driver. While Imagen 4.0 Ultra has sharper textures, GPT Image 1 captures a far more convincing atmosphere and better captures the mundane, bored expression of the passenger.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
GPT Image 1
- + Excellent adherence to the 'vintage' and 'dark parchment' aesthetic.
- + Perfect text layout with clear, legible event details.
- + Highly atmospheric and moody lighting that feels cinematically grounded.
- − The thorns requested in the border are very subtle and blend into the branches.
- − The color palette is somewhat monochromatic compared to the requested 'glowing' effect.
Imagen 4.0 Ultra Generate 001
- + Strong inclusion of all elements, including prominent thorns and a clear full moon.
- + Vibrant glowing effects and a sharp, clean digital art style.
- + Excellent rendering of gothic-style typography for the main title.
- − The aesthetic leans more toward modern vector illustration than 'vintage gothic'.
- − The scroll banner is slightly warped and the shadow placement feels less realistic.
- − The blue-toned background feels less like 'dark parchment' and more like a standard night sky.
Verdict: GPT Image 1 better captures the 'vintage gothic' and 'dark parchment' atmosphere requested in the prompt, looking like an authentic old poster. While Imagen 4.0 Ultra illustrates the specific elements like thorns and the moon more clearly, its clean digital style misses the gritty cinematic quality achieved by GPT Image 1.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
GPT Image 1
- + Excellent adherence to the 'cartoon scene' style with soft, rounded textures
- + Perfectly centered text and flag icon layout
- + Charming miniature diorama aesthetic with a consistent visual language
- − The sushi variety is limited to just two types
- − Diorama base color is a bit flat compared to the lighting on the sushi
Imagen 4.0 Ultra Generate 001
- + High diversity of sushi types including nigiri, maki, and ikura
- + Superior material rendering on the fish and roe showing better PBR qualities
- + More complex and visually interesting composition
- − The text layout is slightly off-center and the flag is placed to the side rather than centered as requested
- − The square plate appears to be floating awkwardly above the diorama base
Verdict: Both models followed the prompt well, but GPT Image 1 captured the '3D cartoon' and 'isometric miniature' aesthetic with more stylistic consistency and better text placement. Imagen 4.0 Ultra provided more detailed food models and better material realism, but failed to center the text/icon and had a less cohesive diorama base construction.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
GPT Image 1
- + Excellent photographic realism with natural depth of field.
- + Dynamic and natural-looking motion in the animals' poses.
- + Masterful lighting and atmospheric god rays that blend into the scene smoothly.
- − The fox kit has slightly anatomically questionable paws/legs during the run.
- − The kitten's facial expression is a bit distorted.
Imagen 4.0 Ultra Generate 001
- + Successfully includes all requested animal types clearly.
- + Vibrant colors and high-contrast details in the flowers and butterflies.
- + Includes distinct dew sparkles on the grass as requested.
- − The image looks more like a digital illustration than a hyper-photorealistic scene.
- − Animals appear stiff and 'planted' in the scene rather than interacting naturally.
- − Anatomical issues with the kitten's front paws being overly large and cartoonish.
Verdict: GPT Image 1 far exceeds the other in the 'hyper-photorealistic' requirement, creating a scene that feels like a real captured moment with beautiful lighting and natural fur textures. While Imagen 4.0 Ultra included more explicit details like dew sparkles and colorful butterflies, the overall execution feels like a composite digital painting rather than a masterpiece photograph. GPT Image 1 is the clear winner for its superior composition, atmosphere, and realism.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
GPT Image 1
- + Excellent typography with correct accentuation
- + Strong adherence to the banner and steaming cloche requirements
- + Features a nice grainy texture that evokes a vintage feel
- − Failed the light background instruction, producing a black background instead
- − The brown-on-black color scheme provides poor contrast
Imagen 4.0 Ultra Generate 001
- + Correctly followed the light background instruction
- + Elegant vector line-art style with sophisticated shading on the dome
- + Balanced use of cream and brown tones as requested
- − Missed the grainier 'subtle texture' visible in the other model
- − The typography is a bit thin compared to classic restaurant branding
Verdict: Imagen 4.0 Ultra Generate 001 is the clear winner as it followed all prompt instructions, including the light background and specific color palette. While GPT Image 1 produced excellent typography and a nice texture, its failure to use a light background makes it a poor fit for the requested vintage minimalist aesthetic.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
GPT Image 1
- + Excellent adherence to the color palette and flat-vector style
- + Mostly legible text with accurate historical names (Armstrong, Aldrin, Collins)
- + Clear, large-scale iconography that is easy to interpret
- − Typos in text ('EARLLUNAR')
- − The sequence/layout of the steps is confusing and doesn't follow a logical linear path
- − Missing 'Lunar Orbit' as a specific distinct step
Imagen 4.0 Ultra Generate 001
- + Strong composition with a clear, vertical infographic layout
- + Includes a proper title and subtitle as requested
- + Matches the NASA-inspired color theme perfectly
- − Text is almost entirely gibberish ('BEOMBERS', 'SPUSTUR', 'MEKERNLAIDING')
- − Icons are messy and don't clearly represent the specific mission steps requested
- − Small supporting text is illegible 'lorem ipsum' style scribbles
Verdict: GPT (Image 1) is the preferred choice because the primary text elements are legible and the icons are clean and recognizable, despite a few typos and a scattered layout. Imagen 4.0 Ultra (Image 2) creates a superior aesthetic composition for an infographic, but fails completely on the 'information' aspect with unintelligible text and ambiguous icons.
Explore each model
Google's Imagen 4.0 Ultra model offering the highest fidelity and resolution for professional-grade image generation