OpenAI's previous image generation model that accepts both text and image inputs and produces image outputs
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
GPT Image 1
#32 of 62 in Text-to-Image
Imagen 4.0 Generate 001
#55 of 62 in Text-to-Image
Where the votes landed
GPT Image 1
0%
win rate
Ties
0%
Imagen 4.0 Generate 001
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
GPT Image 1
- + Excellent adherence to the 'behind' and 'partially visible' instruction for the plant.
- + Realistic glass textures with convincing refractions at the edges.
- + Accurate lighting and shadows consistent with the specified window source.
- − The blue sphere appears slightly too large to be described as 'small'.
- − The sphere appears to be floating without a visible support mechanism.
Imagen 4.0 Generate 001
- + High-quality textures on the red book, including realistic spine details.
- + Good use of reflections within the glass cube surfaces.
- + The sphere is floating centrally, which creates a clean, symmetrical composition.
- − The plant is completely obscured by the cube rather than being 'partially visible through' it.
- − The left face of the cube looks more like a mirror than transparent glass.
Verdict: GPT Image 1 followed the complex spatial instructions much better than Imagen 4.0, specifically the requirement for the plant to be visible through the glass cube. While Imagen 4.0 produced impressive textures on the book, it failed to render the transparency of the cube correctly, making it look like a mirror in some spots and completely blocking the background plant.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
GPT Image 1
- + Excellent skin texture and hyper-realistic facial details
- + Perfect representation of cinematic shallow depth of field
- + Natural, moody color grading that matches the 'light rain' setting
- − The bike anatomy is slightly warped near the rear derailleur
- − Background cars lack the requested 'motion blur' effect
Imagen 4.0 Generate 001
- + Strong environment details with bright reflections on wet pavement
- + Good sense of place with the yellow taxi in the background
- + Captures the activity of repairing more dynamically
- − Nonsensical text on the bike frame
- − Skin texture looks slightly processed and less natural than Model A
- − The tool held in the hand is blending into the bike's chain/frame
Verdict: GPT Image 1 captures the 'no stylization' and 'natural skin texture' requirements much better than the competing model, offering a truly realistic portrait. While Imagen 4.0 provides more vibrant urban atmosphere, it suffers from typical AI artifacts like garbled text on the bike and slightly waxy skin, making GPT Image 1 the superior choice for a photorealistic request.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
GPT Image 1
- + Exceptional photographic realism in skin texture and eyes
- + Sophisticated use of lighting and bokeh to create a cinematic atmosphere
- + Natural-looking hair braids with integrated beads
- − Very dark overall, obscuring some details in the armor and underlayers
- − Less focus on the leather straps mentioned in the prompt
Imagen 4.0 Generate 001
- + Clear visibility of ornate leather straps and buckles as requested
- + Stronger direct lighting from the torch source
- + Excellent engraving detail on the pauldrons
- − Skin texture appears somewhat plastic or CGI-like compared to Model A
- − The hair beads look like modern metallic caps rather than integrated beads
- − Visible artifact on the torch flame where it meets the edge
Verdict: GPT Image 1 produces a more lifelike and moody portrait with superior skin rendering and atmosphere, whereas Imagen 4.0 captures more of the technical prompt details like leather straps and specific armor textures. Ultimately, GPT Image 1 is the preferred choice for its higher visual quality and realistic interpretation of a battle-worn character.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
GPT Image 1
- + Excellent font legibility for the primary headers and prices
- + Food photography looks appetizing and consistent in lighting and style
- + Very clear categorization aligned with the prompt's requested sections
- − Placeholder text beneath titles is repetitive and misspelled ('Apperoiation descrigion')
- − The layout cuts off at the bottom, making it feel incomplete as a full menu design
Imagen 4.0 Generate 001
- + Stronger adherence to the 'grid' layout request with geometric design elements
- + Includes all three requested sections (Appetizers, Pizza, Mains) clearly within the frame
- + Better 'modern minimalist' aesthetic with creative use of color accents and whitespace
- − Text rendering is significantly worse with distorted and illegible characters
- − Prices are somewhat ambiguous in currency or scale compared to Model A
Verdict: GPT Image 1 produces a more functional and realistic menu with high-quality food photography and legible fonts, though it lacks the complete 'grid' feel. Imagen 4.0 Generate 001 provides a superior design layout and better adherence to the minimalist grid prompt, but fails on text clarity. GPT Image 1 is the likely winner for its professional finish and high-quality assets despite minor spelling errors.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
GPT Image 1
- + Perfect adherence to the glowing fiery text effect for all elements.
- + Excellent food photography aesthetic with high textural detail on the patty and bun.
- + Clean and impactful composition suitable for a professional advertisement.
- − Incorrect price rendering, showing '.99' instead of '6.99'.
- − The motion feels slightly more static compared to the diagonal explosion of the other model.
Imagen 4.0 Generate 001
- + Accurate text rendering for the price and all requested phrases.
- + Dynamic diagonal composition creates a stronger sense of mid-air motion.
- + Includes more varied ingredients like onions and pickles which add to the visual complexity.
- − The white text lacks the requested 'fiery, glowing effect' specified in the prompt.
- − The food textures appear slightly more like 3D renders than photorealistic photography.
- − The starburst effect looks a bit low-resolution compared to the rest of the image.
Verdict: GPT Image 1 captures the 'fiery' aesthetic much better, applying the requested glow effect to all text elements consistently, though it failed to render the full price correctly. Imagen 4.0 Generate 001 provides a more dynamic explosion and accurate text content, but missed the stylistic prompt for fiery text on the main title and secondary message. GPT Image 1 is preferred for its superior visual polish and adherence to the overall atmosphere, despite the minor price typo.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
GPT Image 1
- + Excellent text rendering with no spelling errors.
- + Realistic chalk texture and authentic handwritten variation.
- + Perfect adherence to the requested menu items and prices.
- − The 'elegant cursive' request for the title was interpreted as simple script rather than flowing cursive.
- − The lighting is a bit dark and monochromatic.
Imagen 4.0 Generate 001
- + Correctly included all requested menu items and prices.
- + Good framing with a wooden chalkboard border.
- − Included meta-commentary like 'Tittle Menu' and 'Footer' into the actual image text.
- − Significant text hallucinations and gibberish between menu items.
- − The handwriting looks more like a digital marker font than authentic chalk.
Verdict: GPT Image 1 followed the instructions exceptionally well, providing clean, accurate text with a convincing chalk texture. In contrast, Imagen 4.0 Generate 001 included the prompt's structural instructions (like 'Tittle' and 'Footer') as literal text on the board and suffered from significant gibberish text throughout the composition.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
GPT Image 1
- + Excellent high-detail rendering of the horse and astronaut textures
- + Cinematic lighting and composition
- − Completely failed the semantic inversion in the prompt, placing the astronaut on top
- − Lacks the requested surrealism beyond the basic concept
Imagen 4.0 Generate 001
- + Vibrant, surreal color palette with galaxy textures on the horse
- + High visual quality and clear details
- − Failed the negative constraint to have the horse on top of the astronaut
- − Generative artifacts present in the extra glowing lines and fuzzy hoof contact
Verdict: Both GPT Image 1 and Imagen 4.0 failed the specific logic test in the prompt, which requested a 'horse on top' of the astronaut. However, GPT Image 1 is slightly preferred for its superior cinematic rendering and realistic textures, despite both models defaulting to the standard 'astronaut on horse' trope.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
GPT Image 1
- + Excellent photorealism and cinematic lighting
- + Natural-looking textures on the capybara's fur and the jacket
- + Perfectly captures the 'bored' expression of the businesswoman
- − The paws are merged with the steering wheel in a physically impossible way
- − The composition is a bit tight, cutting off some of the taxi environment
Imagen 4.0 Generate 001
- + Comprehensive layout showing more of the car interior and Manhattan street
- + Good interpretation of the 'taxi driver cap' with specific TAXI branding
- + Clearer separation of the front and back seat areas
- − The passenger's phone is unnaturally small and the hands are slightly warped
- − The capybara's paws have sharp, protruding talons that look more like a monster than a capybara
- − Lower level of photorealism compared to Model A, appearing more digital
Verdict: GPT Image 1 succeeds in terms of photorealism, textures, and lighting, creating a scene that looks like a high-quality film still. While Imagen 4.0 provides a better wide-angle composition and more background detail, it suffers from anatomical errors in the hands and paws and a slightly less realistic 'digital' finish. GPT Image 1 is the winner for its superior atmospheric quality and more accurate capybara features despite the minor steering wheel artifact.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
GPT Image 1
- + Perfectly captures the requested 'vintage gothic' style with a muted, classic aesthetic.
- + Excellent composition with a very clear and legible font hierarchy.
- + The parchment texture and framing feel authentic to a real physical invitation.
- − The lighting on the pumpkin is a bit flat compared to the surrounding environment.
- − The border thorns are less pronounced than in the other model.
Imagen 4.0 Generate 001
- + Strong use of 'cinematic lighting' with a glowing blue atmosphere and bright jack-o-lantern.
- + Includes all requested textual elements with creative gothic typography.
- + The thorns and webs in the border are highly detailed and dynamic.
- − The overall style leans more toward a modern digital illustration/cartoon than the requested 'vintage' parchment look.
- − The large vertical scroll element on the left feels unbalanced and out of place.
- − The 'You are invited' banner text uses a generic sans-serif font that clashes with the gothic theme.
Verdict: GPT Image 1 is the superior choice as it accurately interprets 'vintage gothic' and creates a cohesive, professional-looking invitation with excellent text layout. While Imagen 4.0 has more vibrant cinematic lighting and intricate border details, its aesthetic feels too much like a modern mobile game asset, and the font choice on the scroll banner is jarring.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
GPT Image 1
- + Perfectly follows all text layout instructions including specific words and flag icon position.
- + Highly consistent 3D cartoon art style with soft, clean textures as requested.
- + Excellent adherence to the color palette, specifically the solid light blue background.
- − The 3D clay-like textures, while 'soft', lean more toward stylization than realistic PBR materials.
- − The rice grains are somewhat oversized for a 'miniature' look.
Imagen 4.0 Generate 001
- + Impressive realistic textures and PBR shaders, especially on the salmon roe and tuna.
- + Captures a genuine miniature diorama aesthetic with realistic lighting and shadows.
- − Completely failed to include the requested text ('JAPAN', 'SUSHI') and flag icon.
- − Failed to provide the requested solid light blue background.
- − The 45-degree isometric projection is less precise than Model A's composition.
Verdict: GPT Image 1 followed every instruction in the prompt, including the complex text and iconography requirements, maintaining a clean 3D cartoon style. Imagen 4.0 provided much higher fidelity in terms of 'realistic PBR materials', but it failed to include the text, the flag, and the specific background color requested. GPT Image 1 is the clear winner for its superior prompt adherence and overall composition.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
GPT Image 1
- + Excellent adherence to the 'god rays' and warm golden lighting request.
- + Achieves a high level of photorealistic detail in the fur texture.
- + Natural dynamic composition that feels like a captured action shot.
- − The fox kit has a slightly distorted front paw.
- − The kitten's pose and anatomy look a bit awkward where it meets the meadow.
Imagen 4.0 Generate 001
- + Successfully includes all requested animals clearly in a 'tumbling' interaction.
- + High level of detail in the dew drops on the flowers.
- + Vibrant and colorful interpretation of the wildflower meadow.
- − Has a more digital painting/illustration style rather than 'hyper-photorealistic'.
- − Lighting feels flat and lacks the atmospheric depth of god rays present in the other image.
- − The kitten's anatomy, specifically its midsection and legs, is confusingly rendered.
Verdict: GPT Image 1 far better captures the 'hyper-photorealistic' and atmospheric lighting requirements of the prompt, creating a believable and cinematic scene. While Imagen 4.0 Generate 001 offers a colorful and charming composition, its aesthetic leans toward a stylized digital illustration rather than a realistic photograph.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
GPT Image 1
- + Excellent typography with correct accents and vintage serif styling
- + Strong texture that gives a tangible letterpress or stamped feel
- + Accurate interpretation of a vintage banner and cloche
- − Failed the prompt's request for a light background by using solid black
- − Monochromatic approach ignores the cream part of the tone request
Imagen 4.0 Generate 001
- + Perfect adherence to the light background and warm brown/cream color palette
- + Clean, minimalist vector aesthetic that is very professional
- + Accurate spelling and correct inclusion of the banner and steam
- − The 'E' in 'EST' is slightly taller than the other letters in the banner
- − Typography is less 'classic' and more modern/clean compared to Model A
Verdict: Imagen 4.0 Generate 001 followed the prompt instructions much more closely, particularly regarding the light background and the color scheme. While GPT Image 1 has superior vintage typography and texture, it failed on the most basic background color instruction, whereas Imagen 4.0 successfully balanced the minimalist style with a warm, professional palette.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
GPT Image 1
- + Excellent adherence to the requested NASA-inspired color palette.
- + Includes all astronauts and specific lunar landing site details.
- + Clean, minimalist flat-vector aesthetic that feels like a cohesive poster.
- − Several spelling errors in the labels (e.g., 'EARLLUNAR').
- − The layout of the steps is somewhat disjointed and non-linear.
Imagen 4.0 Generate 001
- + Near-perfect text rendering with zero spelling errors.
- + Logical diagrammatic flow showing the trajectory from Earth to Moon.
- + Crisp vector graphics with professional iconography.
- − Fails to include all requested steps, missing 'Descent' and 'Landing'.
- − The color palette uses vibrant greens and oranges that deviate from the 'muted' NASA-inspired request.
Verdict: GPT Image 1 captures the artistic spirit and content requirements much better, including specific astronauts and the landing site, though it suffers from typical AI spelling issues. Imagen 4.0 has superior technical execution in terms of text clarity and diagram logic, but it failed to follow the full list of steps and ignored the muted color palette instructions.
Explore each model
Google's latest Imagen 4.0 text-to-image generation model with significantly better text rendering and overall image quality