OpenAI's state-of-the-art image generation model with arbitrary resolution up to 4K and strong instruction following
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
GPT Image 2
#4 of 62 in Text-to-Image
Imagen 4.0 Generate 001
#55 of 62 in Text-to-Image
Where the votes landed
GPT Image 2
0%
win rate
Ties
0%
Imagen 4.0 Generate 001
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
GPT Image 2
- + Excellent adherence to the 'glass cube' description, depicting it as a hollow container.
- + Naturalistic lighting and physics with the sphere resting on the bottom surface.
- + High textural detail on the red book cover and wooden table.
- − The glass refractive properties where the plant meets the cube are slightly simplified.
Imagen 4.0 Generate 001
- + Sophisticated rendering of reflections and refractions on the glass surfaces.
- + Clean, modern aesthetic with sharp geometric lines.
- − Interpreted the cube as a solid block of glass, causing the sphere to appear to float unnaturally inside it.
- − The plant is less visible through the glass compared to Model A, making it feel more like a backdrop than an object behind the cube.
Verdict: GPT Image 2 followed the prompt's spatial logic more accurately by depicting a hollow glass cube where the sphere rests naturally on the base. While Imagen 4.0 Generate 001 features technically impressive reflections, it treats the cube as a solid optical block, resulting in a floating sphere that looks less grounded in reality. GPT Image 2 is the winner for its superior prompt adherence and realistic composition.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
GPT Image 2
- + Excellent 'imperfect framing' that feels like a real street photograph.
- + Very realistic skin texture and age-appropriate facial features for the subject.
- + Subtle but effective motion blur on the passing car in the background.
- − The bike's geometry is slightly warped near the seat post.
- − Physics of the toolbox and the man's seating (hovering) are a bit off.
Imagen 4.0 Generate 001
- + Strong bokeh and cinematic lighting with vibrant reflections on the pavement.
- + Clearer raindrops visible on the subject's clothing and the bike frame.
- + High level of detail in the bicycle's mechanical components.
- − The 'imperfect framing' prompt was ignored for a perfectly centered composition.
- − The tool in the man's hand is melded into his finger (anatomical artifact).
- − The skin texture looks a bit too smoothed and AI-typical compared to Model A.
Verdict: GPT Image 2 captures the 'candid' and 'imperfect framing' requirements much better than Imagen 4.0, resulting in an image that looks like authentic photojournalism. While Imagen 4.0 has more striking colors and rain effects, it feels significantly more staged and suffers from a noticeable anatomical error in the hands.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
GPT Image 2
- + Excellent photorealistic skin texture and lifelike eyes.
- + Masterful use of lighting and shallow depth of field for a cinematic feel.
- + Successfully captures a 'battle-worn' aesthetic with subtle dirt and weathered armor.
- − The beads in the hair are very small and barely noticeable.
- − The engraved patterns on the armor are somewhat faint and messy in some areas.
Imagen 4.0 Generate 001
- + Very distinct and clear engraving on the plate armor.
- + Prominent and creative inclusion of beads in the hair braids.
- + Strong implementation of the bokeh sparks and torchlight effect.
- − The skin and facial features look slightly plastic and CG-like compared to Model A.
- − The scars look like surface scratches rather than integrated skin variations.
- − The lighting on the face is a bit harsh and lacks the subtlety requested in 'faint' details.
Verdict: GPT Image 2 is the superior image due to its incredible photorealism and sophisticated lighting, which perfectly captures the 'battle-worn' and 'lifelike' requirements of the prompt. While Imagen 4.0 does a great job with the specific details like beads and engravings, it has a more artificial, digital-art appearance that lacks the gritty realism and depth found in GPT Image 2.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
GPT Image 2
- + Exceptional text rendering with perfect legibility.
- + Complete professional menu structure including realistic item descriptions and prices.
- + High-quality, appetizing food photography that looks commercially ready.
- − The layout is a bit dense compared to a strictly minimalist aesthetic.
Imagen 4.0 Generate 001
- + Stronger adherence to the 'grid' aspect of the prompt and minimalist geometric style.
- + Balanced use of vibrant accent colors as requested.
- − Text is largely nonsensical and illegible.
- − Food photos are generic and inconsistent in lighting.
- − The layout is confusing for a functional menu, with headers and items separated poorly.
Verdict: GPT Image 2 is significantly better as it produces a functional, professional-grade menu with perfect text and logical organization. While Imagen 4.0 Generate 001 attempts a more artistic grid layout, its inability to render legible text and logical item-to-image associations makes it unusable for the stated purpose.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
GPT Image 2
- + Excellent photorealistic textures on the beef patty, lettuce, and bun.
- + Perfect adherence to text styling with glowing, fiery effects on all required elements.
- + Dynamic and creative composition with sauce droplets and embers enhancing the sense of motion.
- − The composition is a bit crowded, with the text overlapping some background elements.
Imagen 4.0 Generate 001
- + Clean, vertical composition that makes the 'exploded' burger layers very easy to distinguish.
- + Accurate text placement and spelling for all three required components.
- + Good use of negative space to make the burger the central focus.
- − The 'fiery' text effect is much weaker compared to the other model, looking more like standard glows.
- − The burger components look slightly more like 3D renders than photorealistic food photography.
- − The price starburst uses a comma instead of a period, which may not be the standard decimal separator in all contexts.
Verdict: GPT Image 2 is the superior output as it captures the 'photorealistic' and 'fiery' aspects of the prompt with much more intensity and detail. While Imagen 4.0 provides a clean layout, GPT Image 2 manages to integrate the text and the energetic background into a cohesive, high-quality commercial advertisement.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
GPT Image 2
- + Excellent chalk texture that looks truly dusty and hand-drawn.
- + Perfect adherence to the prompted text with no spelling errors.
- + Realistic café environment and professional composition.
- − The slant in the handwriting is very minimal, though present.
Imagen 4.0 Generate 001
- + Successfully rendered most of the requested text.
- + Includes visible variations in slant as requested.
- − Included meta-commentary from the prompt such as 'Tittle Menu' and 'Footer' throughout the board.
- − The handwriting looks like a smooth digital font rather than gritty chalk.
- − Visible spelling errors like 'Berbs' instead of 'Herbs'.
Verdict: GPT Image 2 followed the prompt perfectly, producing a realistic and aesthetically pleasing chalkboard with authentic chalk textures. In contrast, Imagen 4.0 struggled with prompt leakage, writing instructions like 'Tittle Menu' and 'Footer' directly onto the board and using a font that lacked the requested chalk realism.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
GPT Image 2
- + Perfectly adhered to the complex spatial instruction of 'horse on top'
- + Hyper-realistic textures on the space suit and horse fur
- + Clever use of stirrups and reins attached to the astronaut
- − The lighting on the horse is slightly inconsistent with the moon surface orientation
- − The astronaut's hands have slightly irregular finger counts
Imagen 4.0 Generate 001
- + Visually beautiful cinematic lighting and color palette
- + High level of detail in the horse's mane and cosmic reflections
- − Failed the primary prompt instruction by placing the astronaut on top
- − A horse leg is anatomically incorrect and blends into the background mist
Verdict: GPT Image 2 is the clear winner as it successfully followed the difficult 'horse on top' spatial instruction, which requires overcoming strong training biases. Imagen 4.0 produced a generic, though beautiful, image that completely ignored the specific reversal requested in the prompt.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
GPT Image 2
- + Excellent photorealism with cinematic lighting and realistic textures
- + Superior composition with a natural candid feel
- + Accurate depiction of the passenger's bored, unfocused expression
- − The capybara's hands look slightly more human-like in their grip than paws
Imagen 4.0 Generate 001
- + Perfect adherence to specific details like 'both front paws on the steering wheel'
- + Clearly legible 'TAXI' text on the cap and car sign
- − Visuals look overly clean and 'CGI' rather than photorealistic
- − Composition is very symmetrical and less dynamic
- − Anatomical oddity with the capybara's claws being overly long and sharp
Verdict: GPT Image 2 provides a much more convincing photorealistic scene with lighting and textures that feel like a real night-time photograph. While Imagen 4.0 follows the paw placement instructions more literally, its final output lacks the gritty, cinematic atmosphere of a New York taxi and appears too digital/clean.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
GPT Image 2
- + Perfectly captures the requested 'vintage gothic' and 'parchment' aesthetic.
- + Excellent text rendering with no spelling errors in any of the required sections.
- + Cohesive composition with high-quality details like the NYC skyline and gothic architecture.
- − The image is quite busy with many overlapping textures.
- − The thorny border blends very heavily into the background elements.
Imagen 4.0 Generate 001
- + Clean, legible layout that follows all text requirements perfectly.
- + Strong color contrast between the blue sky and the glowing orange jack-o-lantern.
- + Good use of the thorny border and twisted tree elements.
- − Stylistically resembles a modern vector illustration rather than the requested 'vintage' or 'cinematic' look.
- − The scroll on the left side feels out of place and cuts off the border awkwardly.
- − Lacks the parchment texture requested in the prompt.
Verdict: GPT Image 2 is the clear winner for its superior interpretation of the 'vintage gothic' and 'dark parchment' theme. While Imagen 4.0 followed the text instructions accurately, its aesthetic is more of a clean digital cartoon, whereas GPT Image 2 delivered a detailed, atmospheric, and cinematic invitation that perfectly fits the intended mood.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
GPT Image 2
- + Perfect adherence to text requirements including 'JAPAN', 'SUSHI', and the flag icon.
- + Exemplary isometric 3D cartoon style with high-quality PBR textures on the food and wood.
- + Excellent composition on a square diorama base with a solid light blue background as requested.
- − Chopsticks are slightly fused together at the tips.
Imagen 4.0 Generate 001
- + Very realistic textures on the sushi ingredients, particularly the fish and roe.
- + Clean, minimalist composition.
- − Completely failed to include requested text and flag icon.
- − Background is a neutral grey/white instead of the requested solid light blue.
- − Style leans toward photographic realism rather than the requested 3D cartoon miniature look.
Verdict: GPT Image 2 followed every instruction in the prompt, including complex text rendering, specific background colors, and the 3D isometric style. In contrast, Imagen 4.0 Generate 001 produced a photorealistic image and ignored nearly all the specific layout and text instructions, making it a poor match for the prompt.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
GPT Image 2
- + Excellent adherence to the 'photorealistic' requirement with natural lighting.
- + Dynamic composition that conveys movement and playfulness well.
- + Superior texture rendering for the fur and grass.
- − One butterfly appears to have multiple sets of wings merged in an unrealistic way.
- − The fox's front right paw looks slightly disconnected from the body mechanics.
Imagen 4.0 Generate 001
- + Includes visible dew sparkles on the grass as specifically requested.
- + The 'tumbling' interaction between the dog and cat is very clear and charming.
- + Vibrant and diverse selection of wildflowers.
- − The lighting and fur texture feel more like a digital illustration than a hyper-photorealistic photo.
- − Proportions of the kitten compared to the puppy are slightly off.
- − The fox's anatomy looks a bit stiff and elongated.
Verdict: GPT Image 2 is the winner because it successfully captures the 'hyper-photorealistic' style requested, whereas Imagen 4.0 Generate 001 leans more toward a stylized digital painting or storybook illustration. GPT Image 2 also handles the depth of field and the golden hour lighting much more naturally, creating a more convincing 8K masterpiece.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
GPT Image 2
- + Sophisticated vintage engraving style that perfectly matches the 'classic' and 'retro' prompt keywords.
- + Excellent typography with high-quality rendering of the accent mark in 'Caffè'.
- + Detailed implementation of the banner and subtle paper texture.
- − The composition is more of a complex label/emblem than a minimalist logo.
- − The frame around the cloche makes the design quite busy.
Imagen 4.0 Generate 001
- + Accurately follows the 'minimalist' and 'vector emblem style' descriptors.
- + Clean, balanced layout with modern-retro aesthetic.
- + Simple and effective use of the warm brown and cream color palette.
- − The 'Est. 1720' banner is quite small and lacks the classic detail requested.
- − The steam is a bit stylized and cartoonish compared to the 'vintage' prompt.
- − Overall feels basic compared to the artistic depth of the other model.
Verdict: GPT Image 2 produces a richly detailed vintage emblem that captures the historical essence of a brand established in 1720, though it leans more into complexity than minimalism. Imagen 4.0 Generate 001 adheres better to the 'minimalist' instruction but loses the 'classic typography' and sophisticated vintage texture in the process. GPT Image 2 is the winner for its superior artistic execution and elegant handling of the text elements.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
GPT Image 2
- + Excellent adherence to all six requested steps of the mission.
- + High-quality, legible typography and authentic branding elements like the NASA logo and mission patch.
- + Strong composition with dedicated sections for crew and landing site.
- − The style is more detailed and illustrative than 'flat-vector', utilizing significant shadows and 3D effects.
Imagen 4.0 Generate 001
- + Captures the requested 'flat-vector' aesthetic perfectly with minimal gradients and clean lines.
- + Matches the specified NASA-inspired color palette accurately.
- − Fails to include all requested steps, stopping after step 4 (Lunar Orbit).
- − The layout is less structured and lacks the 'supporting details' like crew names requested in the prompt.
Verdict: GPT Image 2 is the superior infographic, as it successfully follows the sequential steps of the prompt and provides a complete educational layout including the crew and landing site. While Imagen 4.0 captures the requested 'flat' art style more accurately, it fails on prompt adherence by omitting several mission steps and supporting text.
Explore each model
Google's latest Imagen 4.0 text-to-image generation model with significantly better text rendering and overall image quality