OpenAI's state-of-the-art image generation model with arbitrary resolution up to 4K and strong instruction following
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
GPT Image 2
#3 of 62 in Text-to-Image
Imagen 4.0 Ultra Generate 001
#33 of 62 in Text-to-Image
Where the votes landed
GPT Image 2
0%
win rate
Ties
0%
Imagen 4.0 Ultra Generate 001
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
GPT Image 2
- + Excellent adherence to the 'partially visible through the glass' instruction.
- + Highly realistic texture on the red book and the glass framing.
- + Natural and soft window lighting that feels physically accurate.
- − The glass cube looks more like a hollow glass box with thick frames rather than a solid or thin-walled cube.
Imagen 4.0 Ultra Generate 001
- + Beautifully rendered light rays and shadows on the wooden table.
- + Clean, solid glass material with realistic refraction indices.
- + Detailed book spine with legible, high-quality text.
- − The blue sphere appears to be floating mid-air inside the cube without physical support.
- − The plant is entirely behind the cube and its visibility through the glass is less natural than in model A.
Verdict: GPT Image 2 followed the spatial prompt instructions more naturally, particularly in showing the plant through the glass and placing the sphere on the bottom surface. While Imagen 4.0 Ultra produced a more aesthetically cinematic image with better text rendering on the book, the floating sphere and less convincing transparency make it slightly less grounded in reality.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
GPT Image 2
- + Excellent depiction of motion blur from passing cars as requested.
- + Realistic candid framing with interesting foreground elements like the sign.
- + Very natural skin textures and believable lighting on the wet pavement.
- − The red bicycle frame looks a bit thin and structurally confusing near the seat post.
Imagen 4.0 Ultra Generate 001
- + Beautiful rain droplet effects on the character's jacket.
- + Strong cinematic composition with clear shallow depth of field.
- + High-quality skin texture and detail on the subject's hands.
- − Missed the 'motion blur from passing cars' requirement, as the car is static and sharp.
- − The bicycle is floating against the wall without a kickstand or leaning support.
Verdict: GPT Image 2 captured the 'candid street photo' aesthetic much better, specifically following the complex instruction for motion blur in the background cars. While Imagen 4.0 Ultra had better rain droplet details, GPT Image 2 felt more like a real, imperfect photograph taken with a 50mm lens as requested.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
GPT Image 2
- + Natural, cinematic lighting with a shallow depth of field.
- + Photorealistic skin textures and lifelike eyes.
- + Highly detailed and believable weathered plate armor.
- − The braids and beads are a bit subtle compared to the request.
- − The 'battle-worn' aspect is mostly dirt rather than visible scars.
Imagen 4.0 Ultra Generate 001
- + Excellent adherence to specific details like the beads in the hair and prominent scars.
- + Dynamic composition with visible sparks and a torch.
- + Clear, high-contrast engraving on the armor.
- − The character's skin and hair have a slightly plastic, CGI look.
- − The torch on the shoulder appears to be floating or awkwardly attached.
- − The symmetry and lighting feel less natural and more staged.
Verdict: GPT Image 2 produces a much more realistic and cinematic portrait with superior lighting and texture, though it misses the prominence of the requested beads. Imagen 4.0 Ultra Generate 001 follows the prompt's specific object list more closely (scars, beads, torch) but suffers from a less lifelike, more digital-art style that lacks the photographic depth seen in the other image.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
GPT Image 2
- + Exceptional text rendering with perfect spelling and coherent menu item descriptions.
- + Highly professional layout with sophisticated graphic design elements and branding.
- + Food photography is high-quality and consistent with the types of food described.
- − The font hierarchy is slightly dense in the descriptions, but still very legible.
Imagen 4.0 Ultra Generate 001
- + Successfully follows the grid layout request for food photos.
- + Features a clean white background as requested.
- − Text is comprised of gibberish and artifacts rather than readable English.
- − The layout is imbalanced with excessive empty space and illogical section groupings.
- − Food photos are of lower visual quality and lack consistent plate presentation.
Verdict: GPT Image 2 is significantly superior, producing a production-ready menu with perfect English text, clear pricing, and professional graphic design. In contrast, Imagen 4.0 Ultra fails to generate legible text and creates a sparse, unfinished layout that lacks the professional quality requested.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
GPT Image 2
- + Excellent photorealistic texture on the burger bun, meat, and vegetables.
- + Dynamic and messy 'exploded' effect with realistic sauce splashes and flying embers.
- + Strong adherence to the 'fiery glowing' text effect for all requested elements.
- − The composition feels a bit crowded with large text overlapping the sparks.
Imagen 4.0 Ultra Generate 001
- + Clean, symmetrical composition that works well for a traditional advertisement.
- + Accurate text rendering and placement of the price starburst.
- + Good lighting on the food components.
- − The food assets look slightly more artificial/plastic compared to the other model.
- − The 'exploded' effect is very static and lacks the requested sense of motion.
- − The background is a simple swirl rather than a detailed fiery environment with embers.
Verdict: GPT Image 2 captures the 'dynamic' and 'exploded' request much more effectively, with realistic sauce splashes and high-fidelity textures that make the food look appetizing. Imagen 4.0 Ultra Generate 001 provides a cleaner layout but fails to deliver the sense of motion and the intensity of the fiery background requested in the prompt.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
GPT Image 2
- + Excellent chalk texture throughout all characters
- + Perfect adherence to handwriting request with natural slant and variable line thickness
- + Realistic café background and lighting
- − Minor cursive connection error in the word 'TODAY'S'
Imagen 4.0 Ultra Generate 001
- + Perfectly legible and accurate text rendering
- + Good use of compositional wipes on the chalkboard
- + Clear, high-contrast text
- − Text looks like a digital handwritten-style font rather than organic chalk
- − Edges of letters are too smooth and lack grainy chalk texture
- − The 'cursive' request for the title was not fully met
Verdict: GPT Image 2 perfectly captures the requested 'chalk texture' and organic handwriting feel, making it look authentically hand-drawn. Conversely, Imagen 4.0 Ultra Generate 001 produces very clean text that looks like a digital font overlay, failing to capture the physical properties of chalk on a board.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
GPT Image 2
- + Perfectly follows the specific instruction of 'horse on top'
- + Highly detailed textures on the spacesuit and horse fur
- + Convincing cinematic lighting and lunar environment
- − The horse's front legs and harness anatomy are a bit nonsensical near the astronaut's shoulders
- − The astronaut's hands have slightly irregular finger counts
Imagen 4.0 Ultra Generate 001
- + Clean, illustrative style with vibrant colors
- + Creative use of space gear for the horse
- + Good composition with asteroids and ringed planet
- − Completely failed the negative constraint/specific instruction by putting the astronaut on top
- − Lacks the 'cinematic' realism requested
- − Perspective on the horse's legs is slightly distorted
Verdict: GPT Image 2 followed the difficult and specific instruction to place the horse on top of the astronaut, creating a truly surreal image. Imagen 4.0 Ultra provided a high-quality but generic interpretation that ignored the prompt's core structural constraint.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
GPT Image 2
- + Features a highly photorealistic texture on the capybara's fur and the surrounding taxi interior.
- + Captures the 'bored' expression of the passenger perfectly, contributing to the deadpan humor requested.
- + Atmospheric lighting through the window feels authentic to a rainy night in NYC.
- − The passenger is positioned slightly further away than ideal for a tight interior shot.
- − The capybara's hands are rendered somewhat like human fingers wrapped in fur rather than paws.
Imagen 4.0 Ultra Generate 001
- + Excellent composition that shows both the driver and the passenger clearly in one frame.
- + Very accurate adherence to the 'both paws on the wheel' instruction with distinct, sharp claw details.
- + Clarity and lighting are very strong, giving it a polished cinematic look.
- − The passenger looks slightly distressed or anxious rather than 'bored'.
- − The overall image has a slightly digital, CGI-like finish compared to the photorealism of the other image.
- − The capybara's head shape is a bit elongated/distorted.
Verdict: GPT Image 2 is the winner because it achieves a genuine sense of photorealism and nails the specific 'bored' emotion of the passenger, which is key to the prompt's humor. While Imagen 4.0 Ultra has a cleaner composition and better paw details, its overall style looks more like a high-end 3D render than a real photograph.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
GPT Image 2
- + Perfectly captures the 'vintage gothic' and 'dark parchment' aesthetic requested.
- + Excellent text rendering with no spelling errors across all three sections.
- + Highly detailed composition featuring specific NYC references like high-rise silhouettes and arches.
- − The dark color palette makes the border detail slightly harder to distinguish than model_b.
Imagen 4.0 Ultra Generate 001
- + Clean, legible layout with accurate text rendering.
- + Strong color contrast with vibrant orange and blue tones.
- + Faithfully includes all required elements like the scroll and thorn border.
- − Lacks the 'vintage' and 'dark parchment' feel, appearing more like a modern digital illustration.
- − Thorn border looks somewhat artificial and flat compared to the requested cinematic gothic style.
Verdict: GPT Image 2 is the clear winner as it masterfully balances the 'vintage gothic' parchment aesthetic with sharp, accurate text. While Imagen 4.0 Ultra produces a clean and colorful image, it feels more like a cartoon than the moody, cinematic poster requested in the prompt.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
GPT Image 2
- + Excellent 3D text rendering with shadow and depth
- + High richness of detail in the sushi textures and diorama base
- + Creative inclusion of a stone lantern and diverse nigiri types
- − The scene is a bit more crowded than the requested 'minimal garnish' design
Imagen 4.0 Ultra Generate 001
- + Perfect adherence to the 45-degree isometric camera angle
- + Clean, minimal aesthetic with a clear distinction between the base and plate
- + Soft, clay-like cartoon textures that match the prompt perfectly
- − The text layout is a bit scattered and lacks the 3D 'pop' requested for a cartoon scene
- − Composition is a bit floating and detached compared to a traditional diorama
Verdict: GPT Image 2 provides a much more visually engaging 3D cartoon effect with superior text treatment and material realism. While Imagen 4.0 Ultra Generate 001 captures the isometric perspective and minimalism accurately, its flat text and slightly floating plate make it feel less like a cohesive 3D miniature scene.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
GPT Image 2
- + Excellent photorealistic texture on the fur and grass.
- + Natural lighting with realistic 'god rays' and backlighting.
- + High degree of anatomical accuracy and dynamic posing.
- − The fox kit has black front legs that look slightly like solid sleeves/socks.
Imagen 4.0 Ultra Generate 001
- + Strong composition that captures the 'tumbling' and 'chasing' aspect of the prompt.
- + Clear depiction of dew sparkles on the flowers and grass.
- + Whimiscal, vibrant colors that fit a 'joyful' vibe.
- − The style leans more toward a high-end digital illustration or '3D render' than the requested 'hyper-photorealistic' scene.
- − Anatomical issues on the cat's front right paw and the fox's body connection.
Verdict: GPT Image 2 is the superior image because it successfully achieves the requested 'hyper-photorealistic' look with sophisticated lighting and realistic fur textures. Imagen 4.0 Ultra Generate 001 provides a charming and joyful scene, but it looks more like a stylized digital painting and suffers from minor anatomical artifacts on the kitten's paws.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
GPT Image 2
- + Excellent typography with a sophisticated vintage serif font.
- + Beautifully detailed etching style on the cloche and banner.
- + Strong composition with a professional frame that enhances the logo's impact.
- − The design is quite ornate, leaning more toward 'vintage' than the requested 'minimalist' style.
Imagen 4.0 Ultra Generate 001
- + Successfully captures the requested minimalist vector aesthetic.
- + Clean and legible layout with accurate text rendering.
- + Good adherence to the brown and cream color scheme.
- − The steam element is very simple and lacks the elegant 'vintage' feel found in the other model.
- − The typography is basic and less characteristic of a premium restaurant logo.
Verdict: GPT Image 2 provides a much more professional and high-quality vintage aesthetic with intricate detail and superior typography, though it is more decorative than minimalist. Imagen 4.0 Ultra follows the 'minimalist' constraint more closely, but the resulting design feels relatively plain and lacks the premium branding feel of the former. GPT Image 2 is preferred for its artistic execution and better fit for a classic institution like Caffè Florian.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
GPT Image 2
- + Excellent text rendering with accurate labels and astronaut names.
- + Perfect adherence to all six requested steps with appropriate icons for each.
- + Clean, professional composition that follows the NASA-inspired color palette perfectly.
- − The illustrations use slightly more 3D shading than the 'flat-vector' request.
- − The iconic EAGLE landing patch in the corner includes a stylized eagle that is slightly unrefined.
Imagen 4.0 Ultra Generate 001
- + Successfully captures a flat, minimalist vector aesthetic.
- + Adheres well to the requested color palette.
- − Fails significantly on text rendering, producing gibberish for most labels.
- − Does not follow the 6-step chronological sequence requested in the prompt.
- − The layout is cluttered and the iconography is abstract and confusing.
Verdict: GPT Image 2 is the clear winner as it perfectly follows the complex instructions, providing all six specific mission steps with high-quality icons and readable, accurate text. In contrast, Imagen 4.0 Ultra Generate 001 fails to follow the logical flow of the prompt and produces unintelligible placeholder text throughout the infographic.
Explore each model
Google's Imagen 4.0 Ultra model offering the highest fidelity and resolution for professional-grade image generation