OpenAI's previous image generation model that accepts both text and image inputs and produces image outputs
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
GPT Image 1
#28 of 62 in Text-to-Image
Imagen 4.0 Fast Generate 001
#53 of 62 in Text-to-Image
Where the votes landed
GPT Image 1
0%
win rate
Ties
0%
Imagen 4.0 Fast Generate 001
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
GPT Image 1
- + Excellent adherence to lighting instructions with a soft, diffused window glow
- + Consistent glass thickness and realistic refraction of the plant behind it
- + Accurate spatial placement of all requested elements
- − The sphere appears slightly flat compared to the realism of the rest of the scene
Imagen 4.0 Fast Generate 001
- + Realistic reflective surfaces and a high-quality glass sphere texture
- + Strong contrast and Sharp details on the book spine
- + Good composition with a clear succulent plant behind the cube
- − The glass cube is missing its front-left vertical edge, making it structurally incoherent
- − The light is harsh and direct, rather than the requested soft window light
Verdict: GPT Image 1 followed all prompt instructions perfectly, including the specific lighting conditions and spatial relationships. While Imagen 4.0 Fast Generate 001 produced beautiful textures and materials, it failed on the cube's geometry and ignored the 'soft' aspect of the lighting request. GPT Image 1 is the clear winner for its superior structural coherence and adherence to the art direction.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
GPT Image 1
- + Excellent skin texture and realistic facial details
- + Perfectly captures the moody, cinematic aesthetic requested
- + High-quality rendering of rain droplets on surfaces
- − Missed the request for motion blur from passing cars
- − Framing feels slightly more posed than candid
Imagen 4.0 Fast Generate 001
- + Captures a more genuine 'candid' street photography composition
- + Beautiful reflections on the wet pavement
- + Innovative interpretation of 'imperfect framing' with a possible window border
- − Human anatomy issues, specifically the finger melting into the bike frame
- − Low-resolution appearance on the man's face and jacket
- − Failed to include motion blur for passing cars
Verdict: GPT Image 1 produces a far more technically proficient image with superior skin textures and realistic lighting, though it opted for a clean bokeh over the requested motion blur. Imagen 4.0 Fast Generate 001 captures the street photography 'snapshot' feel better, but is marred by significant anatomical artifacts in the hands and a lack of detail in the subject's face.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
GPT Image 1
- + Excellent adherence to the paladin and armor theme
- + Subtle and realistic scarred, battle-worn skin textures
- + Strong atmospheric lighting with beautiful bokeh sparks
- − The transition where the braid meets the armor is slightly unclear
Imagen 4.0 Fast Generate 001
- + Naturalistic foliage and lighting in the garden environment
- + Well-rendered leather jacket texture
- − Completely failed to follow the prompt instructions regarding armor, paladins, and warm torchlight
- − Full-body shot instead of the requested close portrait
Verdict: GPT Image 1 followed every aspect of the prompt, delivering a high-quality, atmospheric portrait of a paladin with intricate armor details. Imagen 4.0 Fast Generate 001 failed entirely, producing a modern elderly man in a garden which has no relation to the requested text.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
GPT Image 1
- + Excellent typography with clean, readable fonts
- + High-quality, appetizing food photography
- + Consistent layout and professional graphic design elements
- − Nonsense placeholder text for descriptions
- − Missed the 'Mains' section header entirely
Imagen 4.0 Fast Generate 001
- + Includes all three requested sections: Appetizers, Pizza, and Main Courses
- + Captures a grid-based layout effectively
- + Vibrant color accents used across the page
- − Significant spelling errors in headers (e.g., 'APETIERS')
- − Image-to-text alignment is messy and illogical
- − The food photography is repetitive, showing mostly pizzas in every section
Verdict: GPT Image 1 produces a high-fidelity, professional-looking design that closely resembles a real menu with beautiful product photography, despite some minor text artifacts. Imagen 4.0 Fast Generate 001 follows the structural prompt more closely by including the 'Main Courses' section, but fails significantly on text legibility and image variety.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
GPT Image 1
- + Excellent photorealistic texture on the patty and bun
- + All prompt text is present and correctly spelled
- + Stronger sense of internal lighting and heat within the ingredients
- − The price text is missing the '6', showing only '.99'
- − The fiery effect on the text is a bit blurry
Imagen 4.0 Fast Generate 001
- + Perfectly renders the full price as requested in the starburst
- + Very clean and sharp text rendering
- + The composition feels more like a professional advertisement with better spacing
- − The cheese has a slightly plasticky, less photorealistic appearance
- − A bit less 'fiery' glow on the bottom text compared to the top
Verdict: GPT Image 1 excels in the photorealistic rendering of the meat and bun, capturing a grittier, high-heat atmosphere. However, Imagen 4.0 Fast Generate 001 provides a more functional advertisement by correctly rendering the numerical price and maintaining very sharp, legible typography. Imagen 4.0 is the preferred choice for a commercial design due to its text accuracy and clean composition.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
GPT Image 1
- + Excellent text rendering with no spelling errors.
- + Very realistic grain and chalk texture on the individual letters.
- + Consistent handwriting style that truly looks like hand-lettering rather than a font.
- − Failed the requirement for the title to be in 'elegant cursive'.
- − The spacing in the 'Brown Butter Chocolate Chip Cookies' line is slightly uneven.
- − The lighting is a bit dark and muddy towards the bottom.
Imagen 4.0 Fast Generate 001
- + Clear, legible layout with a framing box around the header.
- + Correctly followed the partial prompt for 'Brown Butter...' by completing the cookie item.
- − Multiple spelling errors including 'Octuphus' and 'Cookes'.
- − Text looks like a digital handwritten font rather than authentic chalk strokes.
- − Duplicate entries for the final menu item with inconsistent pricing.
Verdict: GPT Image 1 is much more successful at capturing the specific 'chalk' aesthetic requested, featuring realistic texture and perfect spelling. While Imagen 4.0 Fast Generate 001 provides a cleaner layout, it suffers from several spelling mistakes and the text lacks the hand-drawn variations found in GPT Image 1.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
GPT Image 1
- + Excellent cinematic lighting and texture on both the horse and the space suit.
- + Dynamic composition with a sense of depth and scale against the planet.
- + High level of detail in the horse's anatomy and the worn look of the equipment.
- − The horse's front-right leg (viewer's left) has a slightly distorted hoof/ankle joint.
- − The reigns appear to be merging directly into the horse's neck in some places.
Imagen 4.0 Fast Generate 001
- + Clear facial detail visible through the astronaut's visor.
- + Vibrant galactic background colors that create a high-contrast surreal effect.
- + Solid anatomical muscularity in the horse.
- − The astronaut's left hand is positioned awkwardly on the horse's neck without reigns.
- − The lighting is somewhat flat and digital compared to the cinematic feel of Model A.
- − A visible seam/distortion exists where the astronaut's leg meets the horse.
Verdict: Both models followed the base prompt of an astronaut riding a horse, though neither followed the sub-instruction 'horse on top' as it contradicts 'horse riding astronaut' in most semantic interpretations. GPT Image 1 (Model A) is much more visually compelling due to its cinematic lighting and realistic textures, whereas Imagen 4.0 Fast (Model B) looks more like a standard digital composite.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
GPT Image 1
- + Excellent fur texture on the capybara
- + Captured the 'inside and looking through the window' perspective well
- + Atmospheric cinematic lighting
- − The capybara's paws look more like primate hands than capybara feet
- − The passenger is sitting in the front passenger seat rather than the back seat as requested
Imagen 4.0 Fast Generate 001
- + Successfully placed the passenger in the back seat
- + Higher level of detail on the taxi driver cap and uniform
- + Clearer depiction of the bored expression on the businesswoman
- − The lighting is a bit too bright and clean for a night scene
- − The capybara's paw structure is slightly surreal where it meets the steering wheel
Verdict: Imagen 4.0 Fast Generate followed the complex spatial instructions much better by correctly placing the businesswoman in the back seat, whereas GPT Image 1 placed her in the front. While GPT Image 1 has superior realistic texture on the animal, Imagen 4.0 Fast Generate creates a better narrative scene with clearer details on the outfit and more accurate passenger placement.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
GPT Image 1
- + Perfect text rendering for all requested strings
- + Strong adherence to the vintage gothic atmosphere
- + Sophisticated composition with elegant framing
- − The dark parchment texture makes the lower text slightly harder to read than the title
Imagen 4.0 Fast Generate 001
- + Good contrast and vibrant colors
- + Captures most illustrative elements like the banner and twisted trees
- − Several spelling errors including 'IINVITATION' and 'FNIGITS'
- − More cartoonish style rather than the requested polished cinematic look
- − Formatting error at the bottom with 'LOCATION.' overlapping with the next line
Verdict: GPT Image 1 is the clear winner as it flawlessly follows all text instructions, whereas Imagen 4.0 Fast Generate 001 struggles with spelling and layout. GPT Image 1 also better captures the 'vintage gothic' and 'cinematic' aesthetic requested in the prompt.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
GPT Image 1
- + Excellent typography rendering with perfect centering
- + High-quality soft clay-like textures and lighting
- + Strong composition that fills the square frame well
- − The diorama base is a bit large, making it look more like a block than a miniature platform
Imagen 4.0 Fast Generate 001
- + Accurate interpretation of the miniature diorama base
- + Clean isometric perspective
- + Good use of color contrast in the sushi ingredients
- − Small white borders on the left and right sides of the image
- − Flag icon is placed on the same line as 'SUSHI' rather than below as requested
- − Sushi textures are slightly less refined than Model A
Verdict: GPT Image 1 (Model A) is the clear winner as it followed all layout instructions perfectly, including the specific positioning of the text and flag. It also features superior material rendering, creating a cohesive 3D cartoon aesthetic, whereas Imagen 4.0 Fast Generate 001 (Model B) has minor framing issues and less sophisticated text placement.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
GPT Image 1
- + Excellent adherence to the 'playfully chasing' and 'tumbling' motion in the prompt.
- + Perfect representation of all requested species including a golden retriever and tabby kitten.
- + Beautiful lighting with distinct god rays and a dynamic, joyful composition.
- − The fox kit's back legs/tail area is slightly messy and indistinct.
- − The butterfly in the top left corner is a bit large compared to the others, breaking scale slightly.
Imagen 4.0 Fast Generate 001
- + High level of fur detail and clear facial features on the animals.
- + Sophisticated use of depth of field with the flowers in the foreground.
- − Failed to include butterflies which were a key part of the prompt.
- − The animals are sitting still, failing to capture the 'chasing' and 'tumbling' action requested.
- − Incorrectly generated a black cat and a brown/white dog instead of a tabby kitten and golden retriever.
Verdict: GPT Image 1 followed every detail of the prompt, successfully capturing the specific animal breeds and the energetic, playful atmosphere required. Imagen 4.0 Fast Generate 001 produced a high-quality static portrait, but failed on several key instructions including the dog breed, cat pattern, and the inclusion of butterflies.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
GPT Image 1
- + Excellent typography rendering with the correct accent mark.
- + Clean, minimalist composition that fits a modern vector logo style.
- + Highly accurate adherence to the banner and cloche elements.
- − Ignores the light background request, providing a black background instead.
- − The steam element is a bit thick and stylized compared to the rest of the logo.
Imagen 4.0 Fast Generate 001
- + Follows the 'light background' instruction with a subtle texture.
- + Sophisticated layout with balanced warm brown and cream tones.
- + Accurate text rendering for both the name and the established date.
- − Includes some small, nonsensical text elements ('AFED', 'CARO') above the main name.
- − The horizontal lines crossing through the text area slightly clutter the minimalist aesthetic.
Verdict: GPT Image 1 captures the vector emblem aesthetic perfectly and has cleaner typography, but it failed the background color requirement. Imagen 4.0 Fast Generate 001 followed the color and texture requirements better and provided a more professional layout, despite adding some minor hallucinated text.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
GPT Image 1
- + Excellent vector iconography for Earth and the descent stage.
- + Nearly perfect text rendering for historical names and mission stages.
- + Authentic NASA-inspired color palette and balanced layout.
- − One minor spelling error in 'EARLLUNAR'.
- − The logical flow of the labels doesn't perfectly match the spatial location of icons.
Imagen 4.0 Fast Generate 001
- + Strong composition that feels more like a technical flowchart/infographic.
- + Clean icons for the lunar module and crew silhouettes.
- + Adheres well to the requested flat-vector style with crisp lines.
- − Serious spelling issues including 'APOLO', 'SATURN VICON', 'MOOR', and 'MOO+ON'.
- − Repetitive use of icons that causes logical confusion in the infographic flow.
Verdict: GPT Image 1 is the superior choice because it produces legible, accurate text and high-quality individual icons that represent the Apollo mission steps clearly. While Imagen 4.0 Fast Generate 001 has a creative layout, it suffers from significant spelling errors and hallucinates prompt fragments (like 'Vicon') into the literal text of the image.
Explore each model
Google's Imagen 4.0 Fast model optimized for speed and efficiency, suitable for high-volume image generation tasks