OpenAI's state-of-the-art image generation model with better instruction following and adherence to prompts
Settled by community votes across 6 shared challenges, with an AI judge weighing in on each.
GPT Image 1.5
#7 of 62 in Text-to-Image
Imagen 3.0 Generate 002
#36 of 62 in Text-to-Image
Where the votes landed
GPT Image 1.5
0.0%
win rate
Ties
0.0%
Imagen 3.0 Generate 002
100.0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
GPT Image 1.5
- + Excellent text rendering with clear, legible titles and item descriptions
- + High-quality, appetizing food photography that looks professional
- + Coherent layout with logical sections for Appetizers, Pizza, and Mains
- − The grid is slightly asymmetrical, focusing more on a list-and-image layout than a strict photo grid
Imagen 3.0 Generate 002
- + Successfully creates a strict grid layout as requested
- + Clean, minimalist aesthetic with good use of white space
- − Text is largely illegible/gibberish (e.g., 'APPETIEES', 'MANS')
- − Sectioning is repetitive and confusing, with multiple 'Pizza' and 'Mains' labels scattered about
- − Pizza photos appear repeated or very similar, reducing variety
Verdict: GPT Image 1.5 is the clear winner because it produces a functional, professional-grade menu with perfectly legible English text and high-quality food photography. While Imagen 3.0 follows the 'grid' instruction more literally, its inability to render sensible text or a logical menu structure makes it unusable for the intended purpose.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
GPT Image 1.5
- + Excellent photorealistic texture on the meat and bun
- + Perfect text rendering for all requested strings
- + Highly dynamic composition with a strong sense of an explosion
- − The image is very busy, making the 'Magic Burger' title overlap slightly with the top bun
Imagen 3.0 Generate 002
- + Clean, professional layout suitable for a digital menu
- + Good lighting on the food components
- + Correct rendering of most text elements
- − The secondary text inside the starburst and above the price is gibberish
- − The food looks more like a 3D render than a photorealistic image
- − The 'exploded' effect feels static and neatly stacked rather than dynamic
Verdict: GPT Image 1.5 is the clear winner as it successfully rendered all three requested text phrases perfectly and delivered a much more realistic food texture. While Imagen 3.0 Generate 002 has a clean composition, it failed to produce coherent text inside the starburst and lacked the 'explosive' energy requested in the prompt.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
GPT Image 1.5
- + Excellent text accuracy with perfect spelling of all menu items.
- + Extremely realistic chalk texture and natural handwriting variations.
- + Follows all layout instructions precisely including the specific date and cursive title.
- − The bottom line of text is a bit small and faint compared to the rest.
Imagen 3.0 Generate 002
- + Clear and legible high-contrast text.
- + Includes a nice wooden frame that enhances the 'cozy café' atmosphere.
- − Numerous spelling errors including 'Risotttto', 'Octpsuin', and 'Choouip'.
- − Repeats menu items and footer text unnecessarily.
- − Text lacks the authentic chalk texture requested, looking more like a digital font outline.
Verdict: GPT Image 1.5 is the clear winner as it followed every instruction perfectly, including complex spelling and realistic texture. Imagen 3.0 Generate 002 struggled significantly with spelling accuracy and failed to provide the authentic chalk texture specified in the prompt, resulting in a cluttered and repetitive layout.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
GPT Image 1.5
- + Excellent gritty texture and cinematic lighting
- + Highly detailed background with lunar modules, planets, and asteroids
- + Dynamic sense of motion with the dust/debris kick-up
- − Failed the negative constraint: the astronaut is on top of the horse, not the horse on top of the astronaut
Imagen 3.0 Generate 002
- + Beautiful ethereal lighting on the white horse
- + Clean, balanced composition
- + High clarity and resolution
- − Failed the negative constraint: the astronaut is on top of the horse
- − Less surreal than Image A
Verdict: Both models failed the specific spatial logic requested in the prompt ('horse on top, not vice versa'), instead defaulting to the standard horse-riding-astronaut trope. GPT Image 1.5 is the preferred choice as it offers a more complex, cinematic composition with superior environmental details compared to the simpler backdrop of Imagen 3.0.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
GPT Image 1.5
- + Excellent PBR materials with realistic textures for wood, ceramics, and fish
- + Perfect text rendering and placement as requested
- + Detailed and professional composition with great lighting
- − Includes several extra elements like the teapot and soy sauce bottle not explicitly asked for
- − Leans more towards realism than the 'cartoon scene' requested
Imagen 3.0 Generate 002
- + Captures the '3D cartoon' and 'miniature' aesthetic perfectly
- + Strictly follows the 'minimal garnish and plate' instruction
- + Clean isometric perspective with soft, pleasant textures
- − Text is aligned to the top-left rather than top-center
- − Fish textures are slightly repetitive and look like plastic
- − Chopstick rest merges into the table geometry slightly
Verdict: GPT Image 1.5 produces a much higher quality render with realistic materials, though it adds unrequested elements and misses the stylized cartoon vibe. Imagen 3.0 captures the requested cartoon aesthetic and minimalism better, but fails on simple text placement instructions. GPT Image 1.5 is the preferred choice for its superior visual clarity and perfect typography.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
GPT Image 1.5
- + Excellent adherence to the 'tumbling together' action with dynamic poses.
- + Strong lighting effects with prominent god rays and golden hour ambiance.
- + High level of fur detail and expressive, sparkling eyes.
- − Physical anatomy is slightly compromised in the crowding, such as the kitten's paw structure.
- − The composition feels a bit cramped with all four animals squeezed into the frame.
Imagen 3.0 Generate 002
- + Natural and clear composition that allows each animal to be Seen distinctly.
- + Highly realistic fur textures and anatomically accurate features for each species.
- + Beautiful background depth and subtle dew effects on the grass.
- − The 'gold rays' are less distinct and more of a general glow compared to Model A.
- − The bunny is lying on its back in a slightly unnatural pose compared to its typical behavior.
Verdict: GPT Image 1.5 captures the 'joyful, chaotic energy' and specific lighting effects like god rays much better, making it feel more like a cohesive scene of play. However, Imagen 3.0 Generate 002 offers superior photographic clarity and better individual animal renders, even if the composition feels a bit more staged.
Explore each model
Google's Imagen 3.0 text-to-image generation model, producing high-quality images with improved detail and lighting