OpenAI's previous image generation model that accepts both text and image inputs and produces image outputs
Settled by community votes across 6 shared challenges, with an AI judge weighing in on each.
GPT Image 1
#32 of 62 in Text-to-Image
Imagen 3.0 Generate 002
#36 of 62 in Text-to-Image
Where the votes landed
GPT Image 1
0%
win rate
Ties
0%
Imagen 3.0 Generate 002
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
GPT Image 1
- + Excellent photographic quality and food lighting
- + Highly legible typography with bold sans-serif fonts
- + Extremely clean and professional minimalist layout
- − Text contains minor spelling errors like 'descrigion'
- − The grid is a bit tight with the images taking up most of the space
Imagen 3.0 Generate 002
- + Successfully implements a complex 4x4 grid layout
- + Good use of negative space for a professional feel
- + Matches the intended 'casual dining' vibe with varied food photos
- − Text is mostly gibberish or illegible
- − Food photography is less detailed and realistic than Image A
- − Layout hierarchy is confusing with section headers overlapping in blocks
Verdict: GPT Image 1 is the superior choice because it functions as a usable menu design with clear typography and high-end photography. While Imagen 3.0 Generate 002 creates an interesting grid, the text legibility is very poor and the food images across the grid lack the appetizing clarity found in GPT Image 1.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
GPT Image 1
- + Excellent typography with a consistent fiery, glowing effect as requested.
- + Very high level of photorealistic detail on the meat patty and bun textures.
- + Accurate rendering of the requested text phrases.
- − The price in the starburst is missing the '6', reading as '.99'.
- − The composition is a bit tight at the top and bottom edges.
Imagen 3.0 Generate 002
- + Dynamic composition with a strong sense of explosion and flying ingredients.
- + Good inclusion of extra elements like the sauce and onion rings for visual interest.
- + The starburst price tag accurately includes the full '€6.99' text.
- − The text doesn't strictly follow the 'fiery, glowing' effect, looking more like gold metal.
- − Gibberish text present inside the starburst badge.
- − The burger components look slightly more like a 3D render than a photorealistic advertisement.
Verdict: GPT Image 1 followed the stylistic instructions for the text much better, creating a cohesive fiery glow across all typography. While Imagen 3.0 Generate 002 had a more energetic central 'explosion' and correctly rendered the full price, its use of gibberish text in the badge and the lack of a glowing effect on the main title makes GPT Image 1 the more professional advertisement overall.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
GPT Image 1
- + Excellent text rendering with no spelling errors.
- + Perfectly follows the request for matching chalk handwriting style.
- + High realism with subtle variations in chalk density and texture.
- − The 'elegant cursive' request was interpreted more as a stylised print than true cursive join-up.
- − Composition is a bit tightly cropped against the edge of the board.
Imagen 3.0 Generate 002
- + Successfully captures a cursive script style for the menu items.
- + Composition includes a realistic wooden frame and context.
- + Realistic chalk dust artifacts and erasing marks on the board.
- − Significant spelling errors throughout the menu items such as 'Risottto', 'Octpsuin', and 'Choouip'.
- − The text uses a hollow/outline effect for the title which was not requested.
- − Inconsistent font styles and sizes that make the layout look cluttered.
Verdict: GPT Image 1 is the clear winner as it successfully rendered all requested text with perfect spelling and a highly consistent chalk texture. While Imagen 3.0 Generate 002 provided a better sense of environmental context with the board frame, it failed significantly on text accuracy and legibility with numerous spelling hallucinations.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
GPT Image 1
- + Excellent realization of the cinematic lighting style
- + Strong anatomy and fine details on the horse's coat and muscle structure
- + Better sense of scale with a planet visible in the background
- − The astronaut's hand lacks clarity in how it holds the reins
Imagen 3.0 Generate 002
- + Beautiful use of color and nebula effects in the background
- + Vibrant and sharp details on the astronaut's suit architecture
- + Horse's mane has a graceful flowing motion
- − The horse's legs feature some anatomical distortion near the hooves
- − Slightly less 'cinematic' and more 'digital art' feel compared to the lighting in Model A
Verdict: Both models successfully followed the prompt, placing the astronaut on top of the horse in space. GPT Image 1 (Model A) is the winner due to its superior cinematic lighting and more believable animal anatomy, whereas Imagen 3.0 (Model B) has slight structural issues with the horse's lower legs despite having a more colorful background.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
GPT Image 1
- + Excellent center-aligned typography with a clear flag icon
- + Higher material quality with subsurface scattering effects on the fish
- + Polished 3D render look with soft, professional shadows
- − The chopsticks are placed awkwardly on the base rather than a resting block
- − Simple composition with fewer sushi varieties
Imagen 3.0 Generate 002
- + True isometric 45° perspective as requested
- + Includes a chopstick rest and more diverse sushi varieties
- + The diorama base has more architectural character
- − Text and flag are top-left aligned instead of top-center
- − The material textures look flatter and less realistic compared to Model A
- − Text is slightly less bold
Verdict: GPT Image 1 followed the layout instructions for the text better, placing it top-center, and achieved a much more refined '3D cartoon' aesthetic with superior lighting and textures. While Imagen 3.0 was more accurate regarding the specific 45-degree isometric perspective and included more sushi varieties, its text placement failed the prompt and its visual quality appears more like a basic mobile game asset.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
GPT Image 1
- + Excellent sense of motion and playfulness
- + Superior 'god rays' lighting that feels more cinematic
- + Clean anatomy with clear, expressive facial features for all four animals
- − The fox looks a bit more like a second puppy/dog hybrid than a distinct fox kit
Imagen 3.0 Generate 002
- + Excellent fur textures and dew sparkles in the grass
- + Includes the tumbling action requested in the prompt
- + High detail in the wildflower variety
- − The bunny has a very strange, almost cat-like face while on its back
- − Anatomical issues with the legs of the animal on its back
- − The golden retriever's head shape is slightly asymmetrical
Verdict: GPT Image 1 captures the joyful, active spirit of the prompt more effectively with a dynamic composition and beautiful lighting. While Imagen 3.0 provides excellent texture and environment details, it suffers from significant anatomical distortions in the 'tumbling' animals, particularly the rabbit's face and legs.
Explore each model
Google's Imagen 3.0 text-to-image generation model, producing high-quality images with improved detail and lighting