OpenAI's previous image generation model that accepts both text and image inputs and produces image outputs
Settled by community votes across 15 shared challenges, with an AI judge weighing in on each.
GPT Image 1
#28 of 62 in Text-to-Image
GPT Image 2
#3 of 62 in Text-to-Image
Where the votes landed
GPT Image 1
0.0%
win rate
Ties
0.0%
GPT Image 2
100.0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
GPT Image 1
- + Excellent refraction and lighting through the glass cube.
- + The blue sphere has a pleasing matte texture.
- + Clear adherence to the soft left-side lighting requested.
- − The glass cube has internal glass structural lines that make it look more like a frame than a solid cube.
- − The perspective of the cube base feels slightly tilted relative to the tabletop.
GPT Image 2
- + Natural glass transparency and realistic reflections on the cube surfaces.
- + High-quality texture on the red book cover.
- + Solid composition with realistic plant visibility through the glass.
- − The blue sphere's shadow on the bottom glass panel is slightly inconsistent with the light source.
- − The glass walls look a bit like double-paned glass rather than a single solid sheet.
Verdict: Both GPT Image 1 and GPT Image 2 follow the prompt near-perfectly, successfully rendering complex spatial relationships. GPT Image 2 is slightly superior due to the more realistic rendering of the red book's texture and a more coherent integration of the cube into the background scene.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
GPT Image 1
- + Excellent skin texture and moody photographic lighting.
- + Strong bokeh and rain effects that create a cinematic atmosphere.
- + Composition flows well with the subject’s posture and focal point.
- − The red bicycle frame has some physical inconsistencies and nonsensical parts.
- − The motion blur of the cars in the background looks more like static bokeh than motion.
GPT Image 2
- + Successfully captures the requested 'imperfect framing' and candid street style.
- + Includes realistic Japanese text and urban background elements for authenticity.
- + Better representation of 'motion blur' on the passing vehicles.
- − The lighting is a bit flat compared to the 'cinematic' request.
- − Skin texture on the subject is slightly smoother and less detailed than Model A.
Verdict: Model A produces a more aesthetically pleasing, cinematic image with superior skin detail, though the bicycle mechanics are slightly garbled. Model B better adheres to the 'candid street photo' and 'imperfect framing' requirements, capturing the motion blur and environmental details of a Japanese street more accurately.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
GPT Image 1
- + Excellent depiction of ornate engraved plate armor
- + Authentic battle-worn expression and grim skin textures
- + Dramatic use of warm torchlight and high-contrast shadows
- − Lighting is very dark, obscuring details on the right side
- − The braided hair beads are less distinct
GPT Image 2
- + Exceptional detail on leather straps and cloth underlayer as requested
- + Lifelike eyes and very clear skin textures with faint scars
- + Perfect execution of the braided hair with small beads
- − Lighting feels more like soft sunlight than localized torchlight
- − Overall composition is slightly less 'gritty' than the prompt implies
Verdict: GPT Image 2 is the better realization of the prompt, as it captures the specific technical requests like the leather straps, cloth underlayers, and distinct hair beads with superior clarity. While GPT Image 1 has a more intense atmospheric mood, its dark lighting obscures several requested details.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
GPT Image 1
- + Excellent food photography quality with high-resolution details
- + Clean, minimalist layout that feels modern and spacious
- − Nonsense filler text with frequent spelling errors like 'descridion'
- − Weak composition that cuts off the bottom half of the menu
GPT Image 2
- + Outstanding professional layout with logical sections for Appetizers, Pizza, and Mains
- + Extremely readable text with realistic item names and descriptions
- + Complete branding including a logo, social media icons, and dietary tags
- − Some minor food proportions look a bit small compared to the text blocks
Verdict: GPT Image 2 is the superior design because it functions as a complete, industry-standard menu with accurate spelling, distinct sections, and professional branding elements. While GPT Image 1 has high-quality food photography, its text is nonsensical and the layout is incomplete, cutting off abruptly at the bottom.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
GPT Image 1
- + Excellent centering and clean layout for a professional ad feel.
- + Accurate rendering of the requested text phrases.
- + High photorealism on the vegetable textures and burger bun.
- − Missed the first digit of the price, displaying '.99' instead of '6.99'.
- − The motion is quite static compared to the 'dynamic' request.
GPT Image 2
- + Perfect adherence to all text requirements including the specific '€6.99' price.
- + Highly dynamic composition with sauce splashes and tilted angles that convey motion.
- + Superior 'fiery' textures on the text and background elements.
- − The composition is slightly more cluttered than Model A.
- − Some sauce drips defy physics in a way that slightly lowers photorealism.
Verdict: GPT Image 2 is the superior choice as it fulfills all text requirements, including the specific price that GPT Image 1 missed. Furthermore, GPT Image 2 captures the 'dynamic' and 'motion' aspects of the prompt more effectively through the use of tilted angles and splashing sauce, whereas GPT Image 1 feels like a standard static explosion.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
GPT Image 1
- + Excellent text legibility across all sections
- + Consistent chalk texture throughout the board
- + Accurate rendering of specific menu prices and items
- − The handwriting style is too uniform and lacks natural cursive variation in the title
- − The layout is very tight to the edges, feeling less like a natural cafe scene
GPT Image 2
- + Successfully followed the requirement for elegant cursive for the title
- + Great atmospheric lighting and composition that feels like a real café setting
- + Captures natural variations in letter size and slant more effectively
- − The text is slightly less crisp and clear compared to Image A
- − Includes a decorative underline flourish not specifically requested
Verdict: Both models performed exceptionally well with complex text rendering. GPT Image 2 is the preferred result because it better adheres to the stylistic request for 'elegant cursive' for the title and provides a more realistic, natural handwriting variation, whereas GPT Image 1 is more uniform and looks slightly like a digital font despite the texture.
Pose & Character Mashup
Editing“Use Image 1 as the exact pose reference and Image 2 as the character reference. Recreate the person/character from Image 2 in the exact dynamic pose and body position from Image 1. Keep the exact face, hair, clothing style/details, and expression from Image 2. Match the lighting and environment of Image 1. The final image must show the character from Image 2 performing the precise action/pose from Image 1 with perfect anatomy and natural integration.”
AI Judge Analysis
GPT Image 1
- + Matches the yellow background and red ottoman from the source image
- + Correctly applies the black clothing and scarf from the second image
- − The facial likeness is poor and looks somewhat distorted
- − The hand anatomy on the left is severely deformed and simplified
- − Fails to capture the tilt of the head from the pose reference correctly
GPT Image 2
- + Excellent facial likeness and hair matching from the character reference
- + Perfectly replicates the complex cross-legged pose and arm angles from Image 1
- + Successfully integrates the scarf and sunglasses while maintaining natural-looking skin tones
- − The scarf's length and hanging position are slightly unrealistic given the extreme tilt of the body
Verdict: GPT Image 2 is the clear winner as it maintains an impressive facial likeness to the character reference while perfectly replicating the difficult anatomical pose from the source image. GPT Image 1 fails significantly on hand anatomy and provides a less accurate facial reconstruction. GPT Image 2 also does a better job of integrating the character's style, like the scarf and glasses, into the dynamic positioning.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
GPT Image 1
- + Excellent cinematic lighting and atmosphere
- + Realistic textures on the spacesuit and horse's coat
- + High visual quality with a balanced composition
- − Failed the primary prompt instruction of having the horse on top
GPT Image 2
- + Adhered perfectly to the unusual prompt instruction of having the horse on top
- + Correctly interpreted the surreal nature of the request
- + Impressive rendering of the NASA logo and suit details
- − The transition between the horse's legs and the astronaut is anatomically messy
- − The harness/bridle logic is a bit confused in the center
Verdict: While GPT Image 1 is a more aesthetically pleasing and 'cinematic' image, it completely failed to follow the specific spatial instruction of the prompt. GPT Image 2 successfully rendered the surreal concept of a horse riding an astronaut, making it the superior choice for prompt adherence despite some minor clipping artifacts.
Outfit Transfer Challenge
Editing“Use Image 1 as the base person. Dress them in the exact elaborate outfit from Image 2 (including all layers, accessories, jewelry, and shoes). Carefully adapt the clothing to the body shape and pose in Image 1 while maintaining realistic fabric behavior, correct proportions, and perfect lighting/shadow matching. Keep the person’s exact face, hair, and background completely unchanged.”
AI Judge Analysis
GPT Image 1
- + Excellent fabric texture and realistic lighting on the pea coat.
- + Successfully captures the watch and plaid pattern from the reference image.
- + Maintains a high level of facial detail and similarity to the original person.
- − The position of the hands in the pockets resulted in some slight anatomical awkwardness in the right arm.
- − Cropped the original image significantly, losing some of the background context.
GPT Image 2
- + Near-perfect preservation of the original background and framing.
- + Very accurate replication of the specific plaid scarf and gold watch from Image 2.
- + Maintains the exact vitiligo patterns on the face and the specific hair styling of the original subject.
- − The transition between the neck and the collar of the coat is slightly rough with some minor artifacts.
- − The left hand in the pocket has a slightly blurred, less defined appearance compared to the rest of the image.
Verdict: Both models performed exceptionally well, successfully transferring the complex layers (coat, scarf, shirt, jeans, watch) while preserving the identity of the person in Image 1. Model B is the winner because it maintained the original image's aspect ratio and framing while achieving a slightly better likeness of the original subject's unique features.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
GPT Image 1
- + Excellent fur texture and photorealism on the capybara.
- + The capybara's paws are positioned very naturally on the steering wheel.
- + Very clean rendering of the 'TAXI' text on the cap.
- − The composition feels slightly tighter, cutting off more of the car door and environment.
GPT Image 2
- + Good use of local NYC branding cues like the 'NYC' circle logo on the hat.
- + Dynamic angle that shows more of the car's exterior and the city environment.
- + Captures the 'bored' expression of the passenger very effectively.
- − The left paw (from our perspective) has a slight anatomical clipping issue with the steering wheel.
- − The texture of the capybara's snout is slightly more 'painterly' compared to the high realism of Model A.
Verdict: Both models followed the prompt exceptionally well, capturing the surreal humor of a professional capybara driver. GPT Image 1 is slightly preferred for its superior photorealistic textures and cleaner hand/paw anatomy, though GPT Image 2 offers a more well-balanced composition and better 'boring' facial acting from the human passenger.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
GPT Image 1
- + Excellent typography for the main title and banner scroll
- + Clean, cinematic lighting that draws focus to the jack-o-lantern
- − Confused the event details, merging the location into the 'Time' field and omitting the 7pm text
GPT Image 2
- + Includes all specific event details correctly (Date, Time, Location)
- + Rich, intricate border design that adheres perfectly to the thorns and webs prompt
- + Creative inclusion of 'The Arches' and NYC skyline in the background
- − The parchment texture is a bit busy, making some fine text slightly harder to read against the background
Verdict: GPT Image 2 is the superior choice because it followed all text-based instructions perfectly, whereas GPT Image 1 failed to include the time and glitched the layout of the location. GPT Image 2 also provided a much more detailed and thematic border that captured the 'thorns' and 'gothic' aesthetic more effectively.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
GPT Image 1
- + Excellent adherence to the 'minimal' garnish and plate request
- + Perfect execution of soft, 3D cartoon textures
- + Clean and readable typography and flag icon
- − Simple composition might feel sparse to some users
GPT Image 2
- + Detailed diorama base with stone textures and miniature pagoda
- + Professional 3D graphic design for the text header
- + Vibrant colors and high-quality material rendering on the seafood
- − Ignored the 'minimal garnish' instruction, filling the scene with many elements
- − The chopsticks are incorrectly rendered being partially embedded in the base
Verdict: GPT Image 1 followed the instructions for a minimal diorama more closely, resulting in a cleaner and more focused image. GPT Image 2 offers more visual complexity and better text styling, but it fails the specific prompt requirement for 'minimal garnish' and has minor clipping issues with the chopsticks.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
GPT Image 1
- + Excellent character interaction with the cat's paw reaching out toward the bunny.
- + Warm, cohesive lighting with very prominent god rays.
- + Clean, minimalist background that keeps focus on the animals.
- − The fox's front right leg has an anatomically awkward, dark, stump-like appearance.
- − The background butterfly in the top left is slightly less integrated into the lighting.
GPT Image 2
- + Dynamic composition with a more varied and colorful wildflower meadow.
- + Superior detail on the individual floral elements and grass blades.
- + The animals display more distinct action poses, particularly the pouncing fox.
- − The fox exhibits anatomical issues, with a strange tail/leg hybrid visible behind it.
- − The kitten's raised paw features an irregular number of toes/claws.
Verdict: Both models captured the joyful, golden-hour aesthetic requested, but GPT Image 2 (Model B) offers more vibrant colors and a richer environment with a wider variety of flowers. While both images suffer from minor anatomical glitches common in AI-generated animals—particularly with the fox—GPT Image 1 (Model A) has a slightly more focused and heartwarming composition.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
GPT Image 1
- + Clean vector-style icon suitable for modern branding
- + Accurate spelling of the name including the accent
- + Good balance between the cloche and the typography
- − Ignores the light background request, using a black background instead
- − Minimalist style is a bit too simple, lacking the requested vintage texture
GPT Image 2
- + Excellent adherence to the vintage aesthetic and light textured background
- + High-quality typography with period-accurate flourishes
- + Superior visual appeal with detailed hatching and framing
- − The 'E' in 'Caffè' is stylized in a way that slightly degrades legibility compared to Model A
Verdict: GPT Image 2 perfectly captures the requested 'vintage' and 'warm cream tones' with a high level of detail and artistic flair. While GPT Image 1 provides a functional vector logo, it fails the background color requirement and lacks the sophisticated texture visible in GPT Image 2.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
GPT Image 1
- + Strong minimalist flat-vector aesthetic
- + Follows the muted red and navy color palette effectively
- + Clean typography for crew names
- − Poor layout with disorganized text labels that don't align with icons
- − Spelling errors like 'EARLLUNAR'
- − Missing the Lunar Orbit step entirely
GPT Image 2
- + Perfect adherence to all 6 requested steps in a logical sequence
- + Excellent graphic design composition with clear sections and headers
- + Near-perfect text rendering and iconography
- − The detail level is slightly higher than 'flat-vector', leaning into more complex illustrations
- − Minor artifacts in the NASA meatball logo
Verdict: GPT Image 2 is significantly superior, as it successfully follows the structured 6-step infographic requirement with a logical, professional layout. GPT Image 1 fails the infographic test by scrambling the icons and labels, resulting in a confusing visual hierarchy and misspelled text.
Explore each model
OpenAI's state-of-the-art image generation model with arbitrary resolution up to 4K and strong instruction following