OpenAI's previous image generation model that accepts both text and image inputs and produces image outputs
Settled by community votes across 20 shared challenges, with an AI judge weighing in on each.
GPT Image 1
#28 of 62 in Text-to-Image
GPT Image 1.5
#7 of 62 in Text-to-Image
Where the votes landed
GPT Image 1
100.0%
win rate
Ties
0.0%
GPT Image 1.5
0.0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
GPT Image 1
- + Excellent material rendering with realistic glass thickness and refraction
- + Clean and high-resolution textures on the sphere and book
- + Precisely following the spatial prompt regarding the plant being behind the cube
- − The sphere appears to be floating slightly above the bottom surface of the cube
GPT Image 1.5
- + Beautiful lighting and reflections on the sphere's surface
- + Realistic integration of the sphere sitting on the base including a reflection
- + Good adherence to the requested layout
- − The glass cube is missing its top plane; the book appears to be resting directly on the side walls
- − Lower resolution/clarity in the plant leaves compared to Image A
Verdict: Both models followed the complex spatial instructions well. GPT Image 1 is the winner because it successfully renders a complete six-sided glass cube with visible top-side refraction under the book, whereas GPT Image 1.5 failed to render the top glass panel. GPT Image 1 also features much sharper details on the plant and lighting.
Man and Car in California
Editing“Make a photo of the man driving the car down the California coastline”
AI Judge Analysis
GPT Image 1
- + Excellent preservation of the car's design including the grill and badge.
- + Successful integration of the man's likeness and hairstyle into the driver seat.
- + Great sense of motion and environment with realistic lighting and motion blur on the wheels.
- − The man appears to be looking down rather than at the road.
GPT Image 1.5
- + Natural facial expression and pose for a driving scenario.
- + The background and scenery are very convincing for a coastal drive.
- + Good preservation of the man's facial features and distinctive hair.
- − Crop is a bit tight, losing the front of the car which was a focus of the source image.
- − The hand on the steering wheel has some anatomical distortion.
Verdict: Both models did an excellent job of combining two separate source images into a single cohesive scene. GPT Image 1 (Model A) is the winner because it preserves the full subject (the car) more effectively and captures a cinematic sense of movement, whereas GPT Image 1.5 (Model B) crops the car heavily and has slight issues with hand anatomy.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
GPT Image 1
- + Excellent skin texture and hyper-realistic facial details
- + Strong '50mm lens' look with soft bokeh
- + Great atmospheric handling of rain droplets on the bike and clothing
- − Anatomy and perspective issues where the man's hand meets the bicycle chain/gears
- − The bicycle design is overly simplified and structurally illogical in the rear frame area
GPT Image 1.5
- + More complex and realistic bicycle mechanics including a chain, gears, and a tool tray
- + Better storytelling with the inclusion of the tool tray and secondary umbrellas in the background
- + Dynamic composition with visible rain streaks
- − The facial texture is slightly smoother and less detailed compared to Model A
- − The car in the background looks a bit static despite the prompt for motion blur
Verdict: GPT Image 1 (Model A) delivers a more intimate and texturally detailed portrait with superior skin rendering, but the bicycle's anatomy is flawed. GPT Image 1.5 (Model B) provides a more coherent scene with realistic mechanical details and better environmental storytelling, despite having slightly less impressive skin texture. Model B is the preferred choice for its logical structural integrity and effective use of the surrounding environment.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
GPT Image 1
- + Exceptional engraved metal texture with realistic patina
- + Mature and weary facial expression fits the 'battle-worn' prompt perfectly
- + Natural skin texture and subtle color grading
- − Lighting is somewhat flat compared to Model B
- − Fails to show leather straps requested in the prompt
GPT Image 1.5
- + Includes all elements like beads, leather straps, and cross iconography
- + Dynamic lighting with strong torchlight reflections and bokeh sparks
- + High contrast and sharp details on the hair and fabric
- − Skin texture appears slightly airbrushed or overly smooth under the dirt
- − Scars look a bit tropically placed rather than natural wear
Verdict: GPT Image 1.5 is the winner as it adheres more strictly to the specific details of the prompt, including the leather straps and the silver beads in the hair. While GPT Image 1 offers a more compellingly 'weary' character face, GPT Image 1.5 captures the cinematic lighting and ornamental complexity that better matches the 'ornate' and 'warm torchlight' descriptions.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
GPT Image 1
- + Excellent photographic quality and lighting of food items
- + Bold and clean typography that adheres well to a modern minimalist style
- − Nonsensical placeholder text with frequent typos like 'Apperoiation descrigion'
- − Missing the 'Mains' category header despite showing mains in the grid
GPT Image 1.5
- + Includes all requested text categories: appetizers, pizza, and mains
- + Legible, realistic menu item names and descriptions
- + Professional grid-based layout that feels like a real commercial menu
- − Image resolution/sharpness is slightly lower than Model A
- − Text alignment on the bottom section is slightly cramped
Verdict: GPT Image 1.5 (Image B) is the clear winner for its functional design, providing logical categories, realistic food names, and a professional layout that fits a casual dining context. While GPT Image 1 (Image A) has higher fidelity food photography, its text is nonsensical and it fails to include all the requested sections.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
GPT Image 1
- + Excellent typography rendering with clean, glowing neon effects.
- + Very clean composition with minimal clutter, making the food items stand out.
- + Good photorealistic texture on the bun and patty.
- − Failed to include the price correctly, displaying '.99' instead of '6.99'.
- − Less 'explosive' energy compared to the other image, feeling more static.
GPT Image 1.5
- + Successfully rendered all text prompts including the correct €6.99 price.
- + Highly dynamic composition with a strong sense of motion and chaotic energy.
- + Rich detail with added ingredients like onions and cascading sauce drips.
- − The background is a bit busy, which slightly distracts from the central product.
- − The text 'MAGIC BURGER' is partially cropped at the very top of the frame.
Verdict: GPT Image 1.5 is the clear winner because it followed all text-related instructions perfectly, including the specific price of €6.99 which GPT Image 1 failed to render completely. Furthermore, GPT Image 1.5 captured the 'dynamic' and 'exploded' feel much more effectively with its chaotic embers and complex layering of ingredients.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
GPT Image 1
- + Excellent text legibility throughout the entire menu.
- + Perfect spelling and numerical accuracy for all requested items.
- + Uniform chalk texture that feels consistent across the entire board.
- − The handwriting style is a bit too uniform, bordering on a digital font look.
- − Fails the request for 'elegant cursive' for the title, using a simple print instead.
GPT Image 1.5
- + Successfully uses elegant cursive for the title as requested.
- + Handwriting has a more authentic, variable feel with realistic chalk dust and smudging on the board.
- + Captures the 'cozy café' aesthetic more effectively through board texture.
- − The cursive handwriting is much harder to read compared to Model A.
- − The 'Grilled Octopus' line has a slight visual smudge that makes the text look a bit messy.
Verdict: Model B (GPT Image 1.5) followed the prompt instructions more closely by providing the requested 'elegant cursive' for the title and a more natural, variable handwriting style. While Model A (GPT Image 1) is more legible, it failed the specific stylistic requirements of the prompt and has a more mechanical appearance.
Pose & Character Mashup
Editing“Use Image 1 as the exact pose reference and Image 2 as the character reference. Recreate the person/character from Image 2 in the exact dynamic pose and body position from Image 1. Keep the exact face, hair, clothing style/details, and expression from Image 2. Match the lighting and environment of Image 1. The final image must show the character from Image 2 performing the precise action/pose from Image 1 with perfect anatomy and natural integration.”
AI Judge Analysis
GPT Image 1
- + Successfully replicates the complex leg-crossing pose from Image 1.
- + Accurately represents the scarf and glasses from Image 2.
- + Matches the yellow background and lighting of the source image.
- − The face has significant distortions and does not closely resemble the character in Image 2.
- − Notable anatomical issues with the feet and hands.
GPT Image 1.5
- + Excellent character resemblance, keeping the face and hair almost identical to Image 2.
- + Accurately replicates the pose while maintaining better anatomical proportions.
- + Includes clothing details like the text on the sweatshirt from the character reference.
- − The scarf is slightly merged with the sweater fabric in some areas.
- − One foot has an extra toe/misshapen structure.
Verdict: GPT Image 1.5 is the clear winner as it maintains a much higher fidelity to the character's facial features and clothing details from Image 2 while successfully Maping them onto the pose from Image 1. GPT Image 1 fails to preserve the character's identity, resulting in a distorted face that looks very little like the reference provided.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
GPT Image 1
- + Successfully follows the surreal instruction of a horse riding an astronaut.
- + Clean and cinematic lighting with high contrast.
- + Good anatomical rendering of the horse despite the unusual composition.
- − The astronaut's legs and lower body are awkwardly integrated with the horse's back.
- − Slightly muddy details in the dark space background.
GPT Image 1.5
- + Very high detail in the lunar surface and nebula background.
- + Dynamic action with the dust and lunar lander adding to the scene.
- − Completely failed the negative constraint/spatial instruction: shows an astronaut riding a horse instead of a horse on top.
- − Overly busy composition with many clashing celestial elements.
Verdict: The challenge contained a specific spatial constraint ('horse on top, not vice versa') which GPT Image 1 followed perfectly, creating the requested surreal image. GPT Image 1.5 completely ignored this instruction and produced a standard 'astronaut riding a horse' image. While Image 1.5 has more background detail, Image 1 is the superior response due to prompt adherence.
Outfit Transfer Challenge
Editing“Use Image 1 as the base person. Dress them in the exact elaborate outfit from Image 2 (including all layers, accessories, jewelry, and shoes). Carefully adapt the clothing to the body shape and pose in Image 1 while maintaining realistic fabric behavior, correct proportions, and perfect lighting/shadow matching. Keep the person’s exact face, hair, and background completely unchanged.”
AI Judge Analysis
GPT Image 1
- + Excellent preservation of the person's face, skin markings, and facial features.
- + Accurately replicates the coat texture and the specific plaid pattern of the scarf.
- + Maintains the background lighting and grain of the source image well.
- − The scarf is missing the fringe tassels visible in the target outfit.
- − The watch chosen (leather strap) does not match the gold metal link watch from the target image.
GPT Image 1.5
- + Includes the fringe tassels on the scarf and the gold link watch from the target outfit.
- + Captures the full-body pose including the jeans and leg positioning reasonably well.
- − Crops out the top half of the person's face, failing the instruction to keep the person's exact face unchanged.
- − The skin marking on the chin does not match the original person's vitiligo pattern.
- − Missing the sand on the skin that was present in the source image.
Verdict: GPT Image 1 is the clear winner because it fulfills the core requirement of preserving the person's identity and facial features while successfully transferring the complex wardrobe. GPT Image 1.5 failed the most basic constraint of the prompt by cropping out the subject's face and failing to maintain the specific skin pigment details from the source image.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
GPT Image 1
- + Excellent fur texture and photographic lighting integration
- + Captures the bored expression of the passenger perfectly
- + The taxi cap has a high-quality leather/fabric texture
- − The capybara's right paw is gripping the wheel in a slightly distorted, hand-like way
- − The interior feels a bit dark and lacks specific taxi equipment
GPT Image 1.5
- + The paws are more anatomically correct for a capybara while still gripping the wheel
- + Includes realistic taxi details like the meter on the dashboard
- + The composition provides a wider, more authentic view of the car's interior
- − The passenger's face is slightly less detailed and more blurred than Model A
- − The capybara's hat has a slightly more 'clip-art' feel to the logo compared to the rest of the image
Verdict: Both models followed the prompt exceptionally well, but GPT Image 1.5 is the preferred choice due to its better rendering of the capybara's paws and the inclusion of realistic taxi elements like the meter. GPT Image 1 has slightly better facial clarity for the passenger, but the hand-like anatomy of the capybara's right paw is a bit uncanny.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
GPT Image 1
- + Excellent typography rendering with zero spelling errors.
- + Clean and legible layout suitable for an actual invitation.
- + The lighting on the jack-o-lantern and ground is subtle and cinematic.
- − Failed to include the specific '7pm' time detail.
- − The location is incorrectly labeled as 'TIME: The Arches, NYC'.
GPT Image 1.5
- + Accurately included all requested text details including the date, time, and location.
- + High-quality 'vintage' aesthetic with a detailed thorny border and parchment texture.
- + Dynamic background featuring a graveyard and castle silhouette that adds to the gothic theme.
- − The 'frights' text on the scroll has some minor tailing/artifacting at the end.
- − The layout feels more cluttered than Model A due to the high amount of texture.
Verdict: GPT Image 1.5 is the clear winner because it followed the text prompts much more accurately, including specific time and location details that GPT Image 1 missed or mislabeled. While GPT Image 1 has a cleaner layout, GPT Image 1.5's adherence to the 'vintage parchment' and 'thorns' prompt created a much more visually compelling and complete invitation.
Bald man challenge
Image Editing“Give the person a full, thick head of natural hair with realistic texture, density, and a natural hairline. Preserve facial features and lighting.”
AI Judge Analysis
GPT Image 1
- + Excellent preservation of the jacket and background
- + Adds a very thick head of hair as requested
- − The hairline is unnaturally low, obstructing part of the forehead and distorting the eye area
- − Noticeable texture mismatch between the added hair and the original beard
GPT Image 1.5
- + Natural-looking curly hair texture that blends perfectly with the existing beard
- + Realistic hairline and forehead preservation
- + Maintains original facial structure and eye shape
- − Slightly changes the bridge of the nose and the frames of the glasses compared to the original
Verdict: GPT Image 1.5 is the clear winner because it provides a much more convincing and natural-looking result with a realistic hairline and hair texture that fits the subject's existing features. GPT Image 1, while preserving the background better, fails on the core task by creating an unnaturally low hairline that makes the forehead look cramped and the hair look like a wig.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
GPT Image 1
- + Excellent adherence to the 'cartoon scene' style with soft, rounded shapes.
- + Extremely clean text rendering and layout.
- + Strong execution of the 45-degree isometric perspective and diorama base.
- − Very limited variety of sushi compared to the prompt's potential.
- − Textures are perhaps a bit too simplified/plasticky for 'realistic PBR'.
GPT Image 1.5
- + Superb material quality with realistic PBR textures on the wood, ceramic, and fish.
- + Includes a wider variety of sushi items and accessories like the teapot and soy sauce.
- + Accurate text rendering and placement.
- − A bit cluttered compared to the 'minimal garnish' and 'small raised diorama' request.
- − The 'cartoon' aspect is less pronounced than in Image A.
Verdict: Both models followed the complex prompt instructions for text and layout perfectly. Image A (GPT Image 1) better captures the 'cartoon' and 'minimalist' aesthetic requested, whereas Image B (GPT Image 1.5) excels at the 'realistic PBR materials' and offers a more detailed, premium miniature look. Image B is slightly better overall due to the superior textural detail and material rendering.
Over-the-top cartoon caricature
Editing“Create a caricature of me and my job. Make it exaggerated and humorous, incorporating my profession as a tv show anchor and my love for dogs and hockey.”
AI Judge Analysis
GPT Image 1
- + Excellent realization of the caricature art style with hand-drawn watercolor aesthetics.
- + Great source preservation by keeping the subject's denim jacket from the original photo.
- + Creative integration of the theme by having a dog as the hockey player on the TV screen.
- − The facial exaggeration is quite extreme, bordering on creepy.
- − The background is very simple and less 'news-like' compared to the other model.
GPT Image 1.5
- + Successfully incorporates all elements into a busy, humorous scene including a dog in a hockey helmet.
- + High level of detail in the studio setting with cameras, lighting, and news tickers.
- + The caricature face maintains a better likeness to the source person despite the exaggeration.
- − Changed the subject's clothing from the original denim to a red dress.
- − The hand holding the microphone has anatomical issues with the thumb and finger placement.
Verdict: GPT Image 1.5 is the preferred model because it created a more comprehensive and professional-looking news studio scene that perfectly balanced the requested hobbies. While GPT Image 1 captured the caricature style well and preserved the source clothing, its facial exaggeration was slightly unsettling and the composition was more basic.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
GPT Image 1
- + Excellent dynamic composition with a clear sense of movement and 'tumbling'.
- + Consistent lighting and beautiful god rays that create a cinematic atmosphere.
- + Distinct separation between the four animals, making Each easy to identify.
- − The fox kit has unnatural dark brown/black paws that look like gloves.
- − The kitten's eye anatomy is slightly distorted compared to real cats.
GPT Image 1.5
- + Highly detailed fur texture and adorable 'toe beans' visible on the kitten.
- + Excellent 'dew sparkles' effect and warm golden sunrise lighting.
- + Captures a very charming, cuddly interaction between the animals.
- − More static composition compared to Model A; less of a 'chasing' feel.
- − The bunny looks somewhat squashed and less realistic between the other animals.
- − The kitten has an extra toe or claw-like protrusion on its raised paw.
Verdict: Both models followed the prompt exceptionally well, capturing all four animals with high-quality fur and lighting effects. GPT Image 1 (Model A) is the likely winner because it captures the active energy of 'playfully chasing' and 'tumbling' much more effectively than the static posing in GPT Image 1.5 (Model B).
Studio Ghibli Anime Style
Editing“Transform this photo into a Studio Ghibli–inspired illustration. Use soft pastel colors, hand-painted textures, gentle lighting, dreamy backgrounds, and a warm, nostalgic mood”
AI Judge Analysis
GPT Image 1
- + Expertly captures the hand-painted, textured aesthetic of traditional Ghibli backgrounds.
- + Maintains the exact compositional balance and poses of the source image.
- + The facial expressions perfectly translate the 'distracted boyfriend' meme into an anime style.
- − The colors are a bit muted compared to some of Ghibli's more vibrant outdoor scenes.
GPT Image 1.5
- + Excellent 'dreamy' lighting effect with a beautiful sunlight glow.
- + High visual quality with clean lines and soft gradients.
- + Preserves the plaid pattern on the shirt accurately.
- − The style leans closer to modern digital anime/shoujo than the specific hand-painted Ghibli aesthetic.
- − The woman in the red dress has a somewhat generic face compared to Model A.
Verdict: Both models successfully interpreted the prompt, but GPS Image 1 (Model A) is the clear winner for style accuracy. It perfectly replicates the specific pencil-and-watercolor texture and character design language of Studio Ghibli, whereas GPT Image 1.5 (Model B) feels like a more generic modern digital illustration. Model A also does a better job of capturing the specific frustration/smugness of the facial expressions from the original meme.
Golden Hour Stroll
Image Editing“Add dynamic motion to this photo: make hair blow in the wind, add leaves flying, energetic and lively feel.”
AI Judge Analysis
GPT Image 1
- + Excellent preservation of the woman's face and original features
- + Hair movement looks very natural and follows the direction of the wind well
- + Dog's tail is modified to show realistic movement
- − Flying leaves are small and clumped, looking almost like debris
- − Overall brightness and saturation decreased noticeably from the source
GPT Image 1.5
- + Successfully added a large number of colorful flying leaves
- + Hair movement is dynamic and symmetrical
- + Maintains clarity and color saturation closer to the source image
- − The woman's face has changed slightly, losing some of the likeness to the original
- − Hair strands look a bit more artificial and 'pasted on' compared to the source
Verdict: Both models followed the instructions well, but GPT Image 1 (Model A) did a better job at preserving the specific facial features of the woman from the source image. GPT Image 1.5 (Model B) created a more 'lively' feel with the colorful leaves, but the slight alteration of the subject's face makes it less successful as a pure image edit.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
GPT Image 1
- + Clean minimalist design that fits modern vector standards
- + Precise typography and accurate accent on 'Caffè'
- + Superior texture consistency across the graphic
- − Single color choice is slightly less dynamic than requested 'brown and cream' palette
- − Failed the background requirement by using a black background instead of a light one
GPT Image 1.5
- + Excellent use of warm brown and cream tones as requested
- + Detailed vector emblem style with sophisticated shading on the cloche
- + Captured the 'retro' feel more effectively with the banner design
- − Failed the background requirement by using a black background instead of a light one
- − Typography for 'Caffè' is slightly less legible than Model A
Verdict: Both models failed to adhere to the 'light background' instruction, delivering black backgrounds instead. GPT Image 1 provides a more functional, minimalist vector logo, while GPT Image 1.5 offers a much more visually appealing 'vintage' aesthetic through better use of the requested color palette and sophisticated shading, making it the better choice for the specific prompt.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
GPT Image 1
- + Excellent adherence to the 'navy, white, muted red' color palette
- + Clean flat-vector style with professional-looking silhouettes
- + Accurate spelling of the crew names
- − Confusing layout where labels do not clearly align with their respective icons
- − Nonsense word 'EARLLUNAR' replaces 'Lunar Orbit' in the footer area
- − Lacks the specific steps/numbers requested in a logical sequence
GPT Image 1.5
- + Highly organized grid layout that clearly follows the requested 6-step sequence
- + Strong visual narrative with distinct panels for each phase of the mission
- + Perfect spelling for all main phase labels and crew names
- − Introduced green on Earth which was not in the 'NASA-inspired' specific palette
- − Minor spelling deviation in 'Tranquillity' (double 'L' is British variant, but 'Tranquility' was in prompt)
Verdict: GPT Image 1.5 is a significantly better infographic as it follows the requested chronological 6-step sequence with a clear, logical grid. While GPT Image 1 has a slightly more sophisticated art style, its layout is chaotic and includes a major spelling error in a primary label. GPT Image 1.5 preserves the infographic intent much more effectively with clear labeling for each specific step.
Explore each model
OpenAI's state-of-the-art image generation model with better instruction following and adherence to prompts