Head to head
Esc

Models · slot A

to navigate to pick

GPT Image 1.5 OpenAI GPT Image 2 OpenAI

Settled by community votes across 15 shared challenges, with an AI judge weighing in on each.

GPT Image 1.5

27.1 arena score

#7 of 62 in Text-to-Image

Top 3 in Image Editing
Skill signature · Text-to-Image

GPT Image 2

28.1 arena score

#3 of 62 in Text-to-Image

Top 3 in Text-to-Image
Vote tally

Where the votes landed

GPT Image 1.5

0%

win rate

Ties

0%

GPT Image 2

0%

win rate

Shared challenges 15

Challenge by challenge

The strongest take from each model on every shared challenge, with the AI judge's read.

Geometric Composition

Text-to-Image

“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”

GPT Image 1.5
GPT Image 2

AI Judge Analysis

GPT Image 1.5

  • + Excellent adherence to lighting instructions with a clear left-to-right gradient.
  • + Realistic glass refraction and reflections on the sphere and table.
  • + The plant is visible both behind and through the glass cube as requested.
  • The sphere is slightly large for the described 'small sphere' prompt.

GPT Image 2

  • + High resolution and crisp texture on the book cover and table wood grain.
  • + Clean, modern composition with a very clear glass box.
  • + The sphere size better represents the prompt 'small blue sphere'.
  • The lighting is more front-lit than specifically coming from the left window.
  • The refraction of the plant through the glass is less pronounced than in Model A.

Verdict: Both models followed the spatial instructions perfectly. GPT Image 1.5 (Image A) is the winner because it better captured the specific lighting request and the complex visual effect of the plant being seen through the glass, whereas GPT Image 2 (Image B) felt slightly more sterile and missed the directional light nuance.

Candid Street Photography

Text-to-Image

“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”

GPT Image 1.5
GPT Image 2

AI Judge Analysis

GPT Image 1.5

  • + Excellent depiction of rain with visible droplets falling and realistic wet-surface reflections.
  • + Great bokeh and 50mm feel that matches the cinematic request.
  • + Superior skin texture and an 'imperfect framing' that feels like a genuine candid photo.
  • The rear wheel of the bicycle has some structural inconsistencies where the spokes meet the cassette.

GPT Image 2

  • + Successfully incorporates motion blur from a passing car in the background.
  • + Good rendering of a tool kit and surrounding street elements like the Japanese signage.
  • + Accurate elderly facial features and natural hair rendering.
  • The lighting feels flat and less 'cinematic' compared to Image A.
  • The bike's handlebars and frame have significant geometric distortions and floating parts.
  • The rainy atmosphere is very subtle, lacking the texture of light rain requested in the prompt.

Verdict: GPT Image 1.5 is the clear winner as it captures the atmosphere of 'light rain' and a 'cinematic' look much more effectively than GPT Image 2. While GPT Image 2 follows the 'motion blur' instruction well, it suffers from significant AI artifacts in the bicycle's structure and lacks the photographic depth of Image 1.5.

Fantasy Warrior

Text-to-Image

“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”

GPT Image 1.5
GPT Image 2

AI Judge Analysis

GPT Image 1.5

  • + Exceptional lighting with realistic warm torchlight reflections on the metal surfaces.
  • + Highly detailed cloth and leather textures that meet all prompt requirements.
  • + Strong atmosphere with visible bokeh sparks and realistic battle-worn details such as blood and grime.
  • The scars look slightly like surface paint or fresh smeared blood rather than healed tissue.

GPT Image 2

  • + More delicate hair braiding and realistic skin texture with subtle freckles.
  • + Good use of shallow depth of field for a professional portrait look.
  • + Excellent engraving detail on the pauldrons.
  • Lighting is more diffused and lacks the specified 'warm torchlight' intensity seen in the metal reflections.
  • The overall color palette is a bit muted and less cinematic than requested.

Verdict: GPT Image 1.5 is the winner as it perfectly captures the cinematic lighting and 'battle-worn' aesthetic requested in the prompt, especially the warm reflections on the armor. While GPT Image 2 has lovely detail in the hair and face, it feels a bit too clean and softly lit compared to the high-contrast atmosphere specified.

Modern Clean Menu

Text-to-Image

“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”

GPT Image 1.5
GPT Image 2

AI Judge Analysis

GPT Image 1.5

  • + Excellent typography rendering with zero spelling errors
  • + Clean, high-quality food photography that looks realistic
  • + Strict adherence to the bold sans-serif requirement
  • The 'grid' layout for photos is a bit basic, just stacked on the right
  • Less branding elements compared to Model B

GPT Image 2

  • + Sophisticated grid-based composition that integrates text and images perfectly
  • + Includes additional professional design elements like logos, icons, and a footer
  • + Very vibrant use of color and 'vibrant accents' as requested
  • Almost perfect text but contains a minor font artifact on one price ($14.99 looks slightly distorted)
  • Slightly more cluttered than Model A, though still professional

Verdict: Both models performed exceptionally well, producing production-ready menu designs. GPT Image 1.5 (Model A) is cleaner and more minimalist with perfect text, but GPT Image 2 (Model B) delivers a much more comprehensive and creative branding package that makes better use of the grid layout and vibrant accents requested in the prompt.

Magic Burger Explosion: Fiery Photorealism Challenge

Text-to-Image

“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”

GPT Image 1.5
GPT Image 2

AI Judge Analysis

GPT Image 1.5

  • + Excellent photorealistic texture on the meat patty and toasted bun
  • + Dynamic use of smoke and flying sparks that enhances the fiery theme
  • + The text is perfectly integrated and follows the fiery effect prompt
  • The '6.99' price tag has a slight typo in the separator
  • The composition is a bit crowded with the large title overlapping the top bun

GPT Image 2

  • + High clarity on individual ingredients like the red onions and dripping sauce
  • + Clean and well-organized layout with text on the left and burger on the right
  • + The fiery glow effect on the text is very distinct and vibrant
  • The burger feels slightly less 'exploded' and more like a tilted stack
  • The lighting on the lettuce looks a bit more artificial compared to Model A

Verdict: Both models followed the prompt exceptionally well, producing high-impact advertising visuals. GPT Image 1.5 wins slightly on the photorealistic quality of the food and the atmospheric integration of smoke, while GPT Image 2 offers a cleaner graphic design layout with better text legibility.

Chalkboard Menu

Text-to-Image

“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”

GPT Image 1.5
GPT Image 2

AI Judge Analysis

GPT Image 1.5

  • + Excellent text legibility and accuracy of the requested menu items.
  • + Consistent chalk texture across all letters.
  • + Good use of letter size variation which makes it look authentically handwritten.
  • The composition is a tight crop that misses the 'cozy café' atmosphere requested.
  • The border of the chalkboard is barely visible.

GPT Image 2

  • + Stronger adherence to the 'cozy café' setting with background elements and a wooden frame.
  • + Includes a small bowl of chalk which adds to the realism of the scene.
  • + Text is well-rendered and follows the handwritten request effectively.
  • The title cursive is slightly less elegant than Model A.
  • Minor artifacting in the very small text at the bottom.

Verdict: GPT Image 1.5 provides excellent text rendering and chalk texture for the menu itself, but GPT Image 2 follows the environmental aspects of the prompt much better by showing the chalkboard within a café setting. Both models handled the specific text requirements and date perfectly, but GPT Image 2 feels more like a complete photograph rather than just a texture scan.

Pose & Character Mashup

Editing
Edit instruction

“Use Image 1 as the exact pose reference and Image 2 as the character reference. Recreate the person/character from Image 2 in the exact dynamic pose and body position from Image 1. Keep the exact face, hair, clothing style/details, and expression from Image 2. Match the lighting and environment of Image 1. The final image must show the character from Image 2 performing the precise action/pose from Image 1 with perfect anatomy and natural integration.”

Source
GPT Image 1.5
GPT Image 2

AI Judge Analysis

GPT Image 1.5

  • + Excellent character preservation for the face and specific clothing details like the orange text.
  • + Matching the lighting and vibrance of the yellow background perfectly.
  • + High visual clarity and resolution in the rendered skin and fabric details.
  • Failed to match the dynamic pose from Image 1, keeping the torso upright instead of bent.
  • Anatomical issues with the feet and how they connect to the legs.

GPT Image 2

  • + Successfully replicated the exact tilt and dynamic body position from Image 1.
  • + High fidelity in preserving the character's clothing, scarf, and accessories.
  • + Excellent integration of the character into the specific lighting and environment of the source image.
  • Minor anatomical artifacts on the lower hand (extra finger/nail detail).
  • The face angle is slightly distorted compared to the reference image.

Verdict: GPT Image 2 is the winner as it correctly followed the instruction to use Image 1 as the exact pose reference, recreating the complex lean and arm positioning. While GPT Image 1.5 preserved the character's face shape slightly better, it completely failed to capture the dynamic pose, providing a standard standing position instead.

The Reversed Rodeo

Text-to-Image

“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”

GPT Image 1.5
GPT Image 2

AI Judge Analysis

GPT Image 1.5

  • + Excellent cinematic lighting and composition
  • + High level of detail in the space background and textures
  • Completely failed the semantic relationship requested in the prompt
  • Standard interpretation of an astronaut riding a horse

GPT Image 2

  • + Successfully followed the specific instruction of horse on top
  • + High visual clarity and convincing lighting on the subjects
  • + Accurate rendering of the NASA logo and space suit textures
  • Anatomical issues where the horse's legs merge into the suit
  • The horse's hooves are holding reins, which is physically nonsensical

Verdict: GPT Image 1.5 failed the negative constraint and logic instruction, providing a standard astronaut on a horse. GPT Image 2 successfully interpreted the surreal 'horse on top' request, making it the clear winner for prompt adherence despite some minor anatomical artifacts.

Outfit Transfer Challenge

Editing
Edit instruction

“Use Image 1 as the base person. Dress them in the exact elaborate outfit from Image 2 (including all layers, accessories, jewelry, and shoes). Carefully adapt the clothing to the body shape and pose in Image 1 while maintaining realistic fabric behavior, correct proportions, and perfect lighting/shadow matching. Keep the person’s exact face, hair, and background completely unchanged.”

Source
GPT Image 1.5
GPT Image 2

AI Judge Analysis

GPT Image 1.5

  • + Excellent adherence to the outfit details including the watch and sunglasses.
  • + High image resolution and sharp texture on the fabric.
  • + Successfully replicated the vitiligo pattern on the hands/wrists to match the subject.
  • Substantially altered the person's face and head shape, losing the 'exact face' requirement.
  • Changed the hair style and texture compared to Image 1.

GPT Image 2

  • + Perfectly preserved the subject's face, hair, and head shape from Image 1.
  • + Accurately transferred the outfit including the plaid scarf and navy coat.
  • + Balanced composition that feels very integrated with the original background.
  • Missed the sunglasses from the second image.
  • The lighting on the subject's face is slightly flatter than the original Image 1.

Verdict: GPT Image 2 is the clear winner because it successfully followed the difficult constraint of keeping the person’s face and hair completely unchanged while applying the new clothing. While GPT Image 1.5 did a great job with the accessories like the sunglasses and watch, it failed the fundamental task by generating a new face that only vaguely resembles the original person.

The Capybara Taxi Driver

Text-to-Image

“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”

GPT Image 1.5
GPT Image 2

AI Judge Analysis

GPT Image 1.5

  • + Excellent photorealism with detailed capybara fur and realistic lighting.
  • + Accurate adherence to the cap with 'TAXI' text and checkered band.
  • + Perfect realization of the woman's bored expression while looking at her phone.
  • The capybara's paws look slightly mutated with too many digits or claws.
  • The perspective through the front windshield feels a bit flat compared to the side views.

GPT Image 2

  • + Natural composition with a dynamic angle looking in through the taxi door.
  • + Strong bokeh effect on the background city lights enhancing the nighttime atmosphere.
  • + High-quality rendering of the capybara's professional posture.
  • The woman's face is slightly blurry and lacks the distinct 'bored' clarity found in the prompt.
  • The 'T' logo on the hat is less iconic than the literal 'TAXI' text requested for a New York theme.

Verdict: Both models followed the prompt exceptionally well, but GPT Image 1.5 is the winner due to the superior clarity and expression of the businesswoman, which was a key part of the prompt's narrative. GPT Image 1.5 also provided a more literal and satisfying interpretation of the taxi driver's cap and the specific 'bored' look, whereas GPT Image 2 had slightly more artistic lighting but softer details on the secondary subject.

The Halloween Invitation

Text-to-Image

“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”

GPT Image 1.5
GPT Image 2

AI Judge Analysis

GPT Image 1.5

  • + Text rendering is exceptionally clean and perfectly legible throughout the invitation.
  • + The thorny border is very prominent and captures the requested gothic texture effectively.
  • + Excellent cinematic lighting on the central jack-o-lantern and background graveyard.
  • The composition feels slightly more generic than its counterpart.
  • The banner scroll under the pumpkin is a bit simple in its design.

GPT Image 2

  • + Highly detailed and intricate gothic border including skulls and ornate filigree.
  • + Includes creative background elements like an actual stone arch bridge, referencing the location 'The Arches'.
  • + Superior scroll banner design with more realistic parchment curls.
  • The text at the bottom is noticeably smaller and slightly harder to read compared to the top text.
  • The jack-o-lantern is placed deeper in the shadows, losing some of the 'glowing' prominence requested.

Verdict: Both models followed the prompt instructions perfectly, with GPT Image 1.5 providing superior text legibility and a very focused, classic composition. However, GPT Image 2 is the winner due to its superior artistic detail, including the intricate border and the thematic inclusion of arches to represent the venue, which makes for a more professional-looking invitation.

Isometric Miniature Diorama Scenes

Text-to-Image

“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”

GPT Image 1.5
GPT Image 2

AI Judge Analysis

GPT Image 1.5

  • + Excellent PBR material rendering on the wood texture and teapot
  • + High level of realistic detail on the sushi textures and soy sauce bottle
  • + Clean and readable typography with a nicely integrated flag
  • The diorama base is a bit cluttered with large objects like the teapot
  • The isometric perspective feels slightly flattened compared to Model B

GPT Image 2

  • + Perfect 45-degree isometric perspective creates a strong 3D diorama feel
  • + Creative inclusion of a miniature stone lantern and garden elements
  • + Beautifully balanced composition and clean spacing between elements
  • The 'JAPAN' text is slightly less clean with its heavy black outline
  • The chopsticks are rendered with slightly inconsistent thickness

Verdict: Both models followed the prompt exceptionally well, but GPT Image 2 captures the 'miniature 3D diorama' aesthetic more effectively with its superior isometric perspective and charming environment details. While GPT Image 1.5 has slightly more realistic textures on the food, GPT Image 2 presents a more cohesive and visually appealing composition.

Adorable Baby Animals in Sunny Meadow

Text-to-Image

“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”

GPT Image 1.5
GPT Image 2

AI Judge Analysis

GPT Image 1.5

  • + Excellent expressive facial features on all animals
  • + Wonderful capture of morning dew and soft particles in the light
  • + High level of detail in the fur texture and kitten's paws
  • The kitten has five paws/limbs visible
  • The fox's front right leg is anatomically confusing

GPT Image 2

  • + Better sense of motion and 'chasing' as requested in the prompt
  • + More realistic flower variety and natural meadow depth
  • + Clearer distinction between the individual animals and their actions
  • The fox kit's legs and paws are poorly rendered and distorted
  • The lighting feels slightly more artificial with heavy bloom

Verdict: Both models captured the requested aesthetic well, but GPT Image 2 better followed the instruction for a 'chasing' and 'tumbling' scene, creating a more dynamic composition. GPT Image 1.5 has superior facial expressions and texture detail, but suffers from significant anatomical errors like the kitten having too many legs.

Vintage Cafe Logo

Text-to-Image

“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”

GPT Image 1.5
GPT Image 2

AI Judge Analysis

GPT Image 1.5

  • + Excellent typography with a custom hand-lettered feel
  • + Good use of color blocking for a modern-retro look
  • + Accurate rendering of the decorative banner
  • Failed the light background instruction by using a solid black background
  • The steam effect is a bit chunky and less elegant than model B

GPT Image 2

  • + Perfectly adhered to the light background and subtle texture instructions
  • + High-quality vector engraving style with fine detail on the cloche
  • + Elegant and balanced composition within a decorative frame
  • The 'EST.' text on the banner is slightly off-center
  • Typography is more standard compared to the creative lettering in model A

Verdict: GPT Image 2 is the winner because it followed all prompt instructions, including the light background and subtle texture which GPT Image 1.5 ignored. While GPT Image 1.5 had more interesting typography, GPT Image 2's sophisticated engraving style and overall composition better fit the 'vintage minimalist' and 'vector emblem' aesthetic requested.

Apollo 11: Journey to Tranquility

Text-to-Image

“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”

GPT Image 1.5
GPT Image 2

AI Judge Analysis

GPT Image 1.5

  • + Features very clean, bold flat-vector illustrations that match the requested style perfectly.
  • + The font rendering is crisp and perfectly legible.
  • + The layout uses a comic-style panel approach that is intuitive and engaging.
  • The color of the Earth and the ground plane for Earth Orbit uses a reddish-brown color that feels more like Mars than Earth.
  • Incorrectly spells 'Tranquillity' with an extra 'l' (though this is a valid British variant, it's less standard for NASA contexts).

GPT Image 2

  • + Excellent typography and overall poster composition with a clear 'Apollo 11' header.
  • + Excellent adherence to the NASA-inspired palette including iconic logos and mission patches.
  • + Detailed and accurate icons for each of the six requested steps.
  • The icons for steps 3 and 4 use photorealistic textures that clash with the flat vector style of the other elements.
  • Small text like 'Humanity's first step on the moon' has minor kerning and alignment inconsistencies.

Verdict: GPT Image 2 is the superior infographic, as it successfully incorporates all requested steps into a cohesive, professional poster layout with a clear title and logical flow. While GPT Image 1.5 has more consistent flat-vector styling, its individual panels feel more like separate cards than a unified infographic poster, and its color choice for Earth is confusing.

Next steps

Explore each model