Head to head
Esc

Models · slot A

to navigate to pick

GPT Image 2 OpenAI Wan 2.5 (Preview) Alibaba

Settled by community votes across 14 shared challenges, with an AI judge weighing in on each.

GPT Image 2

27.7 arena score

#4 of 62 in Text-to-Image

Skill signature · Text-to-Image

Wan 2.5 (Preview)

23.4 arena score

#27 of 62 in Text-to-Image

Vote tally

Where the votes landed

GPT Image 2

0%

win rate

Ties

0%

Wan 2.5 (Preview)

0%

win rate

Shared challenges 14

Challenge by challenge

The strongest take from each model on every shared challenge, with the AI judge's read.

Geometric Composition

Text-to-Image

“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”

GPT Image 2
Wan 2.5 (Preview)

AI Judge Analysis

GPT Image 2

  • + Excellent structural geometry of the cube
  • + Realistic glass thickness and refraction
  • + Clean and professional photographic quality
  • The scale of the sphere is very small compared to the prompt's focus
  • Composition is a bit static

Wan 2.5 (Preview)

  • + Better lighting with dynamic window shadows on the table
  • + The blue sphere is more prominent and visually appealing
  • + Natural-looking dust motes add to the realism of the scene
  • The glass cube geometry is slightly warped on the top left corner
  • The red book looks a bit worn compared to the clean aesthetic of the other objects

Verdict: Both models followed the spatial instructions perfectly. GPT Image 2 produced a cleaner, more structurally accurate glass cube, while Wan 2.5 (Preview) excelled at atmosphere and lighting, creating a more visually engaging scene with better light play.

Candid Street Photography

Text-to-Image

“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”

GPT Image 2
Wan 2.5 (Preview)

AI Judge Analysis

GPT Image 2

  • + Excellent handling of motion blur from passing cars
  • + Features highly realistic, imperfect framing that feels like a genuine street photo
  • + Superb tool and environment details, including legible Japanese text
  • The man's hands are slightly distorted around the bicycle spokes

Wan 2.5 (Preview)

  • + Strong reflections on the wet pavement and visible rain droplets
  • + Good composition with a clear shallow depth of field
  • + Accurate representation of the requested red bicycle
  • The bicycle structure is physically impossible, with spokes and frame components clumping together
  • Lacks the requested motion blur for passing cars
  • The lighting on the man's hair and shirt feels slightly digital rather than natural

Verdict: GPT Image 2 is the superior output because it successfully captures the complexity of the 'candid' and 'motion blur' requirements, resulting in an image that looks like a real 50mm photograph. Wan 2.5 (Preview) has very strong atmospheric effects like rain and reflections, but it fails significantly on the physical anatomy of the bicycle and ignores the motion blur request.

Fantasy Warrior

Text-to-Image

“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”

GPT Image 2
Wan 2.5 (Preview)

AI Judge Analysis

GPT Image 2

  • + Exquisite detailing on the engraved plate armor and worn leather straps.
  • + Masterful handling of natural-looking skin textures, faint scars, and dirt.
  • + Sophisticated lighting and shallow depth of field that creates a lifelike, cinematic atmosphere.
  • The beads in the hair are very subtle and blend in with the hair texture.

Wan 2.5 (Preview)

  • + Literal interpretation of 'hair braided with small beads'.
  • + Strong presence of the torchlight and bokeh sparks as requested.
  • + Good representation of the tattered cloth underlayer and chainmail.
  • The facial features and skin texture look slightly more digital/less lifelike than Model A.
  • The 'battle-worn' effect on the face looks more like painted-on smudges than authentic dirt and scarring.

Verdict: GPT Image 2 is the superior image due to its incredible textural realism, particularly in the weathered metal and the subtle, lifelike details of the skin. While Wan 2.5 (Preview) more clearly depicted the hair beads and cloth underlayer, its overall composition feels slightly more like a CGI character compared to the photographic quality of GPT Image 2.

Modern Clean Menu

Text-to-Image

“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”

GPT Image 2
Wan 2.5 (Preview)

AI Judge Analysis

GPT Image 2

  • + Exceptional text rendering with perfect spelling and legibility.
  • + Highly professional layout that follows graphic design principles.
  • + Logical and diverse food photography that matches the item descriptions.
  • Very dense layout that might feel slightly cluttered compared to extreme minimalism.

Wan 2.5 (Preview)

  • + Strong minimalist aesthetic with plenty of white space.
  • + Consistent lighting and styling across all food photography.
  • Garbled, nonsensical text in titles and descriptions.
  • Incorrect categorizations, such as showing non-pizza items in the pizza section.
  • Significant spelling errors in prominent headings like 'Restormalit Menue'.

Verdict: GPT Image 2 is the clear winner for design and functionality, producing a production-ready menu with perfect text, logical sections, and professional typography. Wan 2.5 (Preview) fails significantly on the textual elements and logical grouping, though it does capture a pleasingly clean aesthetic in its photography.

Magic Burger Explosion: Fiery Photorealism Challenge

Text-to-Image

“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”

GPT Image 2
Wan 2.5 (Preview)

AI Judge Analysis

GPT Image 2

  • + Excellent photorealistic texture on the meat and bun
  • + Perfect adherence to the fiery text effect request
  • + Complex and realistic deconstructed burger arrangement
  • The composition feels slightly cluttered with the large text blocks
  • Less overall sense of 'exploded' centrifugal motion compared to Model B

Wan 2.5 (Preview)

  • + Strong sense of dynamic motion with ingredients flying outward
  • + Clean composition with professional-looking graphic design elements
  • + Accurate rendering of all requested text and price starburst
  • The burger ingredients look slightly more illustrative and less photorealistic than Model A
  • The 'glow' on the price tag is more of a sticker effect than a fiery integration

Verdict: GPT Image 2 (Model A) wins on pure photorealism and textural detail, capturing the lighting and fire effects with superior grit and realism. Wan 2.5 (Model B) offers a more balanced commercial layout and better sense of outward motion, but its textures look slightly artificial in comparison.

Chalkboard Menu

Text-to-Image

“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”

GPT Image 2
Wan 2.5 (Preview)

AI Judge Analysis

GPT Image 2

  • + Excellent adherence to the 'chalk texture' requirement with grainy, realistic strokes.
  • + Perfect spelling and inclusion of all prompt details, including the completed cookie name.
  • + Highly realistic cursive handwriting style that looks authentically human.
  • The lighting in the upper left is slightly harsh but fits the cafe aesthetic.

Wan 2.5 (Preview)

  • + Dynamic composition with a nice perspective angle.
  • + Strong chalk smudge effects on the board surface for realism.
  • The text looks more like a digital chalk-style font rather than organic handwriting.
  • Failed to complete the 'Herbs' text in the octopus dish.
  • Repeats the price for the cookies across two lines awkwardly.

Verdict: GPT Image 2 is the clear winner as it followed every instruction, including the specific text strings and the demand for a realistic, non-digital handwriting style. Wan 2.5 (Preview) produced text that looks like a consistent digital font and failed to include all the menu words correctly.

The Reversed Rodeo

Text-to-Image

“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”

GPT Image 2
Wan 2.5 (Preview)

AI Judge Analysis

GPT Image 2

  • + Excellent adherence to the specific prompt instruction of having the horse on top.
  • + Highly surreal concept that perfectly matches the requested tone.
  • + Clean textures and realistic lighting on both the spacesuit and horse fur.
  • The leather straps and reins have nonsensical physics and connections.
  • The astronaut's hands are splayed in a somewhat unnatural, finger-heavy pose.

Wan 2.5 (Preview)

  • + Dynamic composition with a sense of movement and energy.
  • + Beautiful cosmic background with vibrant colors and a spiral galaxy.
  • + High level of detail in the horse's anatomy and the astronaut's suit.
  • Failed the negative constraint/specific instruction to have the horse on top.
  • Clipped hooves and repetitive tail/mane patterns.

Verdict: GPT Image 2 followed the specific and difficult 'horse on top' instruction perfectly, resulting in a unique and surreal image. Wan 2.5 (Preview) ignored the spatial constraint entirely, producing a standard 'astronaut on horse' image which, while visually beautiful, does not meet the core prompt requirement.

Outfit Transfer Challenge

Editing
Edit instruction

“Use Image 1 as the base person. Dress them in the exact elaborate outfit from Image 2 (including all layers, accessories, jewelry, and shoes). Carefully adapt the clothing to the body shape and pose in Image 1 while maintaining realistic fabric behavior, correct proportions, and perfect lighting/shadow matching. Keep the person’s exact face, hair, and background completely unchanged.”

Source
GPT Image 2
Wan 2.5 (Preview)

AI Judge Analysis

GPT Image 2

  • + Excellent preservation of the subject's unique identity, face, and hair
  • + Higher quality rendering of the scarf pattern
  • + Perfectly preserved the original background details and lighting
  • Missed the sunglasses from the second image

Wan 2.5 (Preview)

  • + Correctly included the sunglasses from the second image
  • Failed to preserve the base person's identity and face
  • Overlaid the person from Image 2 onto Image 1's setting instead of editing Image 1
  • Slightly more compressed/blurry fabric textures

Verdict: GPT Image 2 followed the complex instructions much more effectively by accurately keeping the identity of the person from Image 1 while applying the clothing from Image 2. Wan 2.5 (Preview) failed the primary negative constraint by completely replacing the person's face and hair with the subject from the second image.

The Capybara Taxi Driver

Text-to-Image

“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”

GPT Image 2
Wan 2.5 (Preview)

AI Judge Analysis

GPT Image 2

  • + Excellent photorealism with cinematic lighting and natural depth of field
  • + Accurate depiction of a business traveler with a convincing 'bored' expression
  • + Fur texture and clothing materials look highly realistic
  • The capybara's hands look slightly more like paws than true anatomy, though this fits the prompt's request for paws

Wan 2.5 (Preview)

  • + Strong composition from the front of the vehicle showing the taxi lights
  • + Good adherence to the prompt's details like the cap and jacket
  • + The background lights provide a vibrant sense of Times Square/NYC
  • The capybara's hands are strangely human-like/primate-like rather than paws
  • The businesswoman's face is slightly less detailed and more generic
  • The textures feel slightly more synthetic and less photorealistic compared to the competitor

Verdict: GPT Image 2 (Model A) is the winner due to its superior photorealistic quality and more natural lighting. While both models followed the prompt closely, Model A captured the specific mood of a tired traveler in the back seat more convincingly and avoided the unsettling human-like hands found in Wan 2.5 (Model B).

The Halloween Invitation

Text-to-Image

“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”

GPT Image 2
Wan 2.5 (Preview)

AI Judge Analysis

GPT Image 2

  • + Excellent typography with a genuine vintage gothic feel.
  • + Highly detailed monochromatic 'dark parchment' aesthetic that matches the theme.
  • + Sophisticated incorporation of NYC elements like the bridge in the background.
  • The parchment texture is very dark, which may reduce readability in some areas.

Wan 2.5 (Preview)

  • + Clear text legibility for all event details.
  • + Vibrantly colored central jack-o-lantern with a high-quality flame effect.
  • The 'Halloween Party Invitation' text looks like a modern 3D layer rather than vintage gothic calligraphy.
  • The composition feels like a digital collage rather than a cohesive vintage poster.
  • The blue sky and bright lighting clash with the requested 'moody night sky' and 'dark parchment' feel.

Verdict: GPT Image 2 perfectly captures the vintage gothic aesthetic with a cohesive, atmospheric design that looks like a professionally made invitation. While Wan 2.5 (Preview) has very clear text and a bright pumpkin, it lacks the 'polished, cinematic' and 'vintage' qualities requested, opting for a style that feels more like a generic modern greeting card. GPT Image 2's attention to detail, such as the thorn border and subtle skyline, makes it the superior choice.

Isometric Miniature Diorama Scenes

Text-to-Image

“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”

GPT Image 2
Wan 2.5 (Preview)

AI Judge Analysis

GPT Image 2

  • + Excellent adherence to the isometric miniature diorama request
  • + Highly detailed sushi varieties with realistic PBR material effects
  • + Perfect text rendering and layout placement
  • Background is a bit saturated, almost neon blue compared to the subject
  • The chopsticks are a bit short in proportion to the platform

Wan 2.5 (Preview)

  • + Clean, soft 3D cartoon aesthetic
  • + Effective use of lighting and depth of field
  • + Accurate text rendering
  • Missed the 'isometric' perspective core instruction
  • Failed to include a square diorama base
  • The flag icon is partially merged with the text improperly

Verdict: GPT Image 2 followed the complex prompt instructions perfectly, capturing the isometric 45° angle, the diorama base, and the specific text layout requested. While Wan 2.5 (Preview) produced a charming 3D render, it ignored the isometric perspective and diorama base requirements, resulting in a standard portrait shot.

Adorable Baby Animals in Sunny Meadow

Text-to-Image

“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”

GPT Image 2
Wan 2.5 (Preview)

AI Judge Analysis

GPT Image 2

  • + Excellent anatomical accuracy for all four animals.
  • + Beautiful lighting with natural-looking god rays and backlight through the fur.
  • + High level of detail in the textures of the fur and the surrounding wildflowers.
  • The 'dew sparkles' are present but less pronounced than in the competing image.

Wan 2.5 (Preview)

  • + Clear inclusion of dew drops and sparkles as requested.
  • + Dynamic composition with all animals appearing to be in mid-leap.
  • + Bright and vibrant color palette that enhances the joyful vibe.
  • The fox's eyes appear unnaturally blue and glowing, bordering on uncanny.
  • Anatomical issues with the animals, particularly the kitten's front right leg and the fox's facial structure.
  • The water droplets appearing in the air look more like floating glass orbs than natural dew.

Verdict: GPT Image 2 is the superior image due to its consistent anatomical realism and sophisticated lighting. While Wan 2.5 (Preview) captures a more energetic movement, it suffers from several artifacts and an 'uncanny valley' effect in the fox's eyes and the kitten's limbs.

Vintage Cafe Logo

Text-to-Image

“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”

GPT Image 2
Wan 2.5 (Preview)

AI Judge Analysis

GPT Image 2

  • + Excellent typography with correct accents and distinct font weights
  • + High-quality etching detail and texture
  • + Flawless rendering of all text including the date banner
  • The frame shape is slightly more complex than a minimalist vector might require

Wan 2.5 (Preview)

  • + Stronger vector aesthetic consistent with a modern-vintage logo
  • + Excellent alignment of the cloche and steam elements
  • Visual artifacts and distortions on the right-hand side of the banner
  • Inconsistent letter spacing and sizing in 'Caffè Florian'
  • The paper texture background is a bit generic

Verdict: GPT Image 2 is the superior generation as it perfectly renders the orthography of 'Caffè' and maintains high-quality engraving details throughout the emblem. While Wan 2.5 (Preview) captures a good vector style, it suffers from significant structural warping on the right side of the ribbon and poor kerning in the main text.

Apollo 11: Journey to Tranquility

Text-to-Image

“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”

GPT Image 2
Wan 2.5 (Preview)

AI Judge Analysis

GPT Image 2

  • + Excellent layout with a clear chronological timeline and distinct sections.
  • + Highly legible, accurate typography and iconography across all six requested steps.
  • + Matches the NASA-inspired aesthetic and modern vector style perfectly.
  • The illustrations of the Lunar Module are a bit more detailed than 'flat-vector' style.
  • The Apollo 11 mission patch in the corner has minor bird-rendering artifacts.

Wan 2.5 (Preview)

  • + Strong adherence to the flat-vector, minimalist aesthetic.
  • + Good use of white space and a clean color palette that matches the prompt.
  • Layout is cluttered and non-linear, making the mission steps hard to follow.
  • Text for 'Descent' and 'Landing' is floating without associated step icons.
  • Iconography is inconsistent, such as the inclusion of a shuttle-like craft for the descent stage.

Verdict: GPT Image 2 (Model A) is the clear winner as it successfully follows the structured infographic request, providing all six chronological steps with professional-grade layout and typography. Wan 2.5 (Preview) struggles with the informational hierarchy, failing to clearly define the later mission steps and providing confusing icons that don't match the historical context of Apollo 11.

Next steps

Explore each model