Head to head
Esc

Models · slot A

to navigate to pick

GPT Image 2 OpenAI Wan 2.7 Alibaba

Settled by community votes across 15 shared challenges, with an AI judge weighing in on each.

GPT Image 2

27.7 arena score

#4 of 62 in Text-to-Image

Skill signature · Text-to-Image

Wan 2.7

20.5 arena score

#38 of 62 in Text-to-Image

Vote tally

Where the votes landed

GPT Image 2

75.0%

win rate

Ties

0.0%

Wan 2.7

25.0%

win rate

75.0% 0.0% ties 25.0%
Shared challenges 15

Challenge by challenge

The strongest take from each model on every shared challenge, with the AI judge's read.

Geometric Composition

Text-to-Image

“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”

GPT Image 2
Wan 2.7

AI Judge Analysis

GPT Image 2

  • + Excellent adherence to the 'partially visible through glass' instruction for the plant.
  • + Very clean glass textures with realistic refractions.
  • + Simple and balanced composition that highlights all prompt elements.
  • The blue sphere appears slightly floating rather than resting on the bottom surface.

Wan 2.7

  • + Features more realistic, aged textures on both the book and the wooden table.
  • + Sophisticated handling of reflections on the glass surfaces.
  • + The sphere has a tangible weight and sits naturally on the surface.
  • The plant is mostly behind the glass but doesn't feel as integrated into the transparency as Image A.
  • There is a ghostly duplicate of the blue sphere occurring on the left side of the cube that isn't a natural reflection.

Verdict: GPT Image 2 followed the prompt with high precision, particularly regarding the plant's visibility through the glass cube. While Wan 2.1 has more impressive organic textures on the wood and book, it suffers from a confusing internal reflection of the sphere that looks like a second object. GPT Image 2 is preferred for its clarity and perfect adherence to the spatial requirements.

Candid Street Photography

Text-to-Image

“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”

GPT Image 2
Wan 2.7

AI Judge Analysis

GPT Image 2

  • + Excellent depiction of motion blur from passing cars as requested.
  • + Captures the textures of the wet environment and skin very realistically.
  • + Shallow depth of field is well-executed with convincing bokeh.
  • The bike anatomy is slightly warped near the rear wheel and spokes.
  • The 'imperfect framing' prompt led to a slightly cluttered foreground with the white sign.

Wan 2.7

  • + Successfully captures a wide-angle candid street atmosphere.
  • + Strong reflections on the wet pavement add to the cinematic feel.
  • + Good portrayal of light rain texture.
  • Fails to include motion blur on the passing cars as requested.
  • The man's hands are fused to the handlebars in a nonsensical way.
  • The background figures and cars lack the 'shallow depth of field' requested.

Verdict: GPT Image 2 is much more successful at capturing the technical photographic elements of the prompt, including the motion blur of passing traffic and a realistic shallow depth of field. Wan 2.7 provides a nice atmosphere but fails on key technical instructions like motion blur and contains significant anatomical errors where the man's hands merge with the bicycle.

Fantasy Warrior

Text-to-Image

“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”

GPT Image 2
Wan 2.7

AI Judge Analysis

GPT Image 2

  • + Excellent skin texture with realistic dirt and fine pores
  • + Atmospheric lighting with high-quality bloom and bokeh
  • + Ornate metal engraving with realistic pitting and wear
  • Hair beads are very small and less prominent than requested
  • Narrower composition hides most of the cloth underlayer textures

Wan 2.7

  • + Strong adherence to the 'beads' and 'scars' requirements with clear visibility
  • + Good balance of plate, leather, and cloth textures
  • + Highly distinct and vibrant bokeh sparks
  • The facial skin has a slightly 'plastic' or smooth digital sheen compared to the armor
  • The armor engraving looks a bit more generic and less finely etched than its counterpart

Verdict: GPT Image 2 provides a more cinematic and photorealistic portrait with superior skin and metal textures, whereas Wan 2.7 adheres more strictly to the specific details of the prompt like the hair beads and scars. Ultimately, GPT Image 2 is the better image due to its more sophisticated lighting and more lifelike eyes.

Modern Clean Menu

Text-to-Image

“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”

GPT Image 2
Wan 2.7

AI Judge Analysis

GPT Image 2

  • + Exceptional text rendering with coherent descriptions and accurate spelling.
  • + Strong professional layout that perfectly follows the requested section headers (Appetizers, Pizza, Mains).
  • + High-quality food photography that looks consistent and appetizing.
  • The layout is a bit dense with text, making it look slightly cluttered in the footer.

Wan 2.7

  • + Aesthetically pleasing presentation with realistic environmental props like the pen and rosemary.
  • + Good use of color accents and rounded UI-style buttons for navigation.
  • + Clean grid layout for the food images.
  • Poor text rendering with numerous spelling errors (e.g., 'Calanrfri', 'Bisge', 'Sannon').
  • Failed to organize by the specific sections requested; items are mixed in a single grid without clear headers.
  • The food descriptions are illegible gibberish.

Verdict: GPT Image 2 is the clear winner as it produces a fully functional, professional menu with perfectly legible and accurate text. While Wan 2.7 creates a lovely mock-up scene, it fails on the specific section requirements of the prompt and suffers from significant spelling errors and nonsensical text.

Magic Burger Explosion: Fiery Photorealism Challenge

Text-to-Image

“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”

GPT Image 2
Wan 2.7

AI Judge Analysis

GPT Image 2

  • + Excellent photorealistic texture on the meat patty and bun.
  • + Text rendering is perfectly integrated with the fiery theme using consistent glowing effects.
  • + Highly dynamic composition with realistic sauce splashes and flying embers.
  • The 'exploded' view is slightly compressed vertically compared to Model B.
  • Large amount of fiery texture can occasionally obscure the finer details of the background.

Wan 2.7

  • + Strong 'exploded' effect with clear separation between all requested ingredients.
  • + Clean, legible typography that follows the prompt instructions well.
  • + Creative use of smoke and sparks to add motion.
  • The textures look more like a digital illustration than the requested 'photorealistic' style.
  • Some gravity-defying elements look a bit static and less integrated into a single cohesive scene.
  • Floating nuts or seeds on the top left feel out of place.

Verdict: GPT Image 2 is the superior choice because it achieves a high level of photorealism that makes the food look appetizing, particularly the texture of the patty and the melting cheese. While Wan 2.7 has a very clear layout, its rendering style is more illustrative and cartoonish, failing to capture the 'photorealistic detail' requested in the prompt as effectively as GPT Image 2.

Chalkboard Menu

Text-to-Image

“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”

GPT Image 2
Wan 2.7

AI Judge Analysis

GPT Image 2

  • + Features highly realistic chalk texture with dusty, varied strokes
  • + Successfully renders all text accurately as requested
  • + Achieves an authentic handwritten cursive style for the title
  • The lighting in the background is slightly dark compared to the subject

Wan 2.7

  • + Perfect text accuracy and alignment
  • + Clean, high-contrast composition
  • Text appears as a clean digital font rather than natural chalk handwriting
  • Fails to provide the 'elegant cursive' requested for the title
  • Lacks the gritty, dusty realism of actual chalk on a board

Verdict: GPT Image 2 captured the handwritten chalk aesthetic perfectly, showing realistic variations in pressure and texture that look authentic to a cafe chalkboard. Wan 2.7 produced very legible text but failed the prompt's requirement for a realistic handwriting style, resulting in a 'digital font' look that lacks character.

Pose & Character Mashup

Editing
Edit instruction

“Use Image 1 as the exact pose reference and Image 2 as the character reference. Recreate the person/character from Image 2 in the exact dynamic pose and body position from Image 1. Keep the exact face, hair, clothing style/details, and expression from Image 2. Match the lighting and environment of Image 1. The final image must show the character from Image 2 performing the precise action/pose from Image 1 with perfect anatomy and natural integration.”

Source
GPT Image 2
Wan 2.7
100% wins 0% ties 0% wins

AI Judge Analysis

GPT Image 2

  • + Excellent adherence to the character reference, including face, sunglasses, and the specific black and white scarf.
  • + Precisely replicates the complex pose and body position from Image 1.
  • + Accurately blends the clothing style from Image 2 onto the body position of Image 1.
  • The transition between the neck and the scarf/shoulders is slightly messy.
  • Left hand has an unnatural finger structure and red nail polish carried over from the source image's subject.

Wan 2.7

  • + High resolution and visual clarity.
  • Completely failed the edit instruction to replace the character.
  • Ignored the character reference (Image 2) entirely.
  • The image is essentially a filtered version of Image 1 with no character change.

Verdict: GPT Image 2 successfully followed the complex multi-image prompt by placing the character from the second image into the exact pose of the first, maintaining key details like the scarf and face. Wan 2.7 failed the task entirely, simply providing a slightly modified version of the pose reference image without incorporating the requested character.

The Reversed Rodeo

Text-to-Image

“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”

GPT Image 2
Wan 2.7

AI Judge Analysis

GPT Image 2

  • + Perfectly follows the specific role-reversal instruction of the horse on top of the astronaut.
  • + High level of texture detail in the astronaut suit and horse fur.
  • + Strong surrealist composition that aligns with the prompt's tone.

Wan 2.7

  • + Excellent cinematic lighting and galactic background details.
  • + High clarity and clean rendering of the horse and astronaut.
  • Failed the core prompt instruction to have the horse on top of the astronaut (role reversal).
  • The horse and astronaut are not wearing any breathing apparatus despite being in space.

Verdict: GPT Image 2 successfully captured the difficult role-reversal request, showing a horse literally riding/sitting on an astronaut, whereas Wan 2.7 followed a generic 'astronaut on a horse' trope. While Wan 2.7 has a more vibrant background, GPT Image 2 is the superior response for its strict adherence to the surreal instruction 'horse on top, not vice versa'.

Outfit Transfer Challenge

Editing
Edit instruction

“Use Image 1 as the base person. Dress them in the exact elaborate outfit from Image 2 (including all layers, accessories, jewelry, and shoes). Carefully adapt the clothing to the body shape and pose in Image 1 while maintaining realistic fabric behavior, correct proportions, and perfect lighting/shadow matching. Keep the person’s exact face, hair, and background completely unchanged.”

Source
GPT Image 2
Wan 2.7
100% wins 0% ties 0% wins

AI Judge Analysis

GPT Image 2

  • + Perfectly replicates the coat, scarf, jeans, and accessories from Image 2.
  • + Maintains the subject's face, hair, and vitiligo patterns with high accuracy.
  • + Provides high-quality realistic fabric texture and lighting integration.
  • The scarf's positioning is slightly stiff and doesn't fully account for the lean against the post.

Wan 2.7

  • + Maintains the background and the subject's facial features and skin patterns well.
  • Completely failed to use the outfit from Image 2, substituting it with a generic fantasy costume.
  • The feet and shoes are poorly rendered and appear distorted into the sand.
  • The proportions of the body feel elongated and unnatural compared to the source.

Verdict: GPT Image 2 followed the instructions perfectly, accurately transferring the specific clothing, scarf, and watch from Image 2 onto the subject while preserving his identity. Wan 2.7 ignored the second source image entirely, generating a completely unrelated outfit and introducing anatomical distortions in the legs and feet.

The Capybara Taxi Driver

Text-to-Image

“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”

GPT Image 2
Wan 2.7
50% wins 0% ties 50% wins

AI Judge Analysis

GPT Image 2

  • + Excellent photorealistic texture on the capybara's fur
  • + Perfectly captures the 'bored' expression of the passenger in the back seat
  • + Cinematic lighting and authentic NYC night atmosphere
  • The capybara's right paw is rendered somewhat ambiguously on the wheel

Wan 2.7

  • + Strong composition showing both the taxi roof sign and interior
  • + Good details on the capybara's paws holding the wheel
  • The human passenger is sitting in the front passenger seat instead of the 'back seat' as requested
  • The lighting on the capybara feels a bit flat compared to the background

Verdict: GPT Image 2 is the superior choice because it correctly places the passenger in the back seat and captures more realistic, cinematic lighting. Wan 2.7 fails on the spatial layout of the prompt by putting the passenger in the front seat, though it does a decent job with the capybara's hands.

The Halloween Invitation

Text-to-Image

“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”

GPT Image 2
Wan 2.7

AI Judge Analysis

GPT Image 2

  • + Excellent atmospheric lighting and cinematic mood.
  • + Flawless rendering of all requested text.
  • + Highly intricate gothic border with visible thorns and skulls.
  • The parchment texture is slightly less literal than Model B.

Wan 2.7

  • + Clean, readable layout with a clear parchment aesthetic.
  • + Accurate rendering of the banner and all requested text.
  • + Follows the framing of the prompt well with twisted trees and bats.
  • The art style is more like a modern illustration than a 'vintage gothic' poster.
  • Lighting feels flat compared to the requested 'cinematic' style.

Verdict: GPT Image 2 perfectly captures the 'vintage gothic' atmosphere with moody cinematic lighting and high-quality details that make it feel like a professional poster. While Wan 2.7 followed all instructions and text accurately, its illustration style is too clean and cartoonish for the requested 'frightful' theme. GPT Image 2's sophisticated textures and depth make it the clear winner.

Isometric Miniature Diorama Scenes

Text-to-Image

“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”

GPT Image 2
Wan 2.7

AI Judge Analysis

GPT Image 2

  • + Excellent typography with a 3D effect that matches the diorama style
  • + Incredibly detailed textures for the fish, rice, and wood
  • + Perfectly follows the layout instructions with a centered diorama base
  • Includes some extra decor like the stone lantern not explicitly requested

Wan 2.7

  • + Clean, minimalist aesthetic that fits the '3D cartoon' prompt well
  • + Simple and clear composition with accurate text placement
  • + Effective soft lighting and smooth PBR materials
  • Text is flat and less integrated into the 3D scene compared to Model A
  • The sushi details are slightly more generic/plastic in appearance

Verdict: GPT Image 2 (Model A) is the clear winner as it provides a much more sophisticated rendering of the requested elements, with high-quality textures on the food and professional 3D-styled typography. While Wan 2.7 (Model B) is very clean and accurate to the layout, its visuals are less detailed and the text feels like a 2D overlay rather than part of a cohesive 3D scene.

Adorable Baby Animals in Sunny Meadow

Text-to-Image

“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”

GPT Image 2
Wan 2.7

AI Judge Analysis

GPT Image 2

  • + Excellent dynamic posing that makes the animals look like they are actively running.
  • + Highly detailed fur textures with convincing rim lighting.
  • + Accurate butterfly anatomy and vibrant color palette.
  • The fox kit has slightly strange anatomy where its left leg meets the body.

Wan 2.7

  • + Beautiful dew sparkle effects that add to the morning atmosphere.
  • + Solid variety of wildflower species as requested in the prompt.
  • + Clear, sharp focus on all four subjects.
  • The tabby kitten's anatomy is distorted around the neck and front legs.
  • The posing is more static and less 'tumbling' compared to the other image.
  • The fox kit's face looks slightly more mature than a 'baby' kit.

Verdict: GPT Image 2 captures the energy of the prompt much better, showing the animals in dynamic, playful poses that truly look like they are tumbling through the field. While Wan 2.7 has lovely atmospheric dew effects, it suffers from anatomical issues on the kitten and a more static composition.

Vintage Cafe Logo

Text-to-Image

“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”

GPT Image 2
Wan 2.7

AI Judge Analysis

GPT Image 2

  • + Excellent typography with a professional, balanced layout.
  • + Authentic vintage texture and woodcut-style engraving.
  • + Highly detailed and elegant banner and cloche design.
  • The frame/shield is slightly off-center relative to the top edge.

Wan 2.7

  • + Clean vector-style execution with a clear circular composition.
  • + Correct adherence to all text elements and the cloche dome requirement.
  • Slight misspelling in the name ('Florion' instead of 'Florian').
  • Visual style feels more like modern clip-art than a 'vintage minimalist' logo.
  • The cloche dome design is a bit simplistic and lacks the requested steam texture.

Verdict: GPT Image 2 (Model A) significantly outperforms Wan 2.7 (Model B) by delivering a sophisticated, authentic vintage aesthetic that perfectly captures the requested 'Caffè Florian' brand identity. While Wan 2.7 provides a clean vector layout, it suffers from a spelling error ('Florion') and lacks the artistic depth found in the engraving-style details of GPT Image 2.

Apollo 11: Journey to Tranquility

Text-to-Image

“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”

GPT Image 2
Wan 2.7

AI Judge Analysis

GPT Image 2

  • + Excellent typography and nearly perfect spelling throughout the entire infographic.
  • + Highly detailed and accurate illustrations of the Saturn V and Lunar Module.
  • + Strong aesthetic adhering to the NASA-inspired color palette with a professional layout.
  • Step 3 includes a miniature moon that makes the scale of the trajectory confusing.
  • The eagle in the mission patch looks slightly more like a literal eagle than the stylized mission insignia.

Wan 2.7

  • + Successfully captures a true 'flat-vector' minimalist icon style as requested.
  • + Good use of factual data points like orbital altitude and times in the supporting text.
  • Multiple spelling errors including 'Tranquiliry', 'Descript', and 'Deccent'.
  • The icons are very small and leave too much empty space in the center compared to the dense star field.
  • The NASA logo in the corner is poorly rendered and inaccurate.

Verdict: GPT Image 2 is the clear winner as it produces a professional-grade infographic with perfect text rendering and high-quality illustrations. While Wan 2.7 follows the 'flat vector' instruction more closely in style, it fails on basic legibility and spelling, whereas GPT Image 2 creates a cohesive, educational, and visually stunning poster.

Next steps

Explore each model