OpenAI's state-of-the-art image generation model with arbitrary resolution up to 4K and strong instruction following
Settled by community votes across 15 shared challenges, with an AI judge weighing in on each.
GPT Image 2
#3 of 62 in Text-to-Image
Wan 2.7 Pro
#38 of 62 in Text-to-Image
Where the votes landed
GPT Image 2
100.0%
win rate
Ties
0.0%
Wan 2.7 Pro
0.0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
GPT Image 2
- + Excellent adherence to lighting instructions with a clear left-side window source
- + Materials look highly realistic, especially the texture of the red book cover
- + Very clean geometry on the glass cube
- − The plant is quite close to the cube, making the 'behind' and 'through' effect less pronounced than Model B
Wan 2.7 Pro
- + Successfully shows the plant clearly through the glass walls of the cube
- + Highly detailed wooden table texture with realistic weathering
- + Accurate sphere reflection on the bottom glass panel
- − The glass cube has some structural inconsistencies in its vertical edges
- − The lighting feels a bit flatter compared to the distinct window light in Model A
Verdict: Both models followed the spatial instructions perfectly. GPT Image 2 is the winner due to its superior lighting and cleaner geometry of the glass cube, whereas Wan 2.7 Pro had slight artifacts in the glass structure and less convincing light direction.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
GPT Image 2
- + Excellent adherence to the 'imperfect framing' and 'candid' aspects of the prompt
- + Captures motion blur from passing cars very effectively
- + Natural skin texture and realistic, weathered blue jacket
- − The bicycle frame geometry is slightly warped where it connects to the wheel
- − The background bokeh has some chunky digital artifacts
Wan 2.7 Pro
- + Beautiful, high-quality reflections on the wet pavement
- + The subject and bicycle are very clearly defined with great detail
- + Excellent rendering of light rain through vertical streaks
- − Failed to include the requested motion blur on the cars
- − Composition is very centered and clean, missing the 'imperfect framing' instruction
- − The bicycle doesn't actually have a kickstand or support, yet stands upright on its own
Verdict: GPT Image 2 followed the specific stylistic cues of the prompt much better, successfully incorporating the motion blur, 50mm feel, and imperfect 'candid' framing. While Wan 2.7 Pro produced a sharper and more aesthetically pleasing image with better rain effects, it ignored the motion blur and framing requirements, resulting in a more 'posed' look.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
GPT Image 2
- + Sublime metal textures and intricate engraving detail
- + High-quality skin texture and realistic, lifelike eyes
- + Excellent cinematic lighting and composition
- − Beads in hair are a bit subtle compared to the prompt
- − Leather strap detail is less prominent than in the competitor
Wan 2.7 Pro
- + Clearer depiction of 'small beads' in braids
- + Strong bokeh sparks and visible torch light sources
- + Distinct scars as requested in the prompt
- − Slightly less realistic skin texture compared to Model A
- − The armor engraving feels a bit more generic or stamped on
Verdict: Both models followed the prompt exceptionally well. GPT Image 2 (Model A) provides a more artistic and technically high-fidelity portrait with superior metal micro-textures and lighting, while Wan 2.7 Pro (Model B) followed the specific hair bead and scar instructions more literally. GPT Image 2 is the winner due to its overall cinematic realism and superior visual quality.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
GPT Image 2
- + Excellent text legibility and correct spelling across multiple menu items.
- + Clear adherence to the specific section requests for appetizers, pizza, and mains.
- + Professional, well-balanced distribution of food photography and whitespace.
- − The 'Mains' category features a mix of plate styles that aren't perfectly uniform.
- − Some slight repetitions in descriptions, though typical for AI.
Wan 2.7 Pro
- + Aesthetically pleasing layout with a nice top-down presentation and props.
- + Successful use of vibrant accents and a clean grid system.
- + Good variety in the food photography shown.
- − Significant text issues including misspellings like 'Garlic Herb Bread' being written as 'Garfic' and 'Calamari' as 'Calariri'.
- − Failed to properly implement the specific requested sections (Appetizers/Pizza/Mains) into the actual body of the list, mixing them all together.
Verdict: GPT Image 2 is the clear winner as it provides a fully functional, professional-grade menu with legible text and accurate categorized sections. Wan 2.7 Pro produces a visually attractive mock-up, but the text is riddled with typos and it ignores the prompt's request for specific section layouts in the body of the menu.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
GPT Image 2
- + Perfect adherence to all text requirements including price and secondary message
- + Superior photorealistic detail in the textures of the meat and bun
- + Excellent sense of dynamic motion through sauce splashes and embers
- − The composition is a bit crowded with large text overlapping embers
Wan 2.7 Pro
- + Clean 'exploded' view that clearly separates components
- + Attractive glowing effect on the main title text
- − Failed to include 'LIMITED TIME ONLY' and '€6.99' text
- − Lighting on bottom bun is inconsistent with the fiery theme
- − The background and elements feel more like a digital illustration than photorealistic
Verdict: GPT Image 2 followed the prompt instructions perfectly, including all requested text and achieving a highly photorealistic, intense aesthetic suitable for a fast-food advertisement. Wan 2.7 Pro missed two-thirds of the required text and produced a simpler, more illustrative image that lacked the 'fiery' atmosphere requested.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
GPT Image 2
- + Excellent adherence to the 'handwritten chalk' requirement with realistic texture and spacing.
- + Correctly rendered all requested text including the truncated line from the prompt.
- + Perfectly captures the aesthetic of a cozy café with a natural, messy chalk board.
- − The date '2026' is slightly less clean than other characters but still legible.
Wan 2.7 Pro
- + High resolution and very clean visual presentation.
- + The chalkboard background has realistic smudge marks and decorative doodles.
- − Text appears to be a digital font rather than natural chalk handwriting.
- − Repeats the phrase '& Herbs - $28' twice on the same line, creating a layout error.
- − Fails the 'no printed or digital fonts' requirement, as letters of the same type are identical.
Verdict: GPT Image 2 followed the prompt instructions much more accurately, especially regarding the 'handwritten' requirement which Wan 2.7 Pro failed by using a uniform digital font. GPT Image 2 also correctly handled the text content without the repetition errors seen in the Wan 2.7 Pro output.
Pose & Character Mashup
Editing“Use Image 1 as the exact pose reference and Image 2 as the character reference. Recreate the person/character from Image 2 in the exact dynamic pose and body position from Image 1. Keep the exact face, hair, clothing style/details, and expression from Image 2. Match the lighting and environment of Image 1. The final image must show the character from Image 2 performing the precise action/pose from Image 1 with perfect anatomy and natural integration.”
AI Judge Analysis
GPT Image 2
- + Excellent adherence to the character reference, including face, hair, and clothing.
- + Perfectly replicates the complex pose from Image 1.
- + Successfully combines the character's black clothing and scarf with the source image's yellow environment.
- − The scarf's physics are slightly stiff for the downward angle.
- − The fingernails on the raised hand retain red polish from the original model in Image 1.
Wan 2.7 Pro
- + Perfectly preserves the source image 1's content.
- − Failed the editing task entirely by providing the original source image 1.
- − Did not incorporate the character or clothing from Image 2.
Verdict: GPT Image 2 successfully completed the image editing task by accurately transposing the character from Image 2 into the specific dynamic pose of Image 1. Wan 2.7 Pro failed to perform any edits, returning a slightly cropped version of the original source image.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
GPT Image 2
- + Excellent adherence to the 'horse on top' instruction
- + Highly detailed textures on the spacesuit and lunar ground
- + Imaginative and surreal interpretation of the prompt
- − The horse's front legs and harness are a bit anatomically confused where they meet the saddle
Wan 2.7 Pro
- + Clean, cinematic composition
- + Good rendering of the horse and astronaut textures
- + Playful use of miniatures planets in the background
- − Completely failed the negative constraint to put the horse on top
- − Common AI artifacting where the horse's back leg blends into the stomach
Verdict: GPT Image 2 followed the complex negative constraint perfectly, depicting a horse literally riding an astronaut. Wan 2.7 Pro ignored the specific instruction 'horse on top, not vice versa' and generated a standard astronaut riding a horse, making it a failure in prompt adherence.
Outfit Transfer Challenge
Editing“Use Image 1 as the base person. Dress them in the exact elaborate outfit from Image 2 (including all layers, accessories, jewelry, and shoes). Carefully adapt the clothing to the body shape and pose in Image 1 while maintaining realistic fabric behavior, correct proportions, and perfect lighting/shadow matching. Keep the person’s exact face, hair, and background completely unchanged.”
AI Judge Analysis
GPT Image 2
- + Excellent adherence to the clothing style, colors, and patterns from Image 2.
- + Maintains the person's face, hair, and sand texture perfectly.
- + Correctly adapts the pose and shadows to the new outfit.
- − Cropped the image closely compared to the original aspect ratio.
- − Failed to include the watch and ring visible in Image 2.
Wan 2.7 Pro
- + Successfully integrated a full-body view onto the original background.
- + Maintains the vitiligo pattern on the hands and chest well.
- − Failed to use the outfit from Image 2, instead generating a generic gold-patterned jacket.
- − The scarf and specific pea coat from Image 2 are completely missing.
Verdict: GPT Image 2 is the clear winner as it followed the specific instructions to use the exact outfit from Image 2, capturing the blue pea coat and plaid scarf with high fidelity. Wan 2.1 Pro failed the core task by generating a completely different, unrelated gold jacket and black trousers, ignoring the visual reference for the clothing entirely.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
GPT Image 2
- + Excellent adherence to the 'back seat' and 'looking at her phone' details for the passenger.
- + Highly realistic texture on the capybara fur and the driver's jacket.
- + Strong cinematic composition that feels like a real photograph from inside the car.
- − The capybara's paw on the steering wheel looks slightly distorted.
- − The lighting on the passenger is a bit dark, making it harder to see her expression.
Wan 2.7 Pro
- + Very clear and professional lighting across the entire scene.
- + The capybara's 'calm, professional expression' is well-captured.
- + Solid rendering of both paws on the steering wheel.
- − The passenger is sitting in the front seat next to the driver, failing the 'back seat' instruction.
- − The passenger is not looking at her phone as requested.
- − Perspective makes the capybara look unusually small in the driver's seat.
Verdict: GPT Image 2 followed the prompt much more accurately, correctly placing the passenger in the back seat and showing her looking at her phone, whereas Wan 2.7 Pro placed her in the front passenger seat. GPT Image 2 also achieved a more authentic 'night taxi' atmosphere with a realistic internal partition and street bokeh.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
GPT Image 2
- + Excellent typography with perfect adherence to the requested text and formatting.
- + Highly detailed and moody atmosphere with realistic lighting and texture.
- + Superior composition that integrates the background elements and border seamlessly into the vintage aesthetic.
- − The parchment texture makes the lower text secondary, though still legible.
Wan 2.7 Pro
- + Successfully includes all requested text with clear legibility.
- + Clean, vector-like illustration style that feels professional as a digital card.
- + Good use of the thorny border and rose motifs.
- − Lighting is flat and less 'cinematic' than requested.
- − The pumpkin and trees look more like modern cartoons than 'vintage gothic'.
- − The scroll banner looks slightly awkward hanging from thin wires.
Verdict: GPT Image 2 (Model A) significantly outperforms Wan 2.7 Pro in terms of atmosphere, detail, and prompt adherence to the 'vintage gothic' and 'cinematic' keywords. While both models handled the complex text requirements perfectly, GPT Image 2's artistic execution feels authentic to the requested style, whereas Wan 2.7 Pro's result looks like a standard modern digital illustration.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
GPT Image 2
- + Excellent text rendering and layout centered vertically.
- + High-quality textures and materials, particularly on the fish and wood.
- + Strong miniature diorama feel with the stone base and lantern.
- − The diorama base is quite complex with many additional elements not requested like the lantern and bushes.
- − The sushi pieces feel slightly crowded on the plate.
Wan 2.7 Pro
- + Perfect adherence to the 'minimal garnish' and 'soft refined textures' request.
- + Very clean, professional infographic-style layout.
- + Accurate sushi anatomy and high-clarity 3D rendering.
- − The text is off-center to the left, violating the 'top-center' and 'perfectly centered' instructions.
- − Small floating artifacts/particles in the top corners and bottom-right corners.
Verdict: GPT Image 2 (Model A) creates a more immersive 3D diorama with impressive textures and perfect text centering. However, Wan 2.7 Pro (Model B) better captures the requested 'minimal' and 'refined' aesthetic, even though it fails to center the text and leaves small floating artifacts in the corners. Overall, Model A is the winner for its superior composition and technical execution of the complex wood and stone materials.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
GPT Image 2
- + Excellent dynamic motion and energy in the poses
- + Vibrant colors and natural-looking backlight with god rays
- + High facial detail and expressive eyes across all four animals
- − The fox's front right paw is a bit blurry and anatomically confusing
Wan 2.7 Pro
- + Strong adherence to all subject requirements
- + Clear rendering of fur textures with good backlighting
- + Whimsical interactions between the animals
- − The kitten has an anatomical error with its tail appearing to sprout from the puppy's back or as a fifth leg
- − The fox has an extra limb/fragmented leg visible between it and the kitten
- − Perspective of the kitten's placement feels slightly flat
Verdict: GPT Image 2 is much more successful as it captures the 'tumbling' and 'chasing' energy of the prompt with better anatomical coherence. While Wan 2.7 Pro features nice lighting and textures, it suffers from significant AI artifacts where the animals overlap, resulting in extra limbs and confusing anatomy.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
GPT Image 2
- + Excellent typography with correct spelling and beautiful classic serif styling.
- + Superior execution of the vintage paper texture and stippled shading on the cloche.
- + Highly professional layout and visual balance appropriate for a formal logo.
- − None identified based on the prompt.
Wan 2.7 Pro
- + Includes additional thematic elements like 'Venezia Italia' and coffee beans in the background.
- + Strong vector-style circular emblem layout.
- − Spelling error in the main brand name, rendering it as 'Florion'.
- − Lower quality linework on the cloche and steam compared to Model A.
- − Clutter in corners detracts from the minimalism requested in the prompt.
Verdict: GPT Image 2 is the clear winner as it perfectly follows all instructions, including correct spelling and a sophisticated vintage aesthetic. Wan 2.7 Pro fails on the primary text by spelling the name incorrectly and includes distracting elements that violate the 'minimalist' requirement.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
GPT Image 2
- + Excellent typography with clean, readable, and mostly accurate text.
- + High-quality, detailed vector illustrations for each mission stage.
- + Includes great additional details like the mission patch and crew names in a professional layout.
- − The style leans more toward high-detail digital illustration than the requested 'flat-vector' style.
- − The landing site graphic features a misspelling ('TRANQUILITY' instead of 'Tranquillity' as per NASA convention, though both are used) and the eagle on the patch looks a bit like a photograph.
Wan 2.7 Pro
- + Perfect adherence to the 'flat-vector' style and icon-heavy layout requested.
- + Very clean use of the NASA color palette (navy, white, red, light gray).
- + Includes specific mission data like burn durations and altitudes, adding to the infographic feel.
- − Significant text rendering issues, including 'DESCRIPT' instead of 'Descent' and 'Tramquilicy'.
- − The icons are very small and some are overly simplified compared to the rest of the poster.
- − The NASA logo in the corner is distorted.
Verdict: GPT Image 2 is the superior overall image due to its professional composition, clear text, and high-quality illustrations, though it pushes the boundaries of 'flat' style. Wan 2.1 Pro nailed the requested flat aesthetic perfectly but suffered from several spelling errors and poor rendering of smaller details like the NASA logo.
Explore each model
Alibaba's Wan 2.7 Pro image generation and editing model with higher-quality outputs and support for 4K image generation