OpenAI's previous image generation model that accepts both text and image inputs and produces image outputs
Settled by community votes across 20 shared challenges, with an AI judge weighing in on each.
GPT Image 1
#29 of 62 in Text-to-Image
Wan 2.7
#39 of 62 in Text-to-Image
Where the votes landed
GPT Image 1
100.0%
win rate
Ties
0.0%
Wan 2.7
0.0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
GPT Image 1
- + Excellent internal reflections of the blue sphere on the bottom glass surface.
- + Clean, modern aesthetic with high resolution and no noticeable artifacts.
- + Perfect adherence to the lighting direction and object placement.
- − The glass cube looks more like a frame than a solid object in some areas, such as the top surface.
- − The blue sphere appears slightly matte and textured compared to what might be expected inside glass.
Wan 2.7
- + Realistic wood texture on the table adds a sense of tangible history.
- + Stronger depiction of the plant being clearly visible through the glass panels.
- + Convincing glass refraction and multiple reflections on the various interior facets.
- − The book appears to be merging slightly with the top of the glass cube.
- − There is a secondary blue sphere reflection on the left that looks more like a ghost object than a reflection.
Verdict: Both models followed the prompt instructions perfectly, including the specific spatial relationships of the objects. GPT Image 1 is cleaner and more aesthetically pleasing for a product-style shot, while Wan 2.7 offers superior realism in textures and complex glass refractions, despite a few minor optical inconsistencies.
Man and Car in California
Editing“Make a photo of the man driving the car down the California coastline”
AI Judge Analysis
GPT Image 1
- + Excellent preservation of the man's identity, including his hairstyle and clothing details like the scarf.
- + Very high-quality integration of the car into a dynamic California coastline environment.
- + Maintains the car's aesthetic and structural details from the source image while adding motion blur.
- − The man's scale within the car seems slightly too large or high up.
Wan 2.7
- + The scenery capture of the coastline is vibrant and aesthetically pleasing.
- + Correctly places the car on a winding coastal road with appropriate lighting.
- − Completely fails to preserve the identity or appearance of the man from the source image.
- − The person in the car is distorted and unrecognizable.
- − The car's proportions feel slightly flattened compared to the source.
Verdict: GPT Image 1 is the clear winner as it successfully merges the two source images into the requested scene while maintaining high fidelity to both. Wan 2.7 provides a nice background, but fails the editing task by replacing the specific man from the source image with a generic and distorted figure.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
GPT Image 1
- + Excellent shallow depth of field and bokeh
- + Highly detailed skin texture and facial features
- + Strong cinematic lighting and color grading
- − Cars in the background are static rather than having motion blur
- − The framing is quite polished, lacking the 'imperfect' snapshot feel
Wan 2.7
- + Captures the 'imperfect framing' and 'candid street' vibe accurately
- + Excellent wet pavement reflections
- + Authentic Japanese street setting with realistic background elements
- − Lacks the requested shallow depth of field
- − Skin texture is slightly smoother and less detailed than Image A
- − The horizontal road lines show some minor AI geometric inconsistency
Verdict: GPT Image 1 produces a more technically impressive and cinematic close-up with superior skin textures and depth of field, though it misses the motion blur requirement. Wan 2.7 better captures the 'candid street photo' atmosphere and imperfect framing requested, providing a wider perspective that feels more like real-world photography. GPT Image 1 is the winner for its superior visual quality and adherence to the character description.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
GPT Image 1
- + Excellent depiction of ornate, dark-wash engraved plate armor.
- + Superior lighting and color palette that feels cinematic and moody.
- + Highly realistic skin textures with believable dirt and faint scarring.
- − The beads in the hair are few and quite dark, making them hard to see.
- − Leather straps and cloth underlayers are mostly obscured by the tight framing.
Wan 2.7
- + Perfect adherence to the 'beads in braids' and 'leather straps' prompts.
- + Clearly visible scars that look more 'battle-worn' than Model A.
- + Stronger contrast on the metal surfaces showing clear torchlight reflections.
- − The bokeh circles in the background are a bit distracting and look digitally inserted.
- − The skin texture on the forehead and the beard rendering feel slightly more artificial compared to Model A.
Verdict: GPT Image 1 creates a far more atmospheric and high-fidelity cinematic shot with incredible detail in the armor engravings and lighting. However, Wan 2.7 followed the specific technical details of the prompt better, specifically including the leather straps, multiple beaded braids, and the visible torch in the background. While Wan 2.7 is more accurate to the checklist, GPT Image 1 is the superior visual piece due to its lifelike eyes and professional-grade composition.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
GPT Image 1
- + Excellent font legibility for headers and prices
- + High-quality, appetizing food photography
- + Clean layout that closely follows the grid request
- − Nonsense placeholder text in descriptions
- − Missing 'Mains' section header despite having a main course item
Wan 2.7
- + Sophisticated and complete branding including logo and contact info
- + Strong adherence to the grid layout with multiple food categories
- + Realistic overall composition including lifestyle props
- − Text is very small and becomes illegible at the bottom
- − Several spelling errors in menu item names like 'Calanrfri' and 'Bisge'
Verdict: GPT Image 1 produces a very clean, professional-looking fragment of a menu with superior image quality and font rendering. However, Wan 2.1 delivers a much more comprehensive and creative design that feels like a finished marketing piece, despite having more spelling artifacts in the small text. Wan 2.1 is the winner for its better interpretation of a 'modern minimalist design' as a cohesive whole.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
GPT Image 1
- + Excellent photorealistic texture on the bun and patty
- + Clean and highly legible glowing typography
- + Sophisticated composition with realistic lighting from embers
- − Missed the '6' in the price starburst rendering it as €.99
- − Wait time for the burger is less dynamic, feeling more stacked than 'exploded'
Wan 2.7
- + Successfully captured the 'exploded' dynamic motion with flying ingredients
- + Correctly included all requested text elements including the price €6.99
- + Vibrant and high-energy illustrative style
- − The 'Magic Burger' text has a busy, clip-art feel with literal flames on top
- − Less photorealistic rendering compared to the competitor
- − Strange floating nuts or seeds on the left side that weren't in the prompt
Verdict: GPT Image 1 offers superior photorealism and a more professional aesthetic suitable for high-end advertising, though it fails on the specific price text. Wan 2.7 follows every prompt instruction perfectly, including the explosion effect and full text, but results in a more chaotic and slightly less realistic image. GPT Image 1 is preferred for visual quality, while Wan 2.7 is preferred for literal prompt accuracy.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
GPT Image 1
- + Text rendering is perfectly accurate to the requested items
- + Genuine chalk texture with realistic grain and smudging
- + Successful variation in letter size and slant as requested
- − Failed to provide 'elegant cursive' for the title
- − The composition is a bit tight with the bottom text cutoff
- − Simpler background lacks the requested 'cozy café' atmosphere
Wan 2.7
- + Excellent 'cozy café' atmosphere with lighting and decor
- + High visual quality and clear layout
- + Text is stylistically consistent across the entire board
- − Text looks like a digital font rather than natural chalk handwriting
- − Title is not in cursive as requested
- − Minor artifacts near the board edges
Verdict: GPT Image 1 followed the technical requirements of the prompt much more effectively, delivering a texture that truly looks like hand-drawn chalk with natural variations. While Wan 2.7 created a much more beautiful and atmospheric scene, its text looks like a clean digital overlay/font, failing the primary instruction for a realistic handwritten style without digital fonts.
Pose & Character Mashup
Editing“Use Image 1 as the exact pose reference and Image 2 as the character reference. Recreate the person/character from Image 2 in the exact dynamic pose and body position from Image 1. Keep the exact face, hair, clothing style/details, and expression from Image 2. Match the lighting and environment of Image 1. The final image must show the character from Image 2 performing the precise action/pose from Image 1 with perfect anatomy and natural integration.”
AI Judge Analysis
GPT Image 1
- + Successfully integrated clothing elements like the patterned scarf and black sunglasses from Image 2.
- + Followed the complex pose from Image 1 with reasonable body positioning.
- + Accurately matched the yellow background and red ottoman from the source environment.
- − The facial likeness to the man in Image 2 is poor and looks like a caricature.
- − Anatomy of the feet and legs is distorted and lacks the clarity of the source image.
- − The hand positioning is awkward and does not match the graceful gesture in Image 1.
Wan 2.7
- + Perfectly preserved the source image 1 without any degradation.
- − Completely failed to perform the edit, ignoring Image 2 entirely.
- − Did not change the character, clothing, or facial features as requested.
Verdict: GPT Image 1 attempted the complex task of merging the character and clothing from Image 2 into the pose of Image 1, resulting in a new, albeit flawed, image. Wan 2.7 failed the task entirely by returning the original pose reference image with no changes. GPT Image 1 is the winner by default for attempting the instruction, despite the low visual quality and poor facial likeness.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
GPT Image 1
- + Excellent cinematic lighting and dark, moody atmosphere
- + Dynamic pose of the horse creates a stronger sense of movement in space
- + Consistent textures on the spacesuit and horse leather
- − The horse's front legs have slightly distorted hoof anatomy
- − The reins appear to blend into the horse's neck in an unrealistic way
Wan 2.7
- + Bright, clear visualization of galaxies and planets in the background
- + Very clean anatomical rendering of the horse's legs and hooves
- + High level of detail on the astronaut's suit and gloves
- − The horse is much Larger than the Earth below, which feels less 'cinematic' and more like a collage
- − Lighting is a bit flat compared to the dark void of space
Verdict: Both models followed the 'horse on top' instruction perfectly. GPT Image 1 is preferred for its superior cinematic quality and atmospheric lighting, whereas Wan 2.7 feels a bit more like a traditional digital illustration with flatter lighting, despite having cleaner anatomy.
Outfit Transfer Challenge
Editing“Use Image 1 as the base person. Dress them in the exact elaborate outfit from Image 2 (including all layers, accessories, jewelry, and shoes). Carefully adapt the clothing to the body shape and pose in Image 1 while maintaining realistic fabric behavior, correct proportions, and perfect lighting/shadow matching. Keep the person’s exact face, hair, and background completely unchanged.”
AI Judge Analysis
GPT Image 1
- + Perfectly replicates the specific clothing items from Image 2, including the pea coat, plaid scarf, and watch.
- + Preserves the subject's facial features and vitiligo patterns with high accuracy.
- + Successfully integrates the clothing into the lighting and atmosphere of the original beach scene.
- − The scale of the subject in the frame is slightly changed from the original wide shot to a medium-up shot.
Wan 2.7
- + Maintains the original full-body composition and background environment.
- + Correctly applies vitiligo patterns to the subject's visible hands.
- − Completely failed the instruction to use the outfit from Image 2, instead generating a generic gothic/regal ensemble.
- − The face is slightly altered and less recognizable than the original person.
Verdict: GPT Image 1 followed the instructions nearly perfectly, correctly identifying and transferring the specific pea coat and plaid scarf from the reference image while maintaining the subject's identity. In contrast, Wan 2.7 generated an entirely different outfit that was not present in any of the source images, failing the prompt's primary requirement.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
GPT Image 1
- + Excellent photorealistic texture on the capybara fur and jacket
- + Great cinematic lighting and shallow depth of field
- + Accurately places the passenger in the back seat as requested
- − The text on the cap is a bit generic
- − The composition is quite tight, showing less of the car interior
Wan 2.7
- + Sharp level of detail across the entire frame
- + Good depiction of the Manhattan background and taxi exterior elements
- − The passenger is placed in the front seat instead of the back seat
- − The capybara's fur has a slightly synthetic, needle-like appearance
- − The anatomical transition between the capybara head and the human-like torso is less seamless
Verdict: GPT Image 1 followed the instructions more accurately by placing the passenger in the back seat and captured a much more convincing photorealistic aesthetic. Wan 2.7 failed on spatial placement by putting the passenger in the passenger seat and has a more 'digitally rendered' look compared to the cinematic quality of GPT Image 1.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
GPT Image 1
- + Excellent moody, cinematic lighting that fits the 'gothic' and 'fright' theme
- + High-quality text rendering with elegant gothic typography
- + Clean and atmospheric composition that feels like a professional poster
- − Confused the 'Time' and 'Location' fields at the bottom
- − The parchment is extremely dark, making some of the background details hard to see
Wan 2.7
- + Perfect text accuracy, including all requested event details and many creative additions
- + Extremely intricate border with webs and thorns as requested
- + Stronger 'vintage parchment' aesthetic with high contrast between elements
- − The illustration style is more 'cartoonish' than 'cinematic'
- − The composition is a bit cluttered with many competing elements
Verdict: GPT Image 1 captures the cinematic and moody atmospheric requirements much better than Wan 2.7, though it makes a minor error by merging the time and location text. Wan 2.7 is technically superior in prompt adherence regarding text accuracy and decorative elements like webs and thorns, but its illustrative style feels less like a gothic invitation and more like a storybook page.
Bald man challenge
Image Editing“Give the person a full, thick head of natural hair with realistic texture, density, and a natural hairline. Preserve facial features and lighting.”
AI Judge Analysis
GPT Image 1
- + Successfully adds a significant amount of thick hair.
- + Retains the background and clothing perfectly.
- − The hairline is very high and unnatural, resembling a wig.
- − Alters the shape of the face and eyebrows significantly, losing the original likeness.
- − The hair texture is somewhat fuzzy and lacks realistic sheen.
Wan 2.7
- + Excellent realism in hair texture and flow.
- + Preserves the subject's facial features and identity much better than GPT Image 1.
- + Natural-looking hairline that blends well with the existing temples and sideburns.
- − The volume of hair is slightly less 'full' compared to the other model, though more realistic.
- − Slight softening of the skin texture compared to the original image.
Verdict: Wan 2.7 is the clear winner as it provides a realistic hairstyle that matches the lighting and texture of the original image while preserving the subject's identity. GPT Image 1 adds more volume, but the result looks like a pasted-on wig and alters the person's face too much, failing the 'preserve facial features' instruction.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
GPT Image 1
- + Excellent typography rendering with clean, centered alignment.
- + Great 'cartoon' but 'PBR' balance with soft, appealing textures.
- + Simpler, cleaner composition that emphasizes the diorama feel.
- − The raised diorama base lacks a plate-like rim, appearing more like a block.
- − The flag icon is slightly simple and lacks the polished feel of the text.
Wan 2.7
- + Features a wider variety of sushi types and more detailed garnish like soy sauce and napkins.
- + Stronger 3D depth in the text and flag element.
- + Beautifully rendered materials with realistic specular highlights on the fish.
- − The text is not centered as requested, skewed to the top right.
- − The diorama base has a slight artifact where the gold trim meets the corner.
Verdict: Both models followed the prompt closely, creating high-quality isometric dioramas. GPT Image 1 followed the layout instructions more accurately, keeping everything perfectly centered with ultra-clean text, whereas Wan 2.7 produced more detailed food models and lighting but failed on the requested centering of the text.
Over-the-top cartoon caricature
Editing“Create a caricature of me and my job. Make it exaggerated and humorous, incorporating my profession as a tv show anchor and my love for dogs and hockey.”
AI Judge Analysis
GPT Image 1
- + Excellent caricature style with watercolor textures and exaggerated facial features.
- + Strong thematic integration with a news desk, anchor papers, and a dog hockey player on the 'over-the-shoulder' graphic.
- + Maintains character likeness while successfully transforming her into a humorous illustration.
- − The transition between her body and the news desk is slightly cluttered with the hockey stick placement.
- − Does not preserve the original 'selfie' pose, opting for a seated anchor position instead.
Wan 2.7
- + Successfully incorporates all elements including multiple dogs, a hockey rink, and news gear.
- + Preserves the 'selfie' arm pose from the original source image.
- + Captures a vibrant, modern digital comic art style.
- − Contains spelling errors in speech bubbles ('Dogs Rolee!').
- − The composition is quite cluttered with overlapping elements and floating text bubbles.
- − The hockey integration feels a bit disjointed with a tiny floating rink and a dog in a helmet.
Verdict: GPT Image 1 (Model A) provides a much higher quality caricature that feels like a professional illustration, cleverly integrating the hockey and dog themes into the newsroom setting. While Wan 2.7 (Model B) attempts to keep the original selfie pose, it suffers from spelling errors and a cluttered composition that lacks the artistic cohesion of its competitor.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
GPT Image 1
- + Excellent dynamic composition showing joyful movement and interaction
- + Captures beautiful 'god rays' and warm golden lighting perfectly
- + Strong emotional expressions on all the animals' faces
- − Anatomy on the fox's front right leg is slightly distorted and blurry
Wan 2.7
- + Features a diverse variety of colorful wildflowers as requested
- + Detailed fur textures and realistic dew sparkles in the grass
- + Includes all requested species clearly within the frame
- − The fox's face looks slightly older and less 'kit-like' compared to the others
- − The composition feels a bit more static and posed rather than tumbling or playful
Verdict: GPT Image 1 captures the 'playful' and 'tumbling' aspect of the prompt much better than Wan 2.7, with dynamic posing and superior atmospheric lighting that truly evokes a sunrise. While Wan 2.7 has lovely detail in the flowers and dew, GPT Image 1 creates a more cohesive and emotionally resonant masterpiece that aligns with the requested 'wholesome vibe'.
Studio Ghibli Anime Style
Editing“Transform this photo into a Studio Ghibli–inspired illustration. Use soft pastel colors, hand-painted textures, gentle lighting, dreamy backgrounds, and a warm, nostalgic mood”
AI Judge Analysis
GPT Image 1
- + Perfectly captures the Studio Ghibli art style including eye shapes and color palette
- + Atmospheric use of soft textures and warm, nostalgic lighting
- + Strong adherence to the 'hand-painted' instruction
- − Faces are generalized and lose the specific likeness of the people in the original meme
- − Pattern on the man's shirt is simplified into lines rather than a plaid pattern
Wan 2.7
- + Excellent preservation of the original subjects' faces and expressions
- + Accurately replicates the plaid pattern of the man's shirt
- + Creates a clean watercolor-inspired aesthetic
- − Art style is more of a westernized watercolor sketch than Studio Ghibli
- − Lighting is flat compared to the dreamy atmosphere requested
Verdict: GPT Image 1 (Model A) successfully executes the 'Studio Ghibli' transformation by adopting specific character design tropes and a warm, nostalgic atmosphere, though it loses some character likeness. Wan 2.7 (Model B) preserves the source image much better but fails to move beyond a standard watercolor filter, missing the specific Ghibli aesthetic requested in the prompt.
Golden Hour Stroll
Image Editing“Add dynamic motion to this photo: make hair blow in the wind, add leaves flying, energetic and lively feel.”
AI Judge Analysis
GPT Image 1
- + Excellent depiction of wind-blown hair with high energy
- + Added a motion-appropriate tail pose for the dog
- + High density of falling leaves consistent with the prompt
- − Visible artifacts where the leash meets the woman's hand
- − Changed the background bridge and some foliage details incorrectly
Wan 2.7
- + Strong source preservation of the background and original character features
- + Smooth and realistic hair motion
- + Subtle but effective motion blur on the falling leaves
- − The leash now appears to float or intersect poorly with the woman's hand
- − Fewer leaves compared to Model A, feeling slightly less energetic
Verdict: Both models followed the instructions well, but GPT Image 1 captured a more 'dynamic and lively' feel by significantly altering the hair and dog's tail to match the wind. However, Wan 2.1 did a much better job of preserving the specific details of the background and the woman's face, whereas GPT Image 1 introduced some minor structural artifacts and changed background elements.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
GPT Image 1
- + Clean vector aesthetic suitable for a modern minimalist logo
- + Accurate spelling of the brand name and date
- − Failed to provide a light background as requested
- − Minimalist style resulted in a very basic cloche design
Wan 2.7
- + Successfully applied the requested warm brown/cream tones and light background
- + Excellent vintage character with subtle paper texture
- + Clear and accurate text rendering for 'Est. 1720'
- − Spelling error in the main brand name ('Florion' instead of 'Florian')
- − Cloche dome rendering is a bit complex for a strictly minimalist logo
Verdict: Wan 2.7 followed nearly all stylistic prompts including color palette and background texture, creating a much more 'vintage' feel, though it suffered a minor spelling error. GPT Image 1 followed the vector and typography requirements well but ignored the background color request, resulting in a stark high-contrast image that lacks the requested warmth.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
GPT Image 1
- + Features clean and legible vector-style icons
- + Maintains a consistent and accurate NASA-inspired color palette
- + The layout is bold and fits the 'infographic poster' style well
- − Several spelling errors like 'EARLLUNAR'
- − Missing specific icon steps such as Lunar Orbit
- − The text alignment is disorganized and lacks a clear flow
Wan 2.7
- + Excellent vertical infographic flow that logically follows the mission steps
- + High level of detail with supporting text for each stage
- + Superior composition with a clear title and professional margins
- − Minor spelling errors in smaller text like 'Tranquiliry' and 'Descript'
- − The astronaut icons are very generic compared to the rest of the graphics
Verdict: Wan 2.7 is the clear winner as it successfully follows the requested infographic structure with a logical vertical timeline and all required steps. GPT Image 1 feels like a randomized collection of icons with confusing text placement and missing steps, whereas Wan 2.7 captures the professional aesthetic of an actual NASA educational poster.
Explore each model
Alibaba's Wan 2.7 image generation and editing model for text-to-image, reference-guided generation, and instruction-based image edits