OpenAI's previous image generation model that accepts both text and image inputs and produces image outputs
Settled by community votes across 19 shared challenges, with an AI judge weighing in on each.
GPT Image 1
#32 of 62 in Text-to-Image
Wan 2.6
#28 of 62 in Text-to-Image
Where the votes landed
GPT Image 1
0%
win rate
Ties
0%
Wan 2.6
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
GPT Image 1
- + Perfect compliance with spatial positioning for all objects.
- + Clean, minimalist aesthetic with high resolution.
- + Excellent rendering of light through translucent curtains.
- − The sphere appears more matte than expected for a 'blue sphere' contrast to the glass.
- − The glass cube edges appear slightly impossible in their thickness/joining.
Wan 2.6
- + Highly realistic textures on the vintage book and weathered wooden table.
- + Beautiful, complex light interactions including caustics and reflections.
- + Very natural-looking plant integration.
- − The plant is more 'next to' the cube than 'behind' it as seen through the glass.
- − Small artifacts in the reflection on the table surface.
Verdict: Both models followed the prompt exceptionally well, but GPT Image 1 achieved a cleaner, more deliberate composition that felt closer to a studio photograph. Wan 2.6 provided superior textures and lighting realism, particularly the weathered look of the book and the warm window sunlight, but missed the subtle 'behind the cube' positioning of the plant.
Man and Car in California
Editing“Make a photo of the man driving the car down the California coastline”
AI Judge Analysis
GPT Image 1
- + Excellent preservation of the car's model and specific design details
- + Great lighting integration between the man, the car, and the environment
- + Realistic camera angle and motion blur on the wheels
- − The man's facial features and dreadlock style are slightly altered from the source
- − The road texture in the foreground is a bit grainy
Wan 2.6
- + Successfully maintains the man's facial likeness and specific plaid coat from the source
- + Vibrant and scenic California coastline background
- + Accurately places the driver in a right-hand drive configuration consistent with the console edit
- − The car's proportions are distorted, making it look much longer and narrower than a Rolls-Royce
- − Loss of detail on the car's front grille and hood area compared to the source
Verdict: Both models performed well on a complex multi-subject editing task. GPT Image 1 is superior in preserving the technical details of the car and creating a more cohesive filmic look, whereas Wan 2.6 did a much better job preserving the specific clothing and facial features of the man from the source image. GPT Image 1 is the likely winner because the distortion of the car in Wan 2.6's output is significant.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
GPT Image 1
- + Excellent skin and hair texture realism
- + Captures a narrow depth of field and bokeh realistically
- + Composed with a strong cinematic focus on the subject's face
- − Lacks the requested motion blur from passing cars; background traffic is static and out of focus
- − The bicycle chain and wheel spokes are simplified and lack structural detail
Wan 2.6
- + Stronger environmental storytelling with tools and a wet street aesthetic
- + Better representation of 'light rain' through visible droplets on skin and clothing
- + Captures the 'imperfect framing' requested in the prompt
- − Anatomical errors in the hands with cluttered, merged fingers
- − The rain droplets on the jacket look like static plastic beads rather than liquid
- − The bicycle geometry is slightly warped at the front wheel
Verdict: GPT Image 1 produces a more aesthetically pleasing portrait with superior skin texture and realistic blurring, whereas Wan 2.6 better incorporates the specific environmental elements like rain droplets and street tools. However, GPT Image 1 is the preferred overall image because Wan 2.6 suffers from significant anatomical distortions in the subject's hands and unnatural-looking rain artifacts.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
GPT Image 1
- + Exquisite engraving detail on the plate armor
- + Highly realistic skin textures with subtle scaring and grime
- + Excellent atmospheric lighting that feels naturally integrated
- − The beads in the hair are very subtle and blend in too much
- − Less variation in the underlayer textures compared to model B
Wan 2.6
- + Strong adherence to the 'beads' prompt with clear, vibrant decorative elements
- + Excellent contrast between the plate, leather straps, and frayed cloth underlayers
- + Dynamic lighting and very expressive, lifelike eyes
- − The dirt on the face looks slightly digital or 'painted on' in patches
- − The torch in the background is a bit distracting and lessens the 'shallow depth of field' effect on the sparks
Verdict: GPT Image 1 excels in realistic skin rendering and the intricate, grounded texture of the metalwork, feeling very cinematic. Wan 2.6 provides a more direct interpretation of all prompt elements, specifically the beads and leather straps, though the facial grime is less naturally integrated than the first image. Both models deliver high-quality results, but Wan 2.6 is preferred for its superior attention to the specific decorative and structural elements requested.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
GPT Image 1
- + Excellent typography rendering with almost perfect spelling
- + High-quality, appetizing food photography
- + Clean layout that closely follows the minimalist prompt
- − The grid is slightly cut off at the bottom
- − Only shows Appetizers and Pizza sections clearly, missing a distinct 'Mains' heading
Wan 2.6
- + Includes all requested headings: Appetizers, Pizza, and Mains
- + Features a distinct grid of photos and vibrant colorful accents in the branding
- + Better overall balance of the entire menu page
- − Poor text quality with significant gibberish and typos
- − Some food images in the grid look repetitive or messy
- − Layout is a bit cluttered compared to the minimalist request
Verdict: GPT Image 1 produces a more professional and usable result due to its superior text rendering and high-fidelity food photography, despite not showing the full page. Wan 2.6 captures the 'vibrant accents' and complete section requirements better, but fails significantly on legibility and clean minimalist aesthetics.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
GPT Image 1
- + Excellent typography rendering with crisp, fiery outlines
- + Very clean and professional composition suitable for a high-end ad
- + Accurately represents all burger layers in a clear, exploded view
- − Missed the '6' in the price tag, displaying '.99' instead
- − The background is relatively static compared to the burger's motion
Wan 2.6
- + High emotional energy with smoke and realistic fire effects
- + Accurately includes the full price '€6.99' in the starburst
- + Dynamic tilted composition enhances the 'exploded' and 'motion' feel
- − Main title text is slightly muddy and less legible than Model A
- − The 'starburst' for the price looks like a 2D sticker slapped onto a 3D scene
Verdict: GPT Image 1 produces a much cleaner and more professional graphic design with superior text rendering and lighting, though it fails the specific price accuracy. Wan 2.6 captures the 'fiery' and 'motion' aspects more dynamically with better background integration, but the main title is harder to read and the starburst is poorly integrated. GPT Image 1 is the better overall advertisement due to its lighting coherence and clarity, despite the numerical typo.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
GPT Image 1
- + Excellent text legibility and spelling accuracy.
- + Very clean, centered composition.
- + Renders a realistic chalk grit texture on the letters.
- − The handwriting feels a bit too uniform, bordering on a digital font look despite the texture.
- − Does not use 'elegant cursive' for the title as requested.
- − The framing is a tight crop that feels less like a 'cozy café' environment.
Wan 2.6
- + Successfully uses a beautiful cursive style for the title text.
- + Much more realistic 'cozy café' atmosphere with depth of field and environmental details like chalk dust.
- + Very natural variation in handwriting slant and pressure.
- − Slightly less crisp text rendering compared to the other model.
- − Minor artifacts in the chalk dust at the bottom of the frame.
Verdict: Wan 2.6 followed the stylistic instructions much better, providing the requested elegant cursive for the title and a more authentic café atmosphere. While GPT Image 1 had excellent legibility, its handwriting looked somewhat mechanical and it failed the cursive requirement.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
GPT Image 1
- + Excellent cinematic lighting and dark, moody atmosphere
- + Strong anatomical rendering of the horse and textures on the spacesuit
- + Higher level of fine detail in the horse's coat and muscle definition
- − The pose is somewhat static for a 'cinematic' scene
- − Failed the negative constraint; the prompt requested 'horse on top' to be surreal, but it showed the standard rider.
Wan 2.6
- + Vibrant color palette with striking nebulae and light rays
- + Dynamic sense of motion with the horse's mane and dust effects
- + Cleaner reflection on the astronaut's visor
- − Failed the logical constraint; specifically ignored the 'horse on top, not vice versa' instruction
- − Slightly less realistic rendering of the horse's legs compared to Model A
Verdict: Both GPT Image 1 and Wan 2.6 failed the difficult spatial reasoning constraint of placing the 'horse on top' of the astronaut, both opting for the conventional astronaut-on-horse image. GPT Image 1 is preferred for its superior texture work and more grounded cinematic aesthetic, whereas Wan 2.6 feels more like a digital composite with a generic space background.
Outfit Transfer Challenge
Editing“Use Image 1 as the base person. Dress them in the exact elaborate outfit from Image 2 (including all layers, accessories, jewelry, and shoes). Carefully adapt the clothing to the body shape and pose in Image 1 while maintaining realistic fabric behavior, correct proportions, and perfect lighting/shadow matching. Keep the person’s exact face, hair, and background completely unchanged.”
AI Judge Analysis
GPT Image 1
- + Excellent replication of the specific plaid pattern and coat style from Image 2
- + Successfully added accessories like the watch that were present in the reference
- + High visual quality with clean lighting and fabric textures
- − Significantly altered the subject's face and changed the vitiligo pattern from Image 1
- − Modified the background landscape, removing various details on the sand
Wan 2.6
- + Perfectly preserved the original background from Image 1
- + Highly accurate preservation of the subject's face, hair, and vitiligo patterns
- + Correctly included the sunglasses from the reference image whereas Model A forgot them
- − The plaid pattern on the scarf is slightly simplified compared to the source
- − The jacket is cropped at the waist, failing to demonstrate the fit of the jeans and full outfit
Verdict: Wan 2.6 is the clear winner for image editing as it followed the preservation instructions almost perfectly, keeping the subject's unique features and the background intact. While GPT Image 1 produced a high-quality stylized image, it essentially generated a new photo of a different person, failing to maintain the identity from the source image.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
GPT Image 1
- + Excellent composition with a focused, realistic close-up on the capybara's expression.
- + Higher photographic quality with realistic textures on the fur and the hat materials.
- + Very accurate adherence to the 'both paws on steering wheel' and professional attire instructions.
- − The passenger is heavily blurred, making it harder to discern her phone and expression.
- − The interior is very dark, which obscures some of the requested taxi details.
Wan 2.6
- + Great environmental storytelling with rainy streets and vibrant Manhattan neon lights visible through the glass.
- + Both characters are in clear focus, making the passenger's 'bored' expression visible.
- + Includes more external taxi details like the checkerboard trim and roof light.
- − The steering wheel placement is physically impossible, appearing to emerge from the center console/passenger side.
- − The capybara's paws look somewhat distorted and do not grasp the wheel realistically.
- − The composition feels a bit crowded and flatter compared to the cinematic depth of Model A.
Verdict: GPT Image 1 is the superior image due to its high photographic realism and logical composition. While Wan 2.6 captures more of the New York atmosphere, it suffers from a major structural error where the steering wheel and the driver's seat are misaligned, whereas GPT Image 1 feels like a coherent, high-quality film still.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
GPT Image 1
- + Excellent typographic consistency and gothic style.
- + Cleaner, more atmospheric lighting that feels cinematic.
- + Perfectly captures the 'dark parchment' texture and moody aesthetic.
- − Failed to include '7pm' and mislabeled the location under the 'TIME' header.
- − The jack-o-lantern and trees are very subtle/dark compared to the prompt's focus.
Wan 2.6
- + Highly accurate adherence to all text details including date, time, and location.
- + Richly detailed border containing both the requested thorns and spiderwebs.
- + Stronger visual hierarchy with a more vibrant jack-o-lantern and twisted trees.
- − The font for the secondary scroll banner is a bit generic compared to the main title.
- − The 'parchment' effect is limited to the edges rather than the whole background.
Verdict: While GPT Image 1 has superior atmosphere and sophisticated typography, it fails significant text components of the prompt by omitting the time and mislabeling the location. Wan 2.6 accurately captures every requested detail, provides a more complex border, and ensures all event information is present and correct, making it the better functional invitation.
Bald man challenge
Image Editing“Give the person a full, thick head of natural hair with realistic texture, density, and a natural hairline. Preserve facial features and lighting.”
AI Judge Analysis
GPT Image 1
- + Excellent preservation of the original facial anatomy and skin texture
- + Correctly applies a thick, dark hair texture that matches the beard style
- + Accurate preservation of the background and original lighting environment
- − The hairline is a bit unnatural and sits too high on the forehead
- − The hair volume is slightly exaggerated and looks somewhat like a wig
Wan 2.6
- + Natural and realistic hair flow and texture
- + Flawless integration of the hair with the ears and temple area
- + Highly convincing hairline and hair density
- − Slightly alters the shape of the person's face/forehead compared to the source image
- − Small artifact where the hair meets the top frame of the glasses
Verdict: Wan 2.6 provides a much more natural and realistic integration of the hair, making it look like part of the subject's original appearance. While GPT Image 1 preserves the facial anatomy more strictly, the hair itself looks stuck-on and lacks the realistic flow and natural hairline found in the Wan 2.6 edit.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
GPT Image 1
- + Excellent typography with clean, centered layout and correct stacking.
- + High-quality 3D clay-like textures match the 'cartoon scene' and 'soft' descriptors perfectly.
- + Superior composition that fills the frame more effectively while maintaining the requested isometric view.
- − The salmon texture is slightly simplified, appearing more like plastic than realistic PBR textures.
Wan 2.6
- + Stronger adherence to 'realistic PBR' for the sushi materials, especially on the shrimp and rice.
- + Good isometric perspective and clean background.
- − Typography layout is flawed, placing the flag icon in the middle of a line rather than below the text.
- − Shadow clipping occurs at the bottom of the diorama base.
- − The text occupies too much of the upper frame compared to the relatively small subject.
Verdict: GPT Image 1 followed the complex multi-part prompt more accurately, particularly regarding the specific layout of the text and flag. While Wan 2.6 had slightly more realistic food textures, its overall composition felt empty and the typography was awkwardly placed. GPT Image 1's soft cartoon aesthetic and perfect alignment make it the superior result for this specific design request.
Over-the-top cartoon caricature
Editing“Create a caricature of me and my job. Make it exaggerated and humorous, incorporating my profession as a tv show anchor and my love for dogs and hockey.”
AI Judge Analysis
GPT Image 1
- + Excellent caricature style with exaggerated facial features that still maintain the subject's likeness
- + Creative integration of all elements including a hockey-playing dog on the TV screen
- + The denim shirt from the original photo is accurately preserved
- − The watercolor texture is a bit messy around the hands
Wan 2.6
- + Clean digital vector art style
- + Includes multiple dogs and professional studio equipment like headsets and lights
- + Clear rendering of the hockey stick and puck
- − The facial features are generic and do not resemble the woman in the source image
- − The 'caricature' aspect is more of a standard cartoon than an exaggerated caricature
Verdict: GPT Image 1 followed the instructions far better by providing a true caricature that captures the specific likeness and personality of the source image while exaggerating it. While Wan 2.6 created a nice cartoon scene, it failed to preserve the subject's likeness, creating a generic character instead.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
GPT Image 1
- + Excellent character composition with all four animals clearly visible and interacting
- + Consistent, soft lighting that backlights the fur beautifully
- + Better anatomical accuracy on the animals' paws and expressions
- − The kitten's tail is missing/obscured
- − Butterflies are slightly less detailed compared to Model B
Wan 2.6
- + Beautiful atmosphere with strong god rays and sparkling dew effects
- + Each animal has a distinct personality and clear tail visible
- + High level of detail in the butterflies and foreground flora
- − The fox's front-right paw is anatomically messy
- − The kitten's front-left paw looks slightly distorted
Verdict: Both models captured the joyful, wholesome prompt perfectly. GPT Image 1 feels like a more cohesive group portrait with superior lighting balance, while Wan 2.6 excels at the atmospheric 'dew sparkles' and environmental effects despite some minor anatomical issues with the animals' paws.
Studio Ghibli Anime Style
Editing“Transform this photo into a Studio Ghibli–inspired illustration. Use soft pastel colors, hand-painted textures, gentle lighting, dreamy backgrounds, and a warm, nostalgic mood”
AI Judge Analysis
GPT Image 1
- + Successfully captures a textured, hand-painted watercolor aesthetic reminiscent of classic anime backgrounds.
- + Accurately recreates the character's facial expressions in a simplified artistic style.
- + Maintains the warm, nostalgic, and soft pastel color palette requested.
- − Simplifies the plaid pattern of the shirt too much compared to the source.
Wan 2.6
- + Stays very close to the source image's composition and specific details like the plaid pattern on the shirt.
- + Features clean linework and clear character recognition.
- + Incorporates a pleasant watercolor wash texture.
- − The white speckle overlay feels like an artificial filter rather than part of the illustration style.
- − The character style leans more towards modern manga/manhwa than the specific 'Studio Ghibli' aesthetic requested.
Verdict: GPT Image 1 is the winner as it better captures the specific 'Studio Ghibli' art style through its soft, painterly textures and dreamy, glowing lighting. While Wan 2.6 preserves more detail from the original photo, such as the exact plaid pattern and sharper facial features, its style is less representative of the requested nostalgic anime aesthetic and relies on a distracting particle overlay.
Golden Hour Stroll
Image Editing“Add dynamic motion to this photo: make hair blow in the wind, add leaves flying, energetic and lively feel.”
AI Judge Analysis
GPT Image 1
- + Strong adherence to all parts of the motion prompt
- + Successfully added movement to the dog's fur/tail and the woman's hair
- + Includes a large quantity of flying leaves providing a sense of wind
- − Introduced some texture inconsistencies on the woman's face and original clothing details
- − Leaves are somewhat blurry and lack high-definition detail
Wan 2.6
- + Excellent preservation of the source image's identity, face, and clothing details
- + Clean, high-quality rendering of the flying leaves
- + Subtle but natural-looking hair movement that respects the original lighting
- − The wind effect is less energetic compared to the prompt's request
- − The dog remains static, failing to include it in the 'lively' feel of the scene
Verdict: GPT Image 1 followed the instructions more comprehensively by applying motion to the hair, dog, and environment, though it slightly altered the character's facial features. Wan 2.6 achieved a cleaner, more high-fidelity result that perfectly preserved the source image, but it was much more conservative with the requested 'energetic' motion. GPT Image 1 is the winner for better capturing the intended mood of the edit.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
GPT Image 1
- + Excellent typography with perfect spelling and correct accent marks.
- + Strong minimalist vector style that feels modern yet vintage.
- + Balanced composition with the banner at the base.
- − Ignored the 'light background' request, providing a black background instead.
- − The steam element is a bit thick and less elegant than Model B.
Wan 2.6
- + Successfully used a light background with subtle vintage texture as requested.
- + The 'Est. 1720' banner is integrated into the icon in a creative way.
- + The steam graphic is delicate and well-rendered.
- − The text 'Caffè' uses a backtick instead of a proper accent grave.
- − The composition feels slightly off-balance with the banner only on the right side.
Verdict: GPT Image 1 captures the minimalist logo aesthetic perfectly with superior typography and spelling, though it failed the background color instruction. Wan 2.6 followed the color and texture prompts better but struggled with the specific accent mark in 'Caffè' and produced a less balanced layout. GPT Image 1 is the preferred logo due to its professional, clean vector execution and correct text.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
GPT Image 1
- + Excellent adherence to the six requested steps and iconography.
- + High-quality vector aesthetic with a consistent NASA-inspired color palette.
- + Legible text and logical layout that functions as an actual infographic.
- − Includes a typo in the footer ('EARLLUNAR' rather than individual steps).
- − The icon-to-label mapping is slightly misaligned in the middle row.
Wan 2.6
- + Features a clean, minimalist design.
- + Accurate spelling for the astronaut names.
- − Completely failed to include the requested infographic steps (Launch, Orbit, etc.).
- − Looks more like a towel or fabric texture than a vector poster.
- − Lacks the specific iconography and flat-vector style requested.
Verdict: GPT Image 1 successfully followed the complex multi-step prompt, delivering a detailed infographic with specific NASA-themed icons and a professional flat-vector aesthetic. Wan 2.6 failed to include the primary content of the prompt, providing only the astronaut names on a textured background. GPT Image 1 is the clear winner despite minor text alignment issues.
Explore each model
Alibaba's multimodal generation model from the Wan AI suite, supporting text-to-video, image-to-video, reference-to-video with audio, and text-to-image, in both Chinese and English