Head to head
Esc

Models · slot A

to navigate to pick

Grok Imagine Image xAI Wan 2.5 (Preview) Alibaba

Settled by community votes across 19 shared challenges, with an AI judge weighing in on each.

Grok Imagine Image

23.4 arena score

#26 of 62 in Text-to-Image

Skill signature · Text-to-Image

Wan 2.5 (Preview)

23.4 arena score

#27 of 62 in Text-to-Image

Vote tally

Where the votes landed

Grok Imagine Image

0%

win rate

Ties

0%

Wan 2.5 (Preview)

0%

win rate

Shared challenges 19

Challenge by challenge

The strongest take from each model on every shared challenge, with the AI judge's read.

Geometric Composition

Text-to-Image

“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”

Grok Imagine Image
Wan 2.5 (Preview)

AI Judge Analysis

Grok Imagine Image

  • + Accurate rendering of glass refractive properties and reflections
  • + Clean and photorealistic lighting that matches the 'left window' requirement
  • + Excellent wood texture on the table surface
  • The blue sphere appears to be floating unnaturally in the center of the cube
  • The cube itself is more rectangular than a true cube

Wan 2.5 (Preview)

  • + Excellent texture on the red book cover and pages
  • + Pleasant atmospheric lighting with dust particles
  • + Good placement of the sphere resting on the bottom surface
  • Perspective issues where the book seems to merge into the top edge of the glass
  • Reflections on the base of the cube are slightly inconsistent with the sphere's position

Verdict: Both models followed the prompt instructions precisely, including the specific lighting and object placement. Grok Imagine Image is slightly better due to its cleaner glass rendering and more realistic table surface, whereas Wan 2.5 (Preview) has a slight clipping issue where the book meets the top of the glass cube.

Man and Car in California

Editing
Edit instruction

“Make a photo of the man driving the car down the California coastline”

Source
Grok Imagine Image
Wan 2.5 (Preview)

AI Judge Analysis

Grok Imagine Image

  • + Excellent photographic quality and motion blur
  • + Successfully captures the California coastline aesthetic
  • + Highly realistic car design and lighting
  • Completely failed to use the specific man provided in the source image
  • Replaced the subject with an older Caucasian man in sunglasses

Wan 2.5 (Preview)

  • + Successfully preserved the identity/appearance of the man from the source image
  • + Accurately places the specific car and man into the requested environment
  • + Good preservation of the source car's specific details
  • Visual quality is slightly lower than Model A
  • Palm trees in the background look somewhat repetitive/artificial

Verdict: This was an image editing task involving two source images. Grok Imagine Image created a high-quality photo but failed the core editing requirement by replacing the specific subject with a generic passenger. Wan 2.5 successfully merged the two source images, placing the correct man in the correct car while maintaining his recognizable features and hairstyle.

Candid Street Photography

Text-to-Image

“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”

Grok Imagine Image
Wan 2.5 (Preview)

AI Judge Analysis

Grok Imagine Image

  • + Excellent depiction of motion blur from passing cars.
  • + Highly realistic, candid street photography aesthetic.
  • + Perfect adherence to the 'imperfect framing' and '50mm' feel.
  • The man's face is obscured and directed downward, making detail assessment difficult.
  • Rain is less visible than in Model B.

Wan 2.5 (Preview)

  • + Clear, high-quality facial details and natural skin texture.
  • + Strong atmospheric rain effects and reflections.
  • + Solid overall composition and clear subject focus.
  • Missing the requested motion blur from passing cars.
  • The bicycle's rear structure is physically impossible through the wheel.
  • Feels more like a staged high-fidelity stock photo than a 'candid street photo'.

Verdict: Grok Imagine Image followed the technical camera prompts much better, successfully incorporating the requested motion blur and 'imperfect' candid framing that gives it a realistic street photography feel. While Wan 2.5 (Preview) has impressive textures and visible rain, it missed the motion blur requirement and contains significant structural errors in the bicycle's anatomy, making Grok the winner.

Fantasy Warrior

Text-to-Image

“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”

Grok Imagine Image
Wan 2.5 (Preview)

AI Judge Analysis

Grok Imagine Image

  • + Exquisite engraving detail on the plate armor
  • + Beautiful warm lighting and sparks creating atmosphere
  • + High-quality rendering of stray hairs and skin texture
  • The beads in the hair are less prominent than requested
  • Armor looks a bit too pristine and clean for a 'battle-worn' description

Wan 2.5 (Preview)

  • + Excellent adherence to hair beads requirement with silver accents
  • + Great depiction of 'battle-worn' through dirt and frayed cloth textures
  • + Strong implementation of leather straps and underlayers
  • Visible AI artifacting/blur on the shoulder plate near the neck
  • The face has a slightly flatter, less lifelike appearance compared to Model A

Verdict: Grok Imagine Image creates a more visually stunning and realistic portrait with superior lighting and armor engraving. However, Wan 2.5 (Preview) follows the prompt's specific details more closely, particularly the hair beads and the 'battle-worn' state of the equipment and leather straps.

Modern Clean Menu

Text-to-Image

“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”

Grok Imagine Image
Wan 2.5 (Preview)

AI Judge Analysis

Grok Imagine Image

  • + Excellent structure showing a wide variety of menu items
  • + Clean, professional typography with large headers
  • + Integrated food photography that feels like a real restaurant menu
  • Has duplicate table entry text like 'Grilled Salmon' and 'Steak Frites'
  • Smaller food images are less detailed and somewhat cluttered
  • Text details are gibberish

Wan 2.5 (Preview)

  • + Perfect grid layout for food photography as requested
  • + High-quality, vibrant food images with consistent lighting
  • + Better adherence to the 'modern minimalist' aesthetic
  • Lacks the volume of menu items found in Image A
  • Headings and sections are a bit sparse for a full menu
  • Text spelling is incorrect for several headers

Verdict: Wan 2.5 (Preview) better followed the aesthetic and structural requirements of the prompt by providing a clean grid of high-quality food photos on a minimalist white background. Grok Imagine followed the section instructions well, but the repeated text entries and more cluttered layout make it feel less polished for a professional design.

Magic Burger Explosion: Fiery Photorealism Challenge

Text-to-Image

“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”

Grok Imagine Image
Wan 2.5 (Preview)

AI Judge Analysis

Grok Imagine Image

  • + Excellent typography with a glowing fire effect
  • + Very dynamic composition with splashing sauces and embers
  • + Strong photorealistic textures on the lettuce and burger patty
  • The starburst for the price is a bit flat compared to the rest of the image

Wan 2.5 (Preview)

  • + Creative melting cheese effect integrated into the main title typography
  • + Very clean center-focused composition
  • + Excellent fiery effect on the price starburst
  • The tomato slices look slightly more CGI than photorealistic
  • The 'Limited Time Only' text is slightly less readable than Model A

Verdict: Both models followed the prompt exceptionally well, producing high-quality ad materials. Grok Imagine Image is the winner due to its superior photorealistic textures and more balanced layout, whereas Wan 2.5 (Preview) had a very creative title design but slightly less realistic food components.

Chalkboard Menu

Text-to-Image

“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”

Grok Imagine Image
Wan 2.5 (Preview)

AI Judge Analysis

Grok Imagine Image

  • + Exceptional text rendering with no spelling errors.
  • + Realistic chalk texture with dusty smudges and subtle grain.
  • + Perfect adherence to the requested items and pricing.
  • The 'elegant cursive' for the title is more of a script than formal cursive.

Wan 2.5 (Preview)

  • + Strong chalk-like aesthetic with natural-looking underlines.
  • + Excellent layout that feels authentic to a café atmosphere.
  • Missed the word 'Herbs' in the second menu item.
  • Repeating prices and awkward line breaks in the cookie section.
  • The handwriting looks slightly more digital/clean than Model A.

Verdict: Grok Imagine followed the prompt instructions perfectly, rendering all text accurately with a very realistic chalk texture and no spelling mistakes. Wan 2.5 (Preview) struggled with the text content, omitting words and repeating prices, though it captured the café environment well. Grok Imagine is the clear winner for its superior text fidelity and adherence to the specific menu items.

The Reversed Rodeo

Text-to-Image

“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”

Grok Imagine Image
Wan 2.5 (Preview)

AI Judge Analysis

Grok Imagine Image

  • + Successfully followed the specific instruction for the horse to be on top of the astronaut
  • + Rich, vibrant colors in the nebula background with a high sense of surrealism
  • + Good anatomical detail on the horse
  • The connection between the horse and astronaut is a bit floaty rather than a 'riding' pose
  • Some minor artifacting where the front hoof meets the astronaut's hand

Wan 2.5 (Preview)

  • + Excellent visual clarity and sharp details on the astronaut suit and horse hair
  • + Great composition with the earth and galaxy in the background
  • + Cinematic lighting and sense of motion
  • Failed the negative constraint/specific instruction by placing the astronaut on top of the horse
  • The ground texture at the bottom feels slightly inconsistent with the space setting

Verdict: Grok Imagine was the only model to successfully follow the logical challenge of placing the horse on top of the astronaut as requested. While Wan 2.5 (Preview) produced a higher quality, more cinematic image, it completely ignored the specific spatial instruction and generated the standard 'cliché' version of an astronaut riding a horse.

Outfit Transfer Challenge

Editing
Edit instruction

“Use Image 1 as the base person. Dress them in the exact elaborate outfit from Image 2 (including all layers, accessories, jewelry, and shoes). Carefully adapt the clothing to the body shape and pose in Image 1 while maintaining realistic fabric behavior, correct proportions, and perfect lighting/shadow matching. Keep the person’s exact face, hair, and background completely unchanged.”

Source
Grok Imagine Image
Wan 2.5 (Preview)

AI Judge Analysis

Grok Imagine Image

  • + Successfully preserved the exact person, face, and unique skin features from Image 1.
  • + High resolution and intricate detail in the clothing applied.
  • + Perfectly preserved the background and wooden structure.
  • Completely failed to use the specified outfit from Image 2, instead generating a generic elaborate costume.
  • Ignored the specific modern coat, scarf, and sunglasses in Image 2.

Wan 2.5 (Preview)

  • + Successfully used the exact outfit, scarf, sunglasses, and watch from Image 2.
  • + Applied the clothing to the pose of the base image very effectively.
  • + Maintained the background lighting and environment from Image 1.
  • Completely failed to preserve the person from Image 1, replacing the face and skin with the man from Image 2.
  • Failed to keep the person's exact face and hair as requested in the instructions.

Verdict: This comparison represents a complete trade-off in prompt adherence: Grok Imagine Image successfully preserved the subject's identity but failed to use the correct clothing, while Wan 2.5 (Preview) successfully extracted the clothing but failed to preserve the subject's identity. Grok Imagine Image is preferred as an edit because it respects the 'base person' instruction, whereas Wan 2.5 essentially performed a background swap of the second person, which is the opposite of the requested task.

The Capybara Taxi Driver

Text-to-Image

“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”

Grok Imagine Image
Wan 2.5 (Preview)

AI Judge Analysis

Grok Imagine Image

  • + The lighting and textures on the capybara are highly realistic.
  • + The woman is sitting in the passenger seat as specified in the prompt.
  • + The interior details, like the dashboard and rear-view mirror, are very convincing.
  • The woman appears to be sitting in the front passenger seat rather than the back seat.
  • The taxi light on top of the car is visible through the roof, which is a structural impossibility.

Wan 2.5 (Preview)

  • + Successfully places the passenger in the back seat as requested.
  • + The capybara's hat is more detailed and resembles a formal driver cap.
  • + The rainy atmosphere and city lights through the window add great aesthetic value.
  • The text on the taxi sign is nonsensical.
  • The steering wheel appears to be melting into the capybara's paws.

Verdict: Wan 2.5 (Preview) followed the positional instructions better by placing the businesswoman in the back seat, whereas Grok Imagine put her in the front. While Grok Imagine has slightly cleaner textures and a more realistic car interior, the structural error of the taxi light appearing to sit on the inside ceiling makes Wan 2.5 (Preview) the more successful interpretation of the scene.

The Halloween Invitation

Text-to-Image

“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”

Grok Imagine Image
Wan 2.5 (Preview)

AI Judge Analysis

Grok Imagine Image

  • + Excellent gothic typography that feels integrated with the theme
  • + Strict adherence to all prompt elements including the thorn and web border
  • + Strong atmospheric composition with cinematic lighting around the jack-o-lantern
  • The parchment edges look slightly clipped against the black background

Wan 2.5 (Preview)

  • + Highly detailed jack-o-lantern with a dynamic fire effect
  • + Vibrant colors and high-quality textures on the twisted tree and scroll
  • + Perfect spelling and legible event details
  • The 'Halloween Party Invitation' text lacks the requested 'gothic' style, looking more like standard 3D serif font
  • The blue sky background feels a bit too comic-like compared to the requested 'dark parchment' vintage look

Verdict: Both models followed the complex text instructions perfectly, which is impressive. Grok Imagine is the likely winner because its aesthetic style aligns much better with the 'vintage gothic' and 'dark parchment' requirements, whereas Wan 2.5 (Preview) produced a more modern, saturated digital illustration. Grok's font choices and border execution feel more cohesive as a party invitation.

Bald man challenge

Image Editing
Edit instruction

“Give the person a full, thick head of natural hair with realistic texture, density, and a natural hairline. Preserve facial features and lighting.”

Before After
Grok Imagine Image
Before After
Wan 2.5 (Preview)

AI Judge Analysis

Grok Imagine Image

  • + Excellent preservation of original facial features and bone structure
  • + Hair texture and color match the existing beard perfectly
  • + Highly realistic hairline and integration with the ears
  • The forehead height remains slightly high, giving a receding look despite the prompt for full hair

Wan 2.5 (Preview)

  • + Very thick and voluminous hair that clearly satisfies the 'full' requirement
  • + Matches the facial lighting well
  • Slightly altered the person's eye shape, making him look like a younger/different person
  • The hair volume is a bit stylized and looks less 'natural' compared to the reference beard texture
  • Minor distortion on the ear-stem of the glasses

Verdict: Grok Imagine Image provides a much more convincing and realistic edit by perfectly matching the new hair to the original's beard texture and preserving every facial detail. While Wan 2.5 successfully adds a thicker head of hair, it slightly changes the person's identity and facial features in the process.

Isometric Miniature Diorama Scenes

Text-to-Image

“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”

Grok Imagine Image
Wan 2.5 (Preview)

AI Judge Analysis

Grok Imagine Image

  • + Excellent adherence to the 45° isometric perspective including the base.
  • + Very clean typography and flag placement as requested.
  • + Realistic PBR-style textures and lighting on the wood and ingredients.
  • The white nigiri has a slightly floating appearance relative to the rice.

Wan 2.5 (Preview)

  • + Beautiful soft-render 3D aesthetic and high-quality textures.
  • + Correct text spelling and inclusion of the flag element.
  • + Nice focus and depth of field effect.
  • Failed the isometric 45° perspective requirement, opting for a standard portrait angle.
  • The flag icon is positioned to the side rather than the requested top-center location.

Verdict: Grok Imagine followed every technical instruction in the prompt, including the specific isometric viewing angle and precise placement of text and UI elements. While Wan 2.5 (Preview) produced a charming, high-quality render, it ignored the core isometric camera instruction and layout specifics.

Over-the-top cartoon caricature

Editing
Edit instruction

“Create a caricature of me and my job. Make it exaggerated and humorous, incorporating my profession as a tv show anchor and my love for dogs and hockey.”

Source
Grok Imagine Image
Wan 2.5 (Preview)

AI Judge Analysis

Grok Imagine Image

  • + Retains a higher level of facial likeness from the source image compared to the other model.
  • + Successfully integrates all elements: TV anchor desk, hockey rink background, and a dog in hockey gear.
  • + Clean, high-quality digital art style that fits the 'humorous caricature' request perfectly.
  • The hands are slightly malformed and undersized for the character's body.
  • The pucks in the background look a bit generic and flat.

Wan 2.5 (Preview)

  • + Incorporates the blue denim shirt from the source image as a nod to the original outfit.
  • + Clearly includes hockey items (stick, jersey) and multiple dogs in the composition.
  • + Strong caricature 'big head' style with exaggerated features.
  • The facial likeness is significantly lower, making the subject look like a generic cartoon character.
  • The background TV shows what appears to be a hockey player on a grass field (likely field hockey/football hybrid).
  • Visible artifacts on the hockey stick and the microphone.

Verdict: Grok Imagine is the superior model for this task because it maintains a clear facial resemblance to the woman in the source image while translating her into a caricature. It also handles the 'TV anchor' prompt more cohesively by placing her in a detailed studio setting, whereas Wan 2.5 (Preview) includes confusing elements like an athlete on a grass field on the monitor behind the subject.

Adorable Baby Animals in Sunny Meadow

Text-to-Image

“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”

Grok Imagine Image
Wan 2.5 (Preview)

AI Judge Analysis

Grok Imagine Image

  • + Excellent depiction of golden hour lighting and sun rays
  • + Cute, stylized character designs that fit a wholesome vibe
  • Artificial, 'AI-illustrated' look with painted fur textures
  • Failed to include butterflies mentioned in the prompt
  • Poorly defined backgrounds and anatomical merging of the animals

Wan 2.5 (Preview)

  • + Much higher realism in fur textures and animal anatomy
  • + Successfully included all elements, including butterflies and dew drops
  • + Superior composition with a more natural depth of field
  • The fox kit has slightly uncanny, human-like eyes
  • Some floating artifacts or seeds in the air look a bit messy

Verdict: Wan 2.5 (Preview) provided a much more realistic interpretation that fully adhered to the prompt elements, most notably the butterflies and intricate fur textures. Grok Imagine produced a charming but overly stylized, illustrative image that felt more like a digital painting than the 'hyper-photorealistic' masterpiece requested.

Studio Ghibli Anime Style

Editing
Edit instruction

“Transform this photo into a Studio Ghibli–inspired illustration. Use soft pastel colors, hand-painted textures, gentle lighting, dreamy backgrounds, and a warm, nostalgic mood”

Source
Grok Imagine Image
Wan 2.5 (Preview)

AI Judge Analysis

Grok Imagine Image

  • + Excellent hand-painted, watercolor-like texture that feels authentic to Ghibli background art.
  • + Accurately replicates the character positions and expressions while translating them into an anime style.
  • + The background buildings and sky are beautifully reimagined in a dreamy, nostalgic style.
  • The man's facial features (beard and eyes) feel a bit more like generic illustration than the specific Ghibli facial archetypes.

Wan 2.5 (Preview)

  • + Strong character designs that very closely mimic the Ghibli 'hero/heroine' face shapes and eyes.
  • + Great use of 'dreamy' elements like floating leaves and glowing dust motes.
  • + Preserves the composition of the original photo perfectly.
  • The textures are a bit too smooth and digital, lacking the requested hand-painted or watercolor feel.
  • The colors are somewhat washed out compared to the vibrant but soft palette expected of Ghibli.

Verdict: Both models did an excellent job of translating the 'Distracted Boyfriend' meme into an anime aesthetic. Grok Imagine Image (Model A) wins on background and texture, providing a rich, hand-painted watercolor look that is a hallmark of Studio Ghibli, whereas Wan 2.5 (Model B) has slightly more accurate Ghibli-style character line work but a more sterile, digital finish.

Golden Hour Stroll

Image Editing
Edit instruction

“Add dynamic motion to this photo: make hair blow in the wind, add leaves flying, energetic and lively feel.”

Before After
Grok Imagine Image
Before After
Wan 2.5 (Preview)

AI Judge Analysis

Grok Imagine Image

  • + Excellent adherence to the 'energetic and lively' instruction with many flying leaves.
  • + Effective wind effect on the hair that looks natural and broad.
  • + Great preservation of the original subjects and background elements.
  • The orange autumn leaves clash slightly with the green trees in the background.
  • Some leaves overlapping the subject appear a bit sharp and flat.

Wan 2.5 (Preview)

  • + Natural-looking wind effect on the woman's hair.
  • + High image quality and preservation of the original scene details.
  • + Subtle and clean editing of the leaves.
  • The green leaves look like stickers and lack natural depth or motion blur.
  • Fewer leaves than Model A makes the scene feel less 'energetic'.

Verdict: Grok Imagine Image followed the 'energetic and lively' instructions more effectively by adding a significant amount of motion throughout the frame with flying leaves and more pronounced wind-blown hair. Wan 2.5 (Preview) produced a higher quality hair effect but the added leaves appeared sparse and lacked the same sense of dynamic movement. Grok Imagine Image is the winner for better capturing the specific mood requested while maintaining the source image's integrity.

Vintage Cafe Logo

Text-to-Image

“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”

Grok Imagine Image
Wan 2.5 (Preview)

AI Judge Analysis

Grok Imagine Image

  • + Excellent typography rendering for both the main name and established date.
  • + Clean vector emblem style with modern minimalist appeal.
  • + High contrast and sharp details on a textured light background.
  • Redundant text as it includes 'Est. 1720' twice.
  • Integration of a spoon/handle into the cloche dome is slightly messy and unconventional.

Wan 2.5 (Preview)

  • + Cohesive use of framing and vintage paper textures for a 'vintage' feel.
  • + Elegant banner design that flows well with the cloche graphic.
  • + Classic typography that matches the historical aesthetic perfectly.
  • The 'Est.' text is slightly less crisp compared to the main title.
  • The steam vector is a bit thinner and less prominent than in Image A.

Verdict: Both models performed exceptionally well on this task, but Wan 2.5 (Preview) produced a more cohesive vintage design by effectively using the banner and parchment textures requested. Grok Imagine Image had cleaner vector lines and sharper text, but it included redundant information and a slightly confusing graphic element merged with the cloche.

Apollo 11: Journey to Tranquility

Text-to-Image

“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”

Grok Imagine Image
Wan 2.5 (Preview)

AI Judge Analysis

Grok Imagine Image

  • + Successfully included all 6 numbered steps requested in the prompt.
  • + Followed the color palette requirements more closely with prominent navy and muted red.
  • + Text rendering is remarkably clean and legible for an AI generation.
  • Step 3 text has gibberish/artifact words ('3rajcoory', 'Transluiory').
  • The Saturn V icon is a bit simplified/stylized compared to a realistic vector.

Wan 2.5 (Preview)

  • + High-quality vector illustrations for the Saturn V and Lunar Module.
  • + Excellent use of space and logical flow for the trajectory lines.
  • + Clean, professional aesthetic suitable for a real poster.
  • Failed to include 6 distinct numbered steps as requested in the instructions.
  • Steps 5 and 6 (Descent and Landing) are just text labels without dedicated icons.
  • Text alignment for the steps is somewhat disorganized compared to Model A.

Verdict: Grok Imagine followed the complex prompt instructions much more accurately, including all six specific steps with their corresponding icons and numbers. While Wan 2.5 (Preview) produced more sophisticated and detailed vector artwork, it failed the instructional adherence by omitting several icons and merging the final steps without the requested structure.

Next steps

Explore each model