An image generation model by xAI designed to generate highly aesthetic images from text descriptions.
Settled by community votes across 19 shared challenges, with an AI judge weighing in on each.
Grok Imagine Image
#26 of 62 in Text-to-Image
Wan 2.6
#28 of 62 in Text-to-Image
Where the votes landed
Grok Imagine Image
51.9%
win rate
Ties
3.7%
Wan 2.6
44.4%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
Grok Imagine Image
- + Excellent photorealistic lighting and textures on the wood and book.
- + Cleverly shows the plant pot visible through the glass as requested.
- + Very clean, modern aesthetic with sharp details.
- − The blue sphere is floating physically impossible in the center of the air.
- − The glass cube looks more like a hollow frame rather than a solid-walled glass object.
Wan 2.6
- + Natural physics with the sphere resting on the bottom surface.
- + Accurate glass refraction and reflections on the internal surfaces.
- + Excellent adherence to all spatial requirements of the prompt.
- − The blue sphere is quite large, bordering on not being 'small' relative to the cube.
- − The book cover has some slightly messy edge artifacts.
Verdict: Both models followed the prompt instructions perfectly, including the specific lighting and placement of the plant. Grok Imagine Image has a slightly cleaner, high-end photographic look, but Wan 2.6 is arguably better due to the realistic placement of the sphere on the floor of the cube rather than levitating in mid-air.
Man and Car in California
Editing“Make a photo of the man driving the car down the California coastline”
AI Judge Analysis
Grok Imagine Image
- + Excellent preservation of the car's model and specific design details
- + High visual quality with realistic motion blur and lighting
- + Successfully replaces the entire environment with a convincing California coastline
- − Completely failed to include the man from the second source image, using a generic placeholder instead
Wan 2.6
- + Excellent character preservation, accurately incorporating the specific man and his clothing from the source image
- + Strong composition that places the subject and car naturally within the requested environment
- + Maintains the luxury convertible aesthetic of the original car
- − The car model changed slightly from the original (interior dashboard and grilles are different)
Verdict: Wan 2.6 is the clear winner because it successfully integrated both source images by placing the specific man from the second image into the car from the first. Grok Imagine ignored the second source image entirely, rendering a generic driver that did not resemble the requested person.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
Grok Imagine Image
- + Excellent adherence to the 'imperfect framing' and 'candid' aspects of the prompt
- + Superior representation of motion blur from passing cars
- + Natural and realistic lighting and colors that feel like real film photography
- − The subject's face is obscured, making it hard to verify 'Japanese man' or skin texture
- − The bicycle structure is slightly nonsensical near the bottom bracket
Wan 2.6
- + Exceptional skin texture and facial detail
- + Very clear rendering of rain droplets on surfaces
- + Great composition showing the subject actively engaged in the repair task
- − Lacks the requested 'motion blur' from passing cars as the background car appears static
- − The raindrops look slightly like physical gems or beads rather than natural rain due to over-sharpening
- − Framing feels too deliberate and 'perfect', missing the candid/imperfect request
Verdict: Grok Imagine captured the 'candid street photo' aesthetic much better, providing realistic motion blur and a convincing 'snapshot' feel that matches the 50mm lens and imperfect framing requests. While Wan 2.6 produced a stunningly detailed portrait with excellent skin textures, it failed to include the requested motion blur and felt more like a posed cinematic still than a candid moment.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
Grok Imagine Image
- + Exquisite detail on the ornate floral engravings of the armor
- + Very sharp and clean facial features with realistic textures
- + Superior lighting contrast and bokeh effect
- − The character looks too clean and polished for the 'battle-worn' description
- − Scars look more like surface scratches than deep battle wounds
Wan 2.6
- + Excellent interpretation of 'battle-worn' with grit, caked dirt, and raw emotion
- + Highly realistic beads and hair braiding that feels integrated into the style
- + Incredible texture on the frayed cloth and weathered leather straps
- − The armor engravings are slightly less intricate than Image A
- − Slightly more chaotic composition than the centered symmetry of Image A
Verdict: Both models performed exceptionally well on this complex prompt. Grok Imagine produced a beautiful, clean, and highly symmetrical portrait with stunning armor detail, but Wan 2.6 captured the 'battle-worn' essence much more effectively with realistic grime, weathered textures, and a more emotive expression.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
Grok Imagine Image
- + Strictly followed the category requirements for Appetizers, Pizza, and Mains.
- + Excellent usage of white space and professional sans-serif typography.
- + Sophisticated integration of food photography with the layout, using varying crops and positions.
- − Considerable amount of text repetition and gibberish in the item descriptions.
- − The 'Pizza' section contains several items that are clearly not pizzas (e.g., Steak Frites).
Wan 2.6
- + Successfully implemented the 'grid' layout requested in the prompt.
- + Higher quality, more realistic food photography with consistent lighting.
- + Includes pricing details which adds to the realism of a restaurant menu.
- − Failed to include a dedicated 'Appetizers' list, only having a heading above pizza photos.
- − Significant text rendering issues including garbled letters and nonsensical prices (e.g., $0.09).
- − Overall layout feels more like a flyer or advertisement than a functional menu.
Verdict: Grok Imagine Image followed the structural requirements of the prompt much better, providing distinct sections for all three requested categories and maintaining a clean, professional aesthetic. While Wan 2.6 has superior photographic quality and followed the 'grid' instruction, the actual content of the menu is disorganized and fails to provide the requested variety of food sections. Grok Imagine Image is the winner for its better adherence to the complex layout and categorization instructions.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
Grok Imagine Image
- + Excellent text rendering with clean, legible fonts
- + High contrast and vibrant colors that make the product pop
- + Accurately includes all requested text elements, including the starburst
- − The 'exploded' effect feels slightly more static than Model B
- − The lettuce and tomatoes look a bit more illustrative/digital than photorealistic
Wan 2.6
- + Superior sense of motion and dynamic explosion of ingredients
- + Excellent atmospheric smoke and fire integration
- + The burger patty and sauce have a more photorealistic texture
- − The primary 'MAGIC BURGER' title is slightly distorted by the fire effects
- − The '6.99' price tag looks like a flat clip-art sticker rather than an integrated graphic
Verdict: Grok Imagine Image provides a much cleaner commercial advertisement with perfectly rendered text and a professional layout, whereas Wan 2.6 captures more realism in the food textures and a better sense of explosive movement. Grok is preferred for its superior adherence to the graphic design elements of the prompt, including the specific placement and legibility of the text.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
Grok Imagine Image
- + Perfectly rendered text with no spelling errors or repetitions.
- + Excellent chalk texture and smear effects for a realistic chalkboard look.
- + Very clean layout and composition that remains legible and professional.
- − The handwriting is slightly more 'neat' than 'elegant cursive' for the title.
- − The background cafe scene is very blurry and relatively generic.
Wan 2.6
- + Authentic-looking chalk dust accumulation at the bottom of the frame.
- + The heading captures the requested cursive style more effectively.
- + Excellent handwriting variation that feels genuinely human-drawn.
- − Redundant text repetition with '- $24' and '- $28' appearing twice for those items.
- − Inconsistent line spacing makes the board feel somewhat cluttered.
- − The 'cookies' item has a slight spelling/character merging issue in 'Chocolate'.
Verdict: Grok Imagine followed the complex text instructions perfectly, providing a clean, error-free menu with great chalk texture. Wan 2.1 had a more artistic and realistic cursive style with authentic dust details, but it failed on text accuracy by repeating prices and having slightly inconsistent spacing. Grok Imagine is the winner for its superior text rendering and adherence to the specific menu contents.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
Grok Imagine Image
- + Successfully follows the logical constraint of the horse 'riding' (on top of) the astronaut
- + Excellent color vibrancy and nebula details
- + Strong surreal atmosphere that aligns with the prompt's tone
- − The posing of the horse's front legs is slightly awkward against the astronaut's hands
- − The tether or rope between them feels chemically unresolved in the background
Wan 2.6
- + High technical quality of the astronaut suit and horse texture
- + Beautiful lighting and cinematic god rays
- − Completely failed the secondary prompt instruction (horse on top, not vice-versa)
- − Classical 'astronaut riding horse' trope which ignores the specific surreal challenge
Verdict: The comparison hinge on prompt adherence: Grok Imagine successfully interpreted the 'surreal' instruction to have the horse on top of the astronaut, whereas Wan 2.6 defaulted to the common trope of an astronaut riding a horse. Despite Wan 2.6 having slightly sharper textures, Grok Imagine is the clear winner for following specific, non-standard spatial instructions.
Outfit Transfer Challenge
Editing“Use Image 1 as the base person. Dress them in the exact elaborate outfit from Image 2 (including all layers, accessories, jewelry, and shoes). Carefully adapt the clothing to the body shape and pose in Image 1 while maintaining realistic fabric behavior, correct proportions, and perfect lighting/shadow matching. Keep the person’s exact face, hair, and background completely unchanged.”
AI Judge Analysis
Grok Imagine Image
- + Excellent preservation of the person's identity and face.
- + High quality rendering of the new textures and embroidery.
- − Completely failed to use the clothing from Image 2.
- − Hallucinated an entirely different 'elaborate' outfit.
Wan 2.6
- + Accurately used the clothes, scarf, and sunglasses from Image 2.
- + Maintained the subject's face, hair, and sand details quite well.
- + Good lighting integration of the new clothing.
- − Did not include the jeans or shoes from the reference as requested.
- − The fit of the coat around the shoulders is slightly bulky.
Verdict: Wan 2.6 followed the instructions much more accurately by transferring the specific clothing (the pea coat, plaid scarf, and sunglasses) from Image 2 onto the person in Image 1. In contrast, Grok Imagine provided a high-quality but completely irrelevant royal outfit that ignored the visual content of the reference image provided for the clothes.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
Grok Imagine Image
- + Excellent character expressions that perfectly match the 'bored' and 'professional' descriptions
- + Strong lighting consistency between the interior cabin and the New York street lights
- + High level of detail on the capybara's fur and the passenger's coat
- − The passenger appears to be in the front seat rather than the back seat as requested
- − The steering wheel placement is slightly awkward relative to the capybara's paws
Wan 2.6
- + Correctly places the human passenger in the back seat
- + More authentic NYC taxi driver hat design and vehicle decals
- + Great use of external rain/wetness effects to add realism to the atmosphere
- − The capybara's hands look more like primate hands than capybara paws
- − The interior of the car looks significantly more worn/dirty than a typical commercial scene
- − The resolution and sharpness are slightly lower than Model A
Verdict: Both models captured the whimsical prompt well, but they succeeded in different areas of composition. Grok Imagine Image produced a higher quality, more photorealistic image with better expressions, but failed to place the passenger in the back seat. Wan 2.6 followed the spatial instructions perfectly and captured the gritty NYC atmosphere, though it suffered from anatomical issues with the capybara's paws.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
Grok Imagine Image
- + Excellent typography rendering with zero spelling errors.
- + Clean, polished composition that feels like a professional graphic design product.
- + Atmospheric lighting on the jack-o-lantern and surrounding trees.
- − The thorns and webs border looks slightly more digital/artificial compared to the rest of the image.
Wan 2.6
- + Beautiful, high-contrast cinematic lighting with a warm orange glow against a cool blue sky.
- + Rich, organic textures on the twisted trees and the jack-o-lantern.
- − The word 'Invitation' contains a spelling error in the letter 't'.
- − The text for the date and time is not centered, creating a slightly unbalanced composition.
Verdict: Grok Imagine followed every instruction perfectly, producing a professional-grade invitation with flawless text rendering and a centered, clean layout. While Wan 2.6 offered more vibrant, painterly lighting and a moodier color palette, it suffered from a spelling error and poor text alignment at the bottom. Grok Imagine is the clear winner for its superior text accuracy and overall design polish.
Bald man challenge
Image Editing“Give the person a full, thick head of natural hair with realistic texture, density, and a natural hairline. Preserve facial features and lighting.”
AI Judge Analysis
Grok Imagine Image
- + Excellent source preservation, keeping the face and background identical.
- + Very natural and believable hair texture and density.
- + Matches the hair color and lighting of the existing beard perfectly.
- − The hair is a bit thin in the front compared to the 'thick' request, though more realistic.
Wan 2.6
- + Successfully added a very thick and full head of hair as requested.
- + Good integration of the hair with the sideburns and beard.
- − Subtly altered facial features, making the person look younger and changing the eye area.
- − The hair thickness looks slightly artificial/wig-like along the top edge.
Verdict: Grok Imagine Image provides a much more successful edit by perfectly preserving the original person's identity and facial structure while adding realistic hair. Wan 2.6 followed the 'thick' instruction more aggressively but at the cost of altering the subject's face and original character.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
Grok Imagine Image
- + Perfect adherence to the 45-degree isometric perspective and diorama base request.
- + Excellent text rendering and layout with consistent font and centered flag icon.
- + Higher variety and density of sushi models while maintaining a clean aesthetic.
- − The lighting is a bit harsh with very dark, sharp shadows compared to the 'gentle' request.
- − The plate cuts off slightly at the edge of the blue base.
Wan 2.6
- + Beautiful soft lighting and 'gentle' shadows that match the prompt perfectly.
- + High-quality PBR-like materials, especially on the wooden board and wasabi texture.
- + Very clean and modern graphic design for the text and flag.
- − The camera angle is slightly lower than the requested 45-degree top-down isometric view.
- − Text placement is a bit tight with the flag squeezed next to 'SUSHI' rather than being a separate small element.
Verdict: Both models followed the prompt well, but Grok Imagine Image (Model A) captured the specific '45-degree isometric' look and 'diorama' feel more accurately. While Wan 2.6 (Model B) had superior soft lighting and material textures that felt more premium, Model A's composition and precise adherence to the requested layout make it the winner for this specific challenge.
Over-the-top cartoon caricature
Editing“Create a caricature of me and my job. Make it exaggerated and humorous, incorporating my profession as a tv show anchor and my love for dogs and hockey.”
AI Judge Analysis
Grok Imagine Image
- + Retains an incredible facial resemblance to the woman in the source image.
- + Successfully interprets the 'caricature' style with the classic big-head/small-body aesthetic.
- + Perfectly integrates all requested themes: TV news desk, hockey rink background, and dogs with hockey gear.
- − The hockey stick in the dog's mouth is slightly warped.
- − The transition between the photographic face and the illustrated body is a bit jarring.
Wan 2.6
- + Strong, cohesive illustrative art style across the entire image.
- + Creative inclusion of a secondary dog (pug) in a hockey jersey.
- + Clear depiction of all requested elements: TV anchor equipment, dogs, and hockey gear.
- − Loses the specific facial likeness of the source image, looking like a generic cartoon character.
- − The left hand holding the hockey stick is anatomicaly awkward relative to the arm position.
Verdict: Grok Imagine Image is the superior choice because it successfully maintains the identity of the person in the source image, transforming her face into a recognizable caricature. While Wan 2.6 creates a fun illustration, it fails the 'caricature of me' aspect by replacing the user with a generic cartoon character. Grok's inclusion of specific details like the 'Sports Scoop' papers and the dogs on ice skates makes for a more clever interpretation of the prompt.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
Grok Imagine Image
- + Strong vibrant colors and clear, bright lighting
- + Whimsical, storybook-like composition with great dynamic movement
- − Has a very stylized 3D animation look rather than the requested hyper-photorealism
- − Failed to include the butterflies requested in the prompt
- − Anatomical issues with the rabbit's placement and the kitten's eyes
Wan 2.6
- + Excellent adherence to the hyper-photorealistic style with realistic fur texture
- + Includes all requested elements including the butterflies and dew sparkles
- + Strong composition with clear god rays and natural-looking interactions
- − Some digital artifacts visible in the floating dandelion seeds/particles
- − The kitten has an extra-long tail that slightly breaks the realism
Verdict: Wan 2.6 is the clear winner as it followed all prompt instructions, including the specific animals and the butterflies, while achieving a much higher level of realism. Grok Imagine produced a high-quality image, but it looks more like a Pixar-style render than a photo and missed the butterfly detail entirely.
Studio Ghibli Anime Style
Editing“Transform this photo into a Studio Ghibli–inspired illustration. Use soft pastel colors, hand-painted textures, gentle lighting, dreamy backgrounds, and a warm, nostalgic mood”
AI Judge Analysis
Grok Imagine Image
- + Captures the iconic Studio Ghibli 'painted' background style with fluffy white clouds and blue skies.
- + Maintains high character fidelity, accurately preserving the poses, expressions, and clothing of the original meme.
- + The color palette is vibrant yet warm, consistent with Ghibli's summer-themed films.
- − The line art on the characters is a bit thin compared to the classic bold Ghibli linework.
- − The man's facial features feel slightly more generic than the source's expressive pout.
Wan 2.6
- + Successfully applies a beautiful watercolor 'wash' texture across the entire image.
- + Excellent hand-drawn aesthetic with soft, expressive line art.
- + The lighting is very gentle and dreamy, perfectly matching the 'soft pastel' request.
- − The addition of white bokeh/sparkle dots feels more like generic shoujo anime than specific Ghibli style.
- − The background is very washed out and loses the architectural detail present in the source.
- − The transition between the man's arm and the woman on the right is slightly messy.
Verdict: Grok Imagine Image is the winner because it perfectly balances the requested Studio Ghibli aesthetic with excellent source preservation. It maintains the specific layout, expressions, and colors of the original 'distracted boyfriend' meme while transforming the environment into a rich, hand-painted Ghibli world. Wan 2.6 provides a beautiful watercolor illustration, but it loses too much background detail and adds distracting sparkles that deviate from the Ghibli art style.
Golden Hour Stroll
Image Editing“Add dynamic motion to this photo: make hair blow in the wind, add leaves flying, energetic and lively feel.”
AI Judge Analysis
Grok Imagine Image
- + Excellent source preservation of the woman and dog
- + Highly energetic feel with numerous leaves
- − Leaves look like floating stickers rather than being part of the environment
- − Dog's ears are slightly distorted to simulate movement
Wan 2.6
- + Natural looking wind effect on the hair
- + Leaves are integrated realistically with motion blur
- + Near-perfect preservation of the original image details
- − Fewer leaves make the scene feel slightly less 'energetic' than instructed
Verdict: Grok Imagine followed the prompt by adding a very large amount of leaves, but they look like a flat overlay over the original image. Wan 2.6 provided a much more subtle and realistic edit, with hair that flows naturally and leaves that have appropriate motion blur, making it the superior technical edit.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
Grok Imagine Image
- + Excellent typography with proper accent on 'Caffè'
- + Superior vector-style illustration with clean linework
- + Accurate rendering of the requested text elements
- − Repeats the 'Est. 1720' text twice, which wasn't specifically requested
- − Includes some ambiguous shapes behind the cloche
Wan 2.6
- + Good use of the requested banner element for the date
- + Pleasant vintage border texture on the background
- + Simple and clear composition
- − Text rendering for 'Caffè' uses a generic font and is slightly poorly spaced
- − The cloche illustration is less refined and feels more like clip-art
Verdict: Grok Imagine Image provides a much more professional-grade logo with superior typography and high-quality vector illustrations that feel like a real brand identity. While Wan 2.6 follows the banner instruction well, the overall execution lacks the polish and typographic elegance found in the Grok output.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
Grok Imagine Image
- + Successfully included all six requested infographic steps with corresponding icons.
- + Followed the color palette and flat-vector style perfectly.
- + Text and labels are mostly legible and properly arranged.
- − Contains minor AI spelling artifacts in the secondary text (e.g., '3rajoory', 'Moom').
- − The Saturn V icon is generic rather than realistic.
Wan 2.6
- + Clean, minimalist aesthetic with clear text rendering.
- + Good use of white space and profile silhouettes for the crew.
- − Complete failure to follow the core instruction of creating a 6-step infographic.
- − Missed all specific icons requested (Saturn V, orbit rings, trajectory arc, lunar module).
- − Lacks the complex informational density required by the prompt.
Verdict: Grok Imagine followed the complex prompt instructions near-perfectly, creating a detailed 6-step infographic with the requested icons and NASA-inspired color scheme. In contrast, Wan 2.6 failed to generate an infographic at all, producing a simple poster with only the crew names and none of the requested technical steps or specific icons.
Explore each model
Alibaba's multimodal generation model from the Wan AI suite, supporting text-to-video, image-to-video, reference-to-video with audio, and text-to-image, in both Chinese and English