ShengShu Technology's text-to-image and reference-to-image model with support for character consistency and multi-reference image processing
Settled by community votes across 19 shared challenges, with an AI judge weighing in on each.
Vidu Q2
#42 of 62 in Text-to-Image
Wan 2.5 (Preview)
#26 of 62 in Text-to-Image
Where the votes landed
Vidu Q2
0.0%
win rate
Ties
0.0%
Wan 2.5 (Preview)
100.0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
Vidu Q2
- + Perfect adherence to spatial prompt instructions, including the blue sphere inside and the plant behind.
- + Excellent rendering of glass physics, reflections, and wood grain texture.
- + Highly realistic lighting and shadows that align with the window source described.
- − The plant foliage is somewhat repetitive and looks slightly synthetic compared to the other objects.
Wan 2.5 (Preview)
- + High visual appeal with a cinematic depth of field and soft lighting.
- + Good texture on the red book, showing realistic wear on the edges.
- − Failed to place the sphere 'inside' the cube properly; it appears to be intersecting the glass or hovering in a broken space.
- − Distracting artifacts like floating dust motes that weren't requested.
- − The plant's position through the glass is distorted in a way that breaks the cube's geometry.
Verdict: Vidu Q2 followed every aspect of the prompt with high precision, particularly the complex layering of objects and their reflections within the glass. Wan 2.5 (Preview) produced a visually soft image but failed on the spatial logic of placing the sphere 'inside' the cube, with the sphere appearing to float through the glass walls. Vidu Q2 is the clear winner for its superior physics and prompt adherence.
Man and Car in California
Editing“Make a photo of the man driving the car down the California coastline”
AI Judge Analysis
Vidu Q2
- + Excellent preservation of the car's specific make and model details.
- + Succesfully places the man from the second source image into the driver's seat.
- + Dynamic motion blur on the road enhances the 'driving' theme.
- − The man's hair appears slightly altered and simplified compared to the source.
- − The perspective of the road behind the car is slightly disjointed.
Wan 2.5 (Preview)
- + Very realistic lighting and integration of the car into the coastal environment.
- + Captures the iconic California coastline aesthetic with palm trees and cliffs.
- + Good motion blur on the wheels and road.
- − The man is very small and lacks detail, making him difficult to identify as the subject from the source.
- − Significant artifacts are visible on the man's face and hair.
Verdict: Vidu Q2 is the winner because it successfully integrates both source elements—the specific man and the car—while maintaining high detail on the subjects. Wan 2.5 (Preview) creates a more beautiful landscape, but the person is barely recognisable and poorly rendered, failing the primary goal of the image editing task.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
Vidu Q2
- + Excellent skin texture and weathered detail on hands and face
- + Effective 'imperfect framing' that feels like a genuine candid snap
- + Great wet pavement reflections and realistic rain effects
- − The bicycle anatomy is physically impossible with multiple frames and mangled chains
- − Motion blur on the car is subtle/lacking compared to the prompt
- − Distracting artifacts near the man's forehead
Wan 2.5 (Preview)
- + Superb composition with deep reflections and visible raindrops
- + High clarity and realistic depth of field
- + Accurate depiction of a red bicycle and repair tools
- − Skin texture is slightly too smooth and looks a bit digital
- − Lacks the requested 'motion blur from passing cars'
- − The framing is very centered and professional rather than 'imperfect'
Verdict: Vidu Q2 captures the 'gritty' and 'candid' feel of the prompt much better with its focus on aged skin and weathered textures, though the bicycle itself is a mess of AI artifacts. Wan 2.5 produces a much cleaner, more aesthetically pleasing image with better overall logic, but it misses the 'imperfect' and 'candid' nuances of the request.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
Vidu Q2
- + Excellent crispness in the ornate engravings on the plate armor.
- + Very detailed texture on the leather straps and the thick woven cloth underlayer.
- + Natural looking face with subtle scarring and consistent lighting.
- − The 'braided hair with beads' is a bit understated compared to the request.
- − Lighting feels slightly more studio-like rather than originating from a distinct torch.
Wan 2.5 (Preview)
- + Superior interpretation of the 'hair braided with small beads' and 'dirt on skin' prompts.
- + The warm torchlight and bokeh sparks are more prominent and atmosphere-enhancing.
- + Excellent realism in the eyes, showing redness and fatigue consistent with a 'battle-worn' character.
- − The plate armor engravings are slightly less sharp than Model A.
- − The tattered cloth edges look a bit more artificial compared to the woven texture in Model A.
Verdict: Both models performed exceptionally well, but Wan 2.5 (Preview) captured the specific 'battle-worn' atmosphere and prompt details like hair beads and dirty skin more effectively. While Vidu Q2 produced sharper armor engravings and leather textures, Wan 2.5 felt more lifelike and character-driven with its tired eyes and superior torchlight integration.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
Vidu Q2
- + Includes pricing column which adds to the menu realism
- + Vibrant and high-quality food photography
- + Good use of accent colors to separate sections
- − Text is largely nonsensical and features poor character rendering
- − Layout feels a bit cluttered and lacks the 'minimalist' feel requested
Wan 2.5 (Preview)
- + Closer adherence to the 'minimalist' aesthetic with clean spacing
- + Clearer text rendering for main headings
- + Excellent grid layout that separates the three requested sections
- − Food photos are repetitive and look very similar across different dishes
- − Included 'Modern minimalist' text at the top as if it were a title rather than a design style
Verdict: Wan 2.5 (Preview) better captured the minimalist grid layout requested in the prompt, resulting in a cleaner and more professional-looking menu. While Vidu Q2 had higher quality individual food images, its layout was chaotic and the typography was significantly more distorted.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
Vidu Q2
- + Excellent typography style with a convincing fiery glow effect.
- + Bright, high-contrast colors that pop against the background.
- + Follows the general composition requirements for a food advertisement.
- − The currency symbol is incorrect, showing a crossed 'E' rather than a Euro sign.
- − The burger is not very 'exploded' as requested, with most components still touching.
Wan 2.5 (Preview)
- + Perfect execution of the 'exploded' burger request with dynamic separation of all ingredients.
- + Accurate text rendering including the correct Euro symbol (€).
- + Superior photorealistic detail on individual ingredients like the lettuce and tomato slices.
- − The 'MAGIC BURGER' text looks slightly drippy/melted rather than strictly fiery.
- − Layout is a bit more scattered, though this fits the 'exploded' prompt better.
Verdict: Wan 2.5 (Preview) is the clear winner as it successfully interpreted the 'exploded' burger instruction, whereas Vidu Q2 simply stacked the ingredients. Wan 2.5 (Preview) also correctly rendered the Euro currency symbol and provided more realistic textures for the food items.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
Vidu Q2
- + Excellent realistic chalk texture with dusty smudges
- + Dynamic and natural-looking handwriting style
- − Numerous spelling errors including 'Lemepun' and 'Browd Botter'
- − Failed price accuracy on multiple items
- − Text becomes nonsensical at the bottom
Wan 2.5 (Preview)
- + Perfect adherence to text spelling and prices
- + Very clean legibility while maintaining a chalkboard aesthetic
- + Handled the complex brown butter chocolate chip cookie item well
- − The handwriting looks slightly too uniform, bordering on a digital font look
- − Missing the requested 'elegant cursive' for the title
Verdict: Wan 2.5 (Preview) significantly outperformed Vidu Q2 by following the spelling and content instructions perfectly, whereas Vidu Q2 struggled with severe typos and hallucinated text. While Vidu Q2 had a more authentic chalk texture, Wan 2.5 (Preview) is much more useful due to its high text accuracy and clarity.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
Vidu Q2
- + Excellent visual coherence with a galactic texture on the horse.
- + Very high cinematic quality and vibrant colors.
- − Failed the negative constraint; the astronaut is riding the horse, not vice versa.
- − The horse's left front leg has an unnatural attachment to the body.
Wan 2.5 (Preview)
- + Clean details on the astronaut's suit and the horse's anatomy.
- + Good depth of field and composition with the planets in the background.
- − Failed the negative constraint; the astronaut is riding the horse.
- − The horse's tail and hind legs show motion blur that feels slightly disconnected from the foreground.
Verdict: Both Vidu Q2 and Wan 2.5 failed the specific 'horse on top' logic constraint, likely due to the strong pre-trained association of astronauts riding horses. Vidu Q2 is slightly more visually appealing for this specific prompt as it uses a surreal galaxy texture for the horse, whereas Wan 2.5 is a more grounded, realistic depiction of a horse in space.
Outfit Transfer Challenge
Editing“Use Image 1 as the base person. Dress them in the exact elaborate outfit from Image 2 (including all layers, accessories, jewelry, and shoes). Carefully adapt the clothing to the body shape and pose in Image 1 while maintaining realistic fabric behavior, correct proportions, and perfect lighting/shadow matching. Keep the person’s exact face, hair, and background completely unchanged.”
AI Judge Analysis
Vidu Q2
- + Successfully preserved the unique vitiligo patterns on the person's face and arm.
- + Accurately included the specific accessories like the watch and sunglasses from Image 2.
- + Maintained the body pose and beach environment of the source image.
- − The right sleeve is missing, exposing the arm which contradicts 'long-sleeve' layers from the outfit.
- − Added a mustache and changed the facial structure slightly.
- − Introduced sand artifacts on the coat that weren't in the original.
Wan 2.5 (Preview)
- + Successfully applied the full outfit including the coat, scarf, and watch.
- + Maintained the general background and lighting of Image 1.
- + High resolution and clear textures on the clothing.
- − Completely failed to preserve the base person's identity, replacing the model from Image 1 with the model from Image 2.
- − Lost the unique vitiligo features that were central to the 'base person' instruction.
- − Changed the person's hair and facial features entirely.
Verdict: Vidu Q2 is the clear winner despite technical flaws like the missing sleeve, as it successfully prioritized preserving the identity and skin features of the person in Image 1. Wan 2.5 (Preview) failed the primary constraint of the task by replacing the base person's face and identity with the model from the outfit reference image.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
Vidu Q2
- + Excellent photorealistic lighting and textures
- + Perfectly captures the passenger's bored expression
- + The capybara's anatomy and fur are rendered with high fidelity
- − The driver's front paw has a slightly strange, hand-like structure
- − The view through the window is slightly less recognizable as Manhattan specifically
Wan 2.5 (Preview)
- + Strong composition with the taxi's exterior and interior visible
- + Excellent city background that feels uniquely like Times Square/NYC
- + Effective use of rain on the windshield for mood
- − The capybara's paws look more like strange human hands with long fingernails
- − The passenger's face is slightly blurry and less detailed than model A
- − The text on the taxi sign is nonsensical
Verdict: Vidu Q2 produces a more convincing photorealistic image with better character expressions and textures. While Wan 2.5 (Preview) has a more atmospheric composition and background, the rendering of the capybara's hands and the passenger's face is lower quality compared to Vidu Q2.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
Vidu Q2
- + Features a classic vintage parchment aesthetic
- + The border of thorns and webs is well-integrated with the artwork
- − Has multiple significant spelling errors in all text fields
- − Failed to include the correct date (30.70.2025 instead of 30.10.2026)
Wan 2.5 (Preview)
- + Excellent text rendering with no spelling mistakes in any of the fields
- + Superior visual depth with cinematic lighting and a clear moody sky inside a portal
- + Accurately followed all specific details including date, time, and location
- − The thorns on the border look a bit repetitive or pattern-like
Verdict: Wan 2.5 (Preview) is the clear winner as it successfully rendered every piece of text perfectly, including the specific date and address. Vidu Q2 struggled significantly with typography, producing numerous garbled words and incorrect date figures.
Bald man challenge
Image Editing“Give the person a full, thick head of natural hair with realistic texture, density, and a natural hairline. Preserve facial features and lighting.”
AI Judge Analysis
Vidu Q2
- + Excellent texture and realistic hair strands
- + Preserves the original facial identity perfectly
- + Maintains consistent lighting across the new hair and the face
- − The hairline on the forehead is slightly too straight and abrupt
Wan 2.5 (Preview)
- + Natural, softer hair volume and styling
- + Excellent integration of temples and sideburns into the existing beard
- + Preserves details and background perfectly
- − The face undergoes a subtle shift in features, making him look like a slightly different person
- − The hair texture is a bit softer/blurry compared to the rest of the image
Verdict: Vidu Q2 successfully adds hair while maintaining the exact facial structure of the subject, though the hairline is a bit sharp. Wan 2.5 (Preview) provides a more natural hairstyle that blends better with the beard, but it slightly alters the subject's face, losing some of the original identity. Vidu Q2 is preferred for its superior preservation of the source subject's features.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
Vidu Q2
- + Excellent variety of sushi types on the plate
- + Accurate rendering of all requested text and the flag icon
- + High level of detail in textures and modeling
- − The diorama base has some strange, ambiguous small shapes on the surface
- − The flag is floating or attached poorly to the text
Wan 2.5 (Preview)
- + Perfectly clean, minimalist aesthetic that fits the 'soft refined textures' prompt
- + Very professional typography and flag placement
- + Superb lighting and soft-focus composition
- − Only depicts a single piece of sushi, which feels slightly empty for a 'scene'
- − The rice texture is slightly repetitive and round like beans
Verdict: Both models followed the prompt exceptionally well, producing clean 3D isometric designs. Vidu Q2 offers more visual interest with a variety of sushi items, while Wan 2.5 (Preview) provides a more professional, polished look with superior lighting and better integrated graphic design elements. Vidu Q2 is slightly preferred for adhering better to the 'scene' aspect by including multiple sushi types.
Over-the-top cartoon caricature
Editing“Create a caricature of me and my job. Make it exaggerated and humorous, incorporating my profession as a tv show anchor and my love for dogs and hockey.”
AI Judge Analysis
Vidu Q2
- + Excellent facial resemblance to the subject in the source image.
- + Creative integration of the hockey rink as the background and paw prints on the desk papers.
- + Preserves the specific denim shirt and black top from the original photo.
- − The fingers on the left hand are anatomically incorrect and poorly rendered.
- − The 'NEWS' tag on the microphone looks slightly amateurish in design compared to the rest of the illustration.
Wan 2.5 (Preview)
- + Stronger caricature style with classic 'big head' exaggeration.
- + Clearer inclusion of hockey elements, including a stick and a TV monitor showing a game.
- + Very clean and balanced composition suitable for a profile picture or avatar.
- − The facial resemblance to the source image is significantly weaker than the other model.
- − The hands holding the microphone are very small and lack detail.
- − The background is more generic and less integrated than the rink design in the other model.
Verdict: Overall, Vidu Q2 is the winner because it maintains a high degree of facial resemblance to the subject while successfully incorporating all elements of the prompt (hockey, dogs, and news anchor). While Wan 2.5 (Preview) captures the classic 'caricature' art style more effectively, its failure to preserve the subject's likeness makes it less successful as a personalized edit.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
Vidu Q2
- + Excellent composition showing motion and multiple animals without crowding.
- + Higher detail in the wildflower meadow and background lighting.
- + More accurate butterfly anatomy and variety.
- − Includes two golden retriever puppies instead of one as requested.
- − The bunny has a slightly strange tail/hind-leg configuration.
Wan 2.5 (Preview)
- + Captures the 'big expressive eyes' and 'ultra-detailed soft fur' perfectly.
- + Follows the count of animals accurately (one of each).
- + Beautiful rendering of dew drops and god rays.
- − The fox's eyes appear slightly unnatural and glowing.
- − The butterfly on the right has disconnected wings.
Verdict: Both models delivered vibrant and heartwarming images, but Wan 2.5 (Preview) is the winner because it adhered more closely to the specific count of animals requested. While Vidu Q2 created a more complex scene, it duplicated the golden retriever, whereas Wan 2.5 captured the specific character and expressions of the baby animals more effectively.
Studio Ghibli Anime Style
Editing“Transform this photo into a Studio Ghibli–inspired illustration. Use soft pastel colors, hand-painted textures, gentle lighting, dreamy backgrounds, and a warm, nostalgic mood”
AI Judge Analysis
Vidu Q2
- + Excellent preservation of the source image's layout and character poses.
- + Effective use of watercolor-style textures that mimic traditional Ghibli background art.
- + Accurately captures the specific facial expressions from the original meme.
- − The character line art is a bit inconsistent in thickness compared to the background.
- − Stronger reliance on digital-looking shading in some areas of the clothing.
Wan 2.5 (Preview)
- + Stronger Ghibli aesthetic with the addition of characteristic floating leaves and soft bloom.
- + The character faces are more successfully stylized into the specific Ghibli 'look'.
- + Beautiful soft lighting and color palette that matches the 'dreamy' prompt.
- − The woman on the right has a slightly more neutral expression, losing some of the 'distracted boyfriend' meme's humor.
- − Some minor loss of structural detail in the background architecture.
Verdict: Both models did an exceptional job at transforming the meme into the requested style while keeping the scene recognizable. Vidu Q2 excels at preserving the exact energy and expressions of the source image, while Wan 2.5 (Preview) provides a more atmospheric and authentic Ghibli-style transformation through its lighting and character design.
Golden Hour Stroll
Image Editing“Add dynamic motion to this photo: make hair blow in the wind, add leaves flying, energetic and lively feel.”
AI Judge Analysis
Vidu Q2
- + Excellent adherence to the 'leaves flying' instruction with realistic autumn colors.
- + Successfully animated the hair to look wind-blown while keeping facial features intact.
- + Excellent source preservation, keeping the person and dog backgrounds virtually unchanged.
- − The amount of leaves is slightly overwhelming, partially obscuring the background landscape.
Wan 2.5 (Preview)
- + Natural and elegant wind-blown hair effect that looks very lively.
- + Good source preservation of the human subject and the dog.
- + Subtle use of leaves preserves the original composition well.
- − The green leaves look like digital stickers/overlays rather than integrated objects.
- − The leaves lack the lighting and texture detail found in Model A.
Verdict: Both models performed well on the core task of adding motion via hair and leaves. Vidu Q2 is the winner because the leaves it added have realistic textures, lighting, and seasonal color that matches the environment, whereas the leaves in Wan 2.5 (Preview) appear as flat, bright green shapes that don't belong in the scene. Vidu Q2 also did a slightly better job of creating a 'dynamic' feel through the volume of the hair movement.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
Vidu Q2
- + Successfully captures the warm brown and cream color scheme.
- + The illustration of the cloche dome includes nice stylized steam.
- − Significant text errors including 'CAFFE FARMIIN' and 'Esttt'.
- − The inclusion of a secondary, garbled logo line below the main banner is redundant and messy.
- − Vector lines are a bit soft/blurry in places.
Wan 2.5 (Preview)
- + Excellent typography with perfect spelling of 'CAFFÈ FLORIAN'.
- + Clean, high-quality vector style that fits the minimalist emblem request.
- + Superior background texture that mimics vintage paper.
- − The cloche dome is a bit small relative to the banner size.
- − The steam inside the dome is a bit thick, losing some delicate detail.
Verdict: Wan 2.5 (Preview) produced a far superior logo by following the text instructions perfectly and maintaining a clean, professional vector aesthetic. Vidu Q2 failed significantly on the text rendering, creating several spelling errors and adding unnecessary, garbled design elements at the bottom.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
Vidu Q2
- + Uses a clean, flat-vector style for the icons.
- + Follows the color palette requested with muted tones.
- − Text is largely gibberish and full of spelling errors.
- − The icons for the landing phase are repetitive and lack distinct 'descent' vs 'landing' variation.
- − The layout is somewhat cluttered and lacks a clear chronological flow.
Wan 2.5 (Preview)
- + Excellent text rendering with correct spelling for labels and astronaut names.
- + Logical infographic flow that traces the trajectory from Earth to Moon.
- + Strong adherence to the 'NASA-inspired' navy and white color palette.
- − The Saturn V rocket icon is used twice, even for the descent phase where a lunar module was requested.
- − One of the astronaut icons appears to be wearing a helmet that resembles a Stahlhelm rather than a flight suit.
Verdict: Wan 2.5 (Preview) is the clear winner due to its ability to render legible, accurate text and create a cohesive infographic layout. While it struggled slightly with icon consistency for the descent phase, Vidu Q2 failed significantly on all text elements, rendering the infographic non-functional.
Explore each model
Alibaba's text-to-image and image-to-image generation model from the Wan AI suite, offering high-quality visual generation capabilities