An image generation model by xAI designed to generate highly aesthetic images from text descriptions.
Settled by community votes across 20 shared challenges, with an AI judge weighing in on each.
Grok Imagine Image
#25 of 62 in Text-to-Image
Vidu Q2
#42 of 62 in Text-to-Image
Where the votes landed
Grok Imagine Image
0.0%
win rate
Ties
0.0%
Vidu Q2
100.0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
Grok Imagine Image
- + Excellent photorealism and depth of field
- + Accurate representation of refraction through the glass cube
- + Subtle and natural lighting that matches the 'soft window light' prompt
- − The cube geometry is slightly elongated, appearing more like a rectangular prism
- − The blue sphere is levitating rather than sitting on the bottom of the cube
Vidu Q2
- + Perfect cube geometry with clear edges
- + Higher detail on the book spine and plant textures
- + Realistic physics with the ball resting on the glass surface
- − The lighting is a bit harsh and contrasted compared to the 'soft' light requested
- − The reflection of the ball on the bottom surface looks more like a mirror than simple glass
Verdict: Both models followed the complex spatial instructions perfectly. Grok Imagine (Model A) produced a more artistic and soft image with superior lighting, while Vidu Q2 (Model B) provided better structural clarity and realistic physics for the object placement. Vidu Q2 is the winner for its superior rendering of the cube's proportions and textures.
Man and Car in California
Editing“Make a photo of the man driving the car down the California coastline”
AI Judge Analysis
Grok Imagine Image
- + Excellent photographic quality and motion blur
- + Perfect adherence to the environment requested
- + High-quality rendering of the Rolls-Royce convertible
- − Failed to use the specific man provided in the source image
- − The driver is a generic older Caucasian man instead of the young Black man in the reference
Vidu Q2
- + Successfully preserved the identity of the specific man from the source image
- + Good preservation of the car's design and features
- + Dynamics of the shot feel appropriate for a coastal drive
- − The driver is positioned on the wrong side for US roads (right-hand drive)
- − Technical quality is slightly lower than Model A, with some harsh edges around the driver
Verdict: Model A failed the multi-image editing task by completely replacing the subject with a different person, despite creating a very high-quality image. Model B correctly identified both the car and the man to combine them into the requested scene. Vidu Q2 is the winner for actually following the instructions to use the specific man provided, whereas Grok Imagine only used the car.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
Grok Imagine Image
- + Excellent depiction of motion blur with a passing car as requested.
- + Highly convincing 'imperfect framing' that mimics a real candid photograph.
- + Strong adherence to the 50mm lens and shallow depth of field instructions.
- − The man's face is obscured, making it hard to judge age/ethnicity details.
- − The presence of a face mask feels slightly dated/specific, though realistic for Japan.
Vidu Q2
- + Exceptional skin texture and detail on the man's hands and face.
- + The bicycle mechanics look more varied and detailed.
- + Good representation of wet pavement and light rain.
- − Lacks the requested motion blur for the passing car.
- − The composition feels more staged/professional than the 'candid street photo' requested.
- − Some anatomical issues with the proximity and angle of the hands relative to the body.
Verdict: Grok Imagine followed the stylistic prompts much more accurately, capturing the 'candid' feel, the 50mm lens aesthetic, and the motion blur of passing cars. Vidu Q2 produced a sharper image with impressive textures, but it ignored the motion blur requirement and felt less like a spontaneous street photo.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
Grok Imagine Image
- + Exquisite engraving detail on the plate armor
- + Beautiful atmospheric lighting and bokeh sparks that frame the subject
- + Strong adherence to the 'beads in hair' prompt with multiple visible beads
- − The armor looks almost too clean and polished despite the facial scars
- − Hair texture looks a bit soft and blurry in some areas
Vidu Q2
- + Superior texture on the cloth underlayer and leather straps
- + The armor reflects a much better 'battle-worn' aesthetic with grime and scratches
- + Very sharp and lifelike skin textures and facial anatomy
- − The 'braided with small beads' requirement is less prominent
- − The torchlight reflection on the metal is a bit more muted than Image A
Verdict: Both models performed exceptionally well, but Vidu Q2 captures the 'battle-worn' and 'highly detailed texture' aspects of the prompt more effectively, particularly on the leather and cloth. While Grok Imagine produced a more traditionally beautiful and ornate image with better lighting, Vidu Q2's grit and realistic material rendering make it the more convincing paladin.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
Grok Imagine Image
- + Strictly follows the requested sections for Appetizers, Pizza, and Mains.
- + Excellent layout with a clear hierarchy and readable bold sans-serif fonts.
- + Integrated high-quality food photography that fits a professional menu style.
- − Contains repetitive text entries like multiple instances of 'Steak Frites' and 'Grilled Salmon'.
- − Minor spelling errors in larger text such as 'Calmon' and 'Saldes'.
Vidu Q2
- + Includes price placeholders which adds to the authenticity of a menu layout.
- + Vibrant and high-resolution food photography with good color saturation.
- − Layout is cluttered and lacks a clear grid structure for images.
- − The text is largely unreadable with heavy garbling and nonsensical words like 'APECIZEN' and 'PPEIZEIS'.
- − Failed to follow the specific sections requested, merging appetziers and pizza haphazardly.
Verdict: Grok Imagine is the clear winner as it adheres much better to the specific layout requirements and section headers requested in the prompt. While Vidu Q2 has vibrant colors, its text rendering and overall organization are poor compared to the professional, structured approach of Grok Imagine.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
Grok Imagine Image
- + Excellent typography with clean, flaming textures
- + Dynamic debris and sauce splashes add a great sense of motion
- + Perfect price tag rendering with the requested currency symbol
- − The starburst graphic for the price is slightly more generic than the rest of the art
Vidu Q2
- + Strong fiery atmosphere and background integration
- + The burger texture looks very appetising and detailed
- − The pricing text contains a symbol error, showing a pound/double-cross hybrid instead of a Euro symbol
- − The font choice for 'Limited Time Only' is less impactful than the headline text
Verdict: Grok Imagine Image followed the prompt more accurately, particularly regarding the text rendering and the Euro currency symbol. While Vidu Q2 produced a high-quality image with great textures, the failure to correctly render the requested currency and the slightly weaker typography makes Grok Imagine Image the superior advertisement.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
Grok Imagine Image
- + Excellent text rendering with perfect spelling and legibility
- + Authentic chalk texture and smudging on the chalkboard
- + Effective completion of the truncated prompt description
- − The 'elegant cursive' for the title is more of a print-script hybrid than true cursive
Vidu Q2
- + Very realistic chalk strokes with layered pressure and varying thickness
- + Dynamic, hand-drawn composition that feels truly organic
- − Numerous spelling errors including 'Truffe Musshoom' and 'Octopd'
- − The text becomes garbled and illegible toward the bottom of the board
- − Incorrect price rendering for the first item
Verdict: Grok Imagine is the clear winner due to its superior linguistic capabilities; it rendered all text perfectly with no spelling errors while maintaining a believable chalk texture. Vidu Q2 captured a more artistically messy 'chalk' aesthetic but failed significantly on legibility and prompt adherence, resulting in nonsensical words and several typos.
Pose & Character Mashup
Editing“Use Image 1 as the exact pose reference and Image 2 as the character reference. Recreate the person/character from Image 2 in the exact dynamic pose and body position from Image 1. Keep the exact face, hair, clothing style/details, and expression from Image 2. Match the lighting and environment of Image 1. The final image must show the character from Image 2 performing the precise action/pose from Image 1 with perfect anatomy and natural integration.”
AI Judge Analysis
Grok Imagine Image
- + Slightly modifies the leg position to be more dynamic and balanced.
- + Maintains the high color saturation of the original environment.
- − Completely failed the character transfer, keeping the person from the pose reference.
- − Applied very minimal changes to the source image, largely ignoring the second input image.
Vidu Q2
- + Excellent character transfer, including sunglasses, scarf, and specific clothing details.
- + Perfectly replicates the complex leg-crossing pose and overall body position from Image 1.
- + Accurately integrates the character into the lighting and background of the target scene.
- − Slight distortion in the anatomy of the hands and fingers.
- − The watch on the wrist is slightly blurry compared to other details.
Verdict: Vidu Q2 successfully completed the complex task of transferring a specific character's identity and clothing onto a difficult pose, maintaining high fidelity to both source images. Grok Imagine Image failed the primary objective by ignoring the character reference and simply outputting a slightly modified version of the pose reference.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
Grok Imagine Image
- + Excellent adherence to the 'horse on top' spatial instruction
- + High cinematic quality with beautiful nebula lighting
- + Strong surreal atmosphere
- − The horse's front hoof is oddly merged with the astronaut's hand
Vidu Q2
- + Vibrant colors and creative galaxy skin on the horse
- + Sharp focus and high detail resolution
- − Completely failed the semantic instruction of 'horse on top'
- − Generic composition that ignores the 'surreal' inversion requested
Verdict: Grok Imagine followed the complex negative constraint to place the horse on top of the astronaut, creating a truly surreal and cinematic image. Vidu Q2 ignored the specific spatial instruction and provided a standard astronaut-riding-horse image, which, while visually appealing, fails as a prompt adherence task.
Outfit Transfer Challenge
Editing“Use Image 1 as the base person. Dress them in the exact elaborate outfit from Image 2 (including all layers, accessories, jewelry, and shoes). Carefully adapt the clothing to the body shape and pose in Image 1 while maintaining realistic fabric behavior, correct proportions, and perfect lighting/shadow matching. Keep the person’s exact face, hair, and background completely unchanged.”
AI Judge Analysis
Grok Imagine Image
- + Excellent preservation of the person's identity and facial features
- + High quality rendering of the new garment
- − Completely ignored the clothing in Image 2, choosing a random regal outfit instead
Vidu Q2
- + Successfully transferred the correct clothing/accessories from Image 2
- + Maintains the background and general aesthetic of Image 1
- − Significant distortion of the model's facial shape and features
- − Inconsistent arm rendering with sand textures appearing as part of the skin or jacket
- − Modified the hairstyle and added a mustache contrary to instructions
Verdict: This is a trade-off between prompt fulfillment and image quality. Grok Imagine Image perfectly preserved the person but failed the primary task of using the specific clothing provided. Vidu Q2 succeeded in transferring the plaid scarf, pea coat, and sunglasses from Image 2, but it failed to preserve the person's face and introduced several anatomical and texture artifacts. Ultimately, Vidu Q2 is the winner for actually performing the requested image-to-image clothing swap, whereas Grok simply generated a new outfit.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
Grok Imagine Image
- + Excellent adherence to the 'bored' expression for the passenger
- + High-quality fur texture and realistic lighting on the capybara
- + Captures the New York night atmosphere well with recognizable street elements
- − The passenger is sitting in the front passenger seat instead of the back seat
- − The capybara's hands look more like paws with prominent claws which might be slightly unsettling
Vidu Q2
- + Correctly places the human passenger in the back seat as requested
- + Features a more detailed driver cap with a badge
- + The capybara's paws look surprisingly natural on the steering wheel
- − The passenger's expression is slightly more vacant than 'bored'
- − Slightly less realistic skin texture on the human compared to Model A
Verdict: While Grok Imagine Image has superior textures and better captured the 'bored' expression of the passenger, it failed the spatial requirement of having the passenger in the back seat. Vidu Q2 followed the prompt's structural instructions more accurately by placing the businesswoman in the rear and maintaining a high level of photorealism overall.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
Grok Imagine Image
- + Perfect text rendering for all requested titles and event details.
- + Highly cohesive composition with a well-balanced border of thorns and webs.
- + Superior atmosphere with cinematic lighting and a moody gothic background.
- − The transition between the dark background and the parchment edges is slightly abrupt.
Vidu Q2
- + Dynamic border elements and a nice illustration style for the Jack-o-lantern.
- + Captures the scroll banner and gothic aesthetic reasonably well.
- − Significant text errors including misspellings in 'Invitation' and the banner.
- − Incorrect date and time details, failing the specific accuracy of the prompt.
- − Lacks the atmospheric depth and background details like the moody sky and twisted trees requested.
Verdict: Grok Imagine significantly outperforms Vidu Q2 by delivering perfect text accuracy and a much more atmospheric, cinematic design that aligns with the 'vintage gothic' prompt. While Vidu Q2 has a nice illustrative style, it suffers from several spelling errors and fails to include the requested date and time correctly.
Bald man challenge
Image Editing“Give the person a full, thick head of natural hair with realistic texture, density, and a natural hairline. Preserve facial features and lighting.”
AI Judge Analysis
Grok Imagine Image
- + Excellent source preservation across the entire image.
- + Realistic hair texture and color that matches the existing beard.
- + Natural, believable hairline that integrates well with the forehead.
- − The hair volume is slightly conservative compared to the request for a 'full, thick' head of hair.
Vidu Q2
- + Successfully added a very thick and voluminous head of hair as requested.
- + Matches the hair color and lighting of the scene quite well.
- − The hairline appears slightly artificial and too straight across the forehead.
- − Minor distortion to the forehead shape compared to the source image.
Verdict: Grok Imagine provided a more seamless and realistic integration, effectively matching the texture and style of the existing beard while perfectly preserving the rest of the image. Vidu Q2 followed the 'thick hair' instruction more literally, but the resulting hairline and forehead structure feel less natural than the Grok version.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
Grok Imagine Image
- + Excellent typography with clean, centered alignment
- + Sophisticated 3D rendering with realistic PBR textures on the fish and wood
- + Perfectly adheres to the isometric 45° perspective
- − The flag icon is a bit large compared to the text
Vidu Q2
- + Creative use of the flag being held by the text
- + Good colors and soft lighting
- − The text is not perfectly centered and looks slightly wonky
- − Sushi components have less defined textures compared to Model A
- − Perspective is slightly more frontal than the requested 45° isometric
Verdict: Grok Imagine Image followed the prompt more precisely, especially regarding the '45° top-down isometric' perspective and the 'ultra-clean' quality. Vidu Q2 had a creative approach to the flag, but Grok Imagine Image's typography and Material (PBR) rendering were significantly more polished and professional.
Over-the-top cartoon caricature
Editing“Create a caricature of me and my job. Make it exaggerated and humorous, incorporating my profession as a tv show anchor and my love for dogs and hockey.”
AI Judge Analysis
Grok Imagine Image
- + Excellent caricature style with a classic large-head-small-body exaggeration.
- + Strong storytelling by placing the character at a news desk with thematic elements on screen.
- + High textual clarity and professional lighting effects.
- − The hands on the desk are very poorly rendered and look deformed.
- − The eyes feel slightly more robotic than the original subject's expression.
Vidu Q2
- + Maintains the subject's original clothing (denim shirt and black top) more accurately.
- + Highly creative integration of dogs, including a small dog in a hockey jersey.
- + Better hand rendering, despite the stylized four-finger cartoon approach.
- − The facial likeness is slightly weaker and more generic than Model A.
- − The microphone is quite massive and overlaps the hair, creating a cluttered composition.
Verdict: Grok Imagine Image creates a more traditional caricature with better environment design, successfully turning the newsroom into a hockey rink. However, Vidu Q2 does a significantly better job at preserving the source image's clothing and provides more charming details like the dog in the hockey jersey. Vidu Q2 is preferred for its harmonious blend of the requested themes and better anatomical coherence in the hands.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
Grok Imagine Image
- + Strong dynamic composition that emphasizes the 'tumbling together' action
- + Excellent lighting effects with clear God rays and golden hour ambiance
- + High level of softness and 'fluffiness' in the textures
- − Failed to include the butterflies mentioned in the prompt
- − Art style leans toward 3D digital illustration rather than 'hyper-photorealistic'
- − Anatomy is slightly stylized with oversized eyes
Vidu Q2
- + Accurately included all requested animals plus the butterflies
- + Achieved a much more 'hyper-photorealistic' look than the competitor
- + Spacious composition captures the 'lush wildflower meadow' perfectly
- − Lighting is a bit flat compared to the dramatic rays in Model A
- − The second golden retriever puppy was not requested (prompt asked for one)
Verdict: Vidu Q2 is the clear winner for its superior adherence to the prompt and more realistic photographic style. While Grok Imagine captured the lighting and 'fluffy' texture well, it missed the butterflies entirely and resulted in a look more akin to a Pixar movie than a photorealistic masterpiece. Vidu Q2 successfully balanced the complex scene with multiple characters and environment details.
Studio Ghibli Anime Style
Editing“Transform this photo into a Studio Ghibli–inspired illustration. Use soft pastel colors, hand-painted textures, gentle lighting, dreamy backgrounds, and a warm, nostalgic mood”
AI Judge Analysis
Grok Imagine Image
- + Successfully captures the hand-painted, watercolor texture characteristic of Studio Ghibli backgrounds.
- + Maintains high character fidelity to the original meme while translating features into the requested art style.
- + Creates a cohesive, nostalgic mood with soft lighting and puffy clouds.
- − The facial expression of the man is slightly more subdued than the original meme.
Vidu Q2
- + Excellent use of pastel colors and dreamy, bright lighting that fits the 'aesthetic' side of anime art.
- + Preserves the iconic facial expressions of the 'distracted boyfriend' meme very effectively.
- + Includes nice line-art details on the clothing and hair.
- − The background feels a bit more like generic digital watercolor rather than the specific, lush Ghibli painting style.
- − The woman in the red dress has slightly inconsistent facial lighting compared to the rest of the scene.
Verdict: Both models did an excellent job of translating the iconic 'Distracted Boyfriend' meme into an anime illustration. Grok Imagine Image is the winner because its background depth, soft watercolor textures, and clouds feel more authentically like a Studio Ghibli film, whereas Vidu Q2 feels slightly more like a standard modern anime style.
Golden Hour Stroll
Image Editing“Add dynamic motion to this photo: make hair blow in the wind, add leaves flying, energetic and lively feel.”
AI Judge Analysis
Grok Imagine Image
- + Excellent wind effect on the hair that feels natural and dynamic.
- + Significant volume of leaves adds a strong sense of chaotic motion.
- + Highly accurate preservation of the subject's face and the dog.
- − The leaves appear somewhat repetitive in shape and color.
- − Almost all leaves were changed to orange, losing the original green summer feel of the park.
Vidu Q2
- + Natural integration of leaves with varied motion blur and sizes.
- + Beautifully rendered hair movement that maintains flow and realism.
- + Excellent preservation of the original image identity and colors.
- − Fewer leaves across the scene compared to Model A, resulting in slightly less 'energetic' feel.
- − Minor distortion on a few leaves in the foreground.
Verdict: Both models successfully interpreted the request with high fidelity to the source image. Vidu Q2 is the winner because the addition of leaves feels more realistic with varying focal depths and motion blur, and it better preserves the original lighting and color temperature of the scene, whereas Grok Imagine tends to turn everything into an autumn-specific palette.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
Grok Imagine Image
- + Perfect text rendering of the restaurant name and establishment date
- + Clean vector illustration with professional logo composition
- + Effective use of the requested brown and cream color palette with a subtle grain texture
- − The addition of a spoon and cup handle merging with the cloche makes the silhouette slightly cluttered
Vidu Q2
- + The banner design is classic and aesthetically pleasing
- + Follows the light background and warm tone instructions well
- − Significant spelling errors in the brand name and 'Est' text
- − The steam effect inside the cloche appears somewhat muddy and lacks clarity
- − Messy typography with inconsistent letter sizing and odd characters
Verdict: Grok Imagine produced a high-quality, professional logo with perfect typography and a cohesive vector style. In contrast, Vidu Q2 failed significantly on text rendering, producing misspelled words and distorted characters, making it unusable as a logo.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
Grok Imagine Image
- + Excellent text legibility for the main headings and crew names.
- + Closely adheres to the specified NASA-inspired color palette and flat-vector style.
- + Contains all six requested steps in a logical, numbered sequence.
- − The 'Translunar' section has visual artifacts and garbled sub-text.
- − The Earth icon mistakenly includes a ring, which is traditionally associated with Saturn.
- − The NASA logo is a slightly deformed approximation rather than a clean vector.
Vidu Q2
- + Features a clean, high-resolution aesthetic with consistent iconography.
- + Good use of the light gray and navy colors from the requested palette.
- + Stronger interpretation of the translunar trajectory and lunar surface 3D perspective.
- − Text is largely nonsensical or misspelled (e.g., 'ALFONCH', 'Aeutth Orot').
- − Fails to follow the requested numbering for the six specific steps.
- − The icons for descent and landing are repetitive rather than distinct steps.
Verdict: Grok Imagine is the superior choice for this challenge because it successfully followed the sequential instructions and produced mostly legible text that reflects the Apollo 11 mission. While Vidu Q2 has a very clean aesthetic, its failure to generate readable english text and its breakdown of the numbered steps makes it ineffective as an infographic.
Explore each model
ShengShu Technology's text-to-image and reference-to-image model with support for character consistency and multi-reference image processing