Head to head
Esc

Models · slot A

to navigate to pick

Grok Imagine Image Pro xAI Vidu Q2 ShengShu Technology

Settled by community votes across 20 shared challenges, with an AI judge weighing in on each.

Grok Imagine Image Pro

24.5 arena score

#17 of 62 in Text-to-Image

Skill signature · Text-to-Image

Vidu Q2

19.8 arena score

#42 of 62 in Text-to-Image

Vote tally

Where the votes landed

Grok Imagine Image Pro

0%

win rate

Ties

0%

Vidu Q2

0%

win rate

Shared challenges 20

Challenge by challenge

The strongest take from each model on every shared challenge, with the AI judge's read.

Geometric Composition

Text-to-Image

“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”

Grok Imagine Image Pro
Vidu Q2

AI Judge Analysis

Grok Imagine Image Pro

  • + Excellent realism in the texture of the wooden table and the glass pane reflections.
  • + Highly accurate depiction of light refraction through the glass cube.
  • + Legible and context-relevant text on the book spine.
  • The plant is more 'behind' the cube than 'visible through' it, as some leaves overlap the top.

Vidu Q2

  • + Vibrant colors and strong lighting contrast that follows the 'light from left' instruction.
  • + Good composition with interesting shadows falling across the table.
  • Physics issues with the reflection of the blue sphere which looks more like a second sphere underneath.
  • The glass cube edges appear slightly distorted and inconsistent in thickness.
  • The book is floating slightly above the glass surface.

Verdict: Grok Imagine Image Pro produces a much more photorealistic result with superior handling of glass physics and material textures. Vidu Q2 struggles with the internal reflections of the glass cube, making the blue sphere's reflection look physically impossible, and the book appears poorly seated on the cube's surface.

Man and Car in California

Editing
Edit instruction

“Make a photo of the man driving the car down the California coastline”

Source
Grok Imagine Image Pro
Vidu Q2

AI Judge Analysis

Grok Imagine Image Pro

  • + Excellent preservation of the car's model and design features.
  • + High visual quality with a realistic sense of motion on the road.
  • + Beautifully composed California coastline background.
  • Completely failed to use the specific man from the source image, replacing him with a different person.
  • The steering wheel position seems slightly off compared to the driver's grip.

Vidu Q2

  • + Excellent subject preservation, correctly identifying and placing the man from the source image behind the wheel.
  • + Good preservation of the source car's identifying features.
  • + Dynamic composition with a sense of speed.
  • The man's hands on the steering wheel have anatomical errors (extra/deformed fingers).
  • The car's proportions feel slightly elongated compared to the source.

Verdict: Vidu Q2 is the winner because it successfully followed the core instruction to place the specific man from the source images into the car, whereas Grok Imagine Image Pro completely replaced him with a generic figure. While Vidu Q2 has some artifacts in the hands, it is the only model that actually fulfilled the complex multi-source editing requirement.

Candid Street Photography

Text-to-Image

“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”

Grok Imagine Image Pro
Vidu Q2

AI Judge Analysis

Grok Imagine Image Pro

  • + Excellent full-body composition and anatomy
  • + Realistic lighting and reflections that match the rainy environment
  • + Strong adherence to the 'candid' and 'motion blur' instructions
  • The wrench he is holding doesn't align correctly with the bike's nut
  • Background storefronts look slightly generic

Vidu Q2

  • + Successfully captures a very shallow depth of field
  • + Grit and realistic texture on the bicycle frame and man's skin
  • Major anatomical errors with three hands visible
  • Strange bicycle geometry with a second wheel appearing in the foreground
  • Failed to include motion blur for passing cars

Verdict: Grok Imagine Image Pro provided a much more coherent and believable scene that followed all technical prompt requirements, including motion blur and a wide candid framing. In contrast, Vidu Q2 suffered from significant AI artifacts including extra hands and nonsensical bicycle construction, despite having high-quality textures.

Fantasy Warrior

Text-to-Image

“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”

Grok Imagine Image Pro
Vidu Q2

AI Judge Analysis

Grok Imagine Image Pro

  • + Excellent adherence to the hair braiding and beads requirement, with detailed placement throughout the hair.
  • + Superior rendering of textures, particularly in the weathered leather, rusted engravings, and cloth underlayer.
  • + Very convincing skin textures with realistic scars, dirt, and lifelike eye detail.
  • The placement of beads on the top of the head appears slightly artificial or stuck on.
  • The bokeh sparks are a bit uniform in size across the frame.

Vidu Q2

  • + Strong dynamic lighting with warm torchlight glints on the edges of the armor plates.
  • + Composition feels slightly more natural with a three-quarter view rather than a direct frontal shot.
  • + Good depiction of battle damage including blood spatter and dents on the metal.
  • Minimal braiding and beads compared to the prompt's specific request.
  • Texture on the cloth underlayer is less defined and detailed than the competing model.
  • Armor engravings appear slightly more generic and less 'ornate'.

Verdict: Grok Imagine Image Pro is the winner due to its superior adherence to individual prompt details, particularly the complex braiding and beads, and the rich, weathered textures of the armor and skin. While Vidu Q2 has excellent lighting and a more cinematic pose, it lacks the fine detail and specific character traits requested in the prompt.

Modern Clean Menu

Text-to-Image

“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”

Grok Imagine Image Pro
Vidu Q2

AI Judge Analysis

Grok Imagine Image Pro

  • + Excellent grid layout that strictly follows the requested sections for Appetizers, Pizza, and Mains.
  • + High-quality, realistic food photography that looks professional and appetizing.
  • + Mostly legible text including prices and descriptions that make sense in a menu context.
  • The placeholder text description for 'Avocado Toast' incorrectly repeats the description for 'Prosciutto Arugula'.
  • Uses a slightly stylized, handwritten-looking font for descriptions instead of a strictly bold sans-serif throughout.

Vidu Q2

  • + Successfully incorporates vibrant accents and a modern minimalist aesthetic.
  • + Good distribution of colorful food images across the page.
  • Text is largely gibberish or contains severe misspellings (e.g., 'APECIZEN', 'MPAIZERS').
  • Layout is chaotic and does not follow a clear, professional logical flow like a real menu.
  • Food images are repetitive and some contain visual artifacts/distortions.

Verdict: Grok Imagine Image Pro is the clear winner as it produces a functional, professional-looking menu with a clean grid and realistic food photography. Vidu Q2 fails significantly on text rendering, generating incoherent gibberish and a disorganized layout that lacks the professional standard requested by the prompt.

Magic Burger Explosion: Fiery Photorealism Challenge

Text-to-Image

“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”

Grok Imagine Image Pro
Vidu Q2

AI Judge Analysis

Grok Imagine Image Pro

  • + Excellent photorealistic texture on the meat patty and fresh vegetables.
  • + Dynamic 'exploded' composition with unique cheese-pull effects between layers.
  • + Highly accurate text rendering, including the specific Euro currency symbol and price format.
  • The starburst for the price is a bit generic and flat compared to the rest of the image.
  • Background feels slightly more like floating embers than a full 'fiery' environment.

Vidu Q2

  • + Stronger 'fiery' background that fills the frame with intense energy.
  • + Good glow effect applied to all text layers as requested.
  • + Vibrant colors that pop against the dark backdrop.
  • The currency symbol is malformed and does not resemble a Euro symbol.
  • The explosion of ingredients feels less layered and physics-accurate than Image A.
  • The bottom bun texture appears somewhat artificial.

Verdict: Grok Imagine Image Pro is the winner because it provides a much higher level of photorealistic detail and manages the complex text and symbol requirements perfectly. While Vidu Q2 captured a more intense fiery atmosphere, it failed on the specific currency symbol and the ingredient separation was less sophisticated.

Chalkboard Menu

Text-to-Image

“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”

Grok Imagine Image Pro
Vidu Q2

AI Judge Analysis

Grok Imagine Image Pro

  • + Perfect text accuracy and spelling for all menu items.
  • + Authentic chalk texture and realistic smudge marks on the board.
  • + Clean composition that remains legible and follows the prompt's formatting perfectly.
  • The handwriting is a bit uniform, arguably leaning toward a 'chalk font' aesthetic rather than purely messy handwriting.

Vidu Q2

  • + Very expressive and varied handwriting style with thick chalk strokes.
  • + Good use of color variance in the chalk highlights.
  • Significant spelling errors throughout the text (e.g., 'Musshoom', 'Lemepun', 'Browd Botter').
  • Changed the price of the first item from $24 to $34.
  • The text becomes garbled and illegible at the bottom of the board.

Verdict: Grok Imagine Image Pro followed the prompt with near-perfect precision, delivering accurate spelling and a clean, professional chalkboard aesthetic. Vidu Q2 struggled significantly with text rendering, resulting in numerous spelling errors and pricing inaccuracies that make the image unusable for the specific request. Grok Imagine Image Pro is the clear winner for its superior text legibility and adherence.

Pose & Character Mashup

Editing
Edit instruction

“Use Image 1 as the exact pose reference and Image 2 as the character reference. Recreate the person/character from Image 2 in the exact dynamic pose and body position from Image 1. Keep the exact face, hair, clothing style/details, and expression from Image 2. Match the lighting and environment of Image 1. The final image must show the character from Image 2 performing the precise action/pose from Image 1 with perfect anatomy and natural integration.”

Source
Grok Imagine Image Pro
Vidu Q2

AI Judge Analysis

Grok Imagine Image Pro

  • + Successfully maintains the yellow background and red ottoman from the source image.
  • Completely failed the character instruction by simply re-rendering the original woman.
  • Did not incorporate any clothing elements, face, or features from Image 2.

Vidu Q2

  • + Excellent character transfers, including the face, sunglasses, scarf, and black sweatshirt.
  • + Perfectly replicates the complex leg-cross and arm-extension pose from Image 1.
  • + Accurately matches the lighting of the yellow studio environment from Image 1.
  • Hand anatomy on the lower left arm has too many fingers.
  • The scarf physics are slightly unnatural given the lean of the character.

Verdict: Grok Imagine Image Pro failed the task entirely, providing an output that is nearly identical to Image 1 with no attempt to integrate the character from Image 2. Vidu Q2 followed the instructions meticulously, accurately swapping the character's face, hair, and clothing while maintaining the specific difficult pose and environment.

The Reversed Rodeo

Text-to-Image

“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”

Grok Imagine Image Pro
Vidu Q2

AI Judge Analysis

Grok Imagine Image Pro

  • + Successfully followed the specific instruction of having the horse on top of the astronaut
  • + Highly detailed rendering of both the spacesuit and the horse's musculature
  • + Strong cinematic composition with the planet and nebula backgrounds
  • The horse is floating just above the astronaut rather than literally 'riding' him
  • Anatomical weirdness where the horse's back leg appears to be emerging from its front or belly

Vidu Q2

  • + High visual quality with vibrant colors and cosmic patterns integrated into the horse's coat
  • + Excellent clarity and lighting coordination between the subjects and the environment
  • Failed the negative constraint entirely by placing the astronaut on top
  • Lacks the surreal 'subversion' requested in the prompt

Verdict: The challenge contained a specific spatial constraint ('horse on top, not vice versa') which Grok Imagine Image Pro followed, whereas Vidu Q2 ignored it in favor of a traditional horse-riding-astronaut image. While Vidu Q2 has a beautiful aesthetic with the celestial horse, Grok Imagine Image Pro is the superior output because it adhered to the complex prompt logic.

Outfit Transfer Challenge

Editing
Edit instruction

“Use Image 1 as the base person. Dress them in the exact elaborate outfit from Image 2 (including all layers, accessories, jewelry, and shoes). Carefully adapt the clothing to the body shape and pose in Image 1 while maintaining realistic fabric behavior, correct proportions, and perfect lighting/shadow matching. Keep the person’s exact face, hair, and background completely unchanged.”

Source
Grok Imagine Image Pro
Vidu Q2

AI Judge Analysis

Grok Imagine Image Pro

  • + Excellent preservation of the subject's facial features and specific vitiligo patterns on the forehead.
  • + High visual quality and realistic fabric rendering.
  • + Keeps the background architectural elements and lighting consistent with Image 1.
  • Failed completely to use the outfit from Image 2, substituting it with a generic royal/fantasy costume.
  • The hands are significantly lightened and do not match the subject's actual skin tone.

Vidu Q2

  • + Successfully transferred the specific outfit from Image 2, including the coat, scarf, sunglasses, and watch.
  • + Accurately adapted the clothing to the subject's pose and environment.
  • + Maintains the subject's skin condition and overall likeness.
  • The glasses and mustache from Image 2 were blended into the subject's face, slightly altering the original face preservation.
  • The hand in the bottom left has anatomical distortions (too many fingers).
  • Added sand textures to the coat that were not present in the original outfit.

Verdict: Vidu Q2 is the clear winner as it actually followed the instruction to use the specific outfit from Image 2, whereas Grok Imagine Image Pro hallucinated an entirely different royal costume. While Vidu Q2 slightly altered the face by adding the mustache and accessories of the second model, it succeeded in the complex task of re-dressing the subject in the correct plaid scarf and pea coat.

The Capybara Taxi Driver

Text-to-Image

“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”

Grok Imagine Image Pro
Vidu Q2

AI Judge Analysis

Grok Imagine Image Pro

  • + Excellent text rendering on the driver cap mentioning 'NYC TLC Medallion'.
  • + Superior photorealistic lighting and interior detail.
  • + The businesswoman is correctly positioned in the back seat.
  • The passenger is sitting in the middle/right rather than fully behind the driver as a taxi usually operates, though it still works.
  • The capybara's paws look slightly claw-like and less like natural paws.

Vidu Q2

  • + Natural side-profile composition showing the depth of the car.
  • + Good adherence to the requested calm expression on the capybara.
  • + The businesswoman captures the 'bored' expression very effectively.
  • The businesswoman is placed in the seat directly behind the driver but appears physically separated or lacks clear spatial coherence at the bottom of the frame.
  • The capybara paws look like human hands covered in fur.

Verdict: Grok Imagine Image Pro is the winner due to its superior lighting and texture, creating a highly believable photorealistic scene with impressive text rendering on the taxi cap. While Vidu Q2 offers a good cinematic angle, it suffers from several anatomical and spatial artifacts compared to the polished output of Grok Imagine Image Pro.

The Halloween Invitation

Text-to-Image

“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”

Grok Imagine Image Pro
Vidu Q2

AI Judge Analysis

Grok Imagine Image Pro

  • + Excellent typography with perfect spelling in all requested text fields.
  • + Superb atmospheric lighting and atmospheric depth in the illustration.
  • + High degree of adherence to the composition requirements, including the scroll banner and parchment texture.
  • The scroll banner placement is a bit central rather than integrated into the top or bottom flow.
  • Small spiders on the web border are slightly repetitive.

Vidu Q2

  • + Strong gothic aesthetic for the thorn and branch border.
  • + Good color contrast between the parchment and the jack-o-lantern.
  • Significant spelling errors in the title and subtitle text.
  • Incorrect date (30.70.2025) and time formatting.
  • Illustration style is less cinematic and feels more like a flat graphic compared to the prompt's request.

Verdict: Grok Imagine Image Pro significantly outperformed Vidu Q2 by following all text instructions with 100% accuracy and providing a high-quality, cinematic illustration. Vidu Q2 struggled with basic spelling, failed to render the requested date correctly, and delivered a less polished composition.

Bald man challenge

Image Editing
Edit instruction

“Give the person a full, thick head of natural hair with realistic texture, density, and a natural hairline. Preserve facial features and lighting.”

Before After
Grok Imagine Image Pro
Before After
Vidu Q2

AI Judge Analysis

Grok Imagine Image Pro

  • + Excellent preservation of the source image's facial features and clothing.
  • + The hairstyle is very grounded and realistic, matching the character's age and style.
  • + Seamless integration with the existing lighting and background.
  • The hair replacement slightly alters the shape of the forehead compared to the original.

Vidu Q2

  • + Successfully adds a full head of hair with thick texture.
  • + Preserves the majority of the background and clothing.
  • The hairstyle is overly voluminous and looks slightly unnatural/AI-generated.
  • The hair rendering looks like a flat overlay in some areas, losing the fine detail seen in the original beard.
  • The forehead and hairline transition is less realistic than the other model.

Verdict: Grok Imagine Image Pro is the winner because it provides a highly realistic, naturally integrated head of hair that perfectly matches the character's existing features and lighting. Vidu Q2 adds more volume, but the result looks stylized and slightly less cohesive with the original photography.

Isometric Miniature Diorama Scenes

Text-to-Image

“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”

Grok Imagine Image Pro
Vidu Q2

AI Judge Analysis

Grok Imagine Image Pro

  • + Excellent 3D material rendering with realistic PBR subsurface scattering on the fish.
  • + Clean, professional typography that feels integrated with the scene.
  • + High geometric clarity and consistent lighting across the sushi pieces.
  • The 'diorama base' is a bit simple, appearing more like a standard wooden tray.
  • The camera angle is slightly more tilted than a true 45° isometric view.

Vidu Q2

  • + Stronger adherence to the 'diorama base' prompt with a stylized raised platform.
  • + Good interpretation of the miniature cartoon aesthetic.
  • + Vibrant colors and creative sushi toppings.
  • The Japanese flag icon is awkwardly floating/attached to a letter.
  • The text rendering is slightly less refined with some aliasing on edges.
  • The sushi models look slightly more 'plastic' compared to the high-quality textures in Model A.

Verdict: Both models followed the prompt well, but Grok Imagine Image Pro produced a significantly more polished result with superior PBR textures and professional-grade typography. While Vidu Q2 captured the 'diorama' concept slightly better, its execution of the text and the hanging flag icon was less clean than Grok's output.

Over-the-top cartoon caricature

Editing
Edit instruction

“Create a caricature of me and my job. Make it exaggerated and humorous, incorporating my profession as a tv show anchor and my love for dogs and hockey.”

Source
Grok Imagine Image Pro
Vidu Q2

AI Judge Analysis

Grok Imagine Image Pro

  • + Excellent adherence to the 'caricature' style with exaggerated facial features.
  • + Includes many specific details like the hockey trophy, sticks, pucks, and numerous dogs.
  • + Clear, legible text that complements the theme.
  • The facial likeness to the source image is somewhat lost in the extreme exaggeration.
  • The composition feels a bit crowded with many competing elements.

Vidu Q2

  • + Maintains a strong facial likeness to the original subject while still being a caricature.
  • + Effectively preserves the original clothing (denim shirt and black top).
  • + Clean, vibrant illustration style.
  • The hands have anatomical errors, particularly the extra finger on the right hand.
  • The background hockey rink and goal are less detailed and humorous than the studio setting in Model A.

Verdict: Grok Imagine Image Pro delivers a more successful 'humorous and exaggerated' caricature by leaning into the typical distorted features of the genre and adding many clever thematic details. Vidu Q2 does a better job of preserving the source person's likeness and outfit, but it suffers from anatomical issues with the hands and a less engaging composition.

Adorable Baby Animals in Sunny Meadow

Text-to-Image

“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”

Grok Imagine Image Pro
Vidu Q2

AI Judge Analysis

Grok Imagine Image Pro

  • + Excellent depiction of god rays and sunrise atmosphere
  • + Highly detailed and consistent fur texture
  • + Expressive and distinct character interactions that feel playful
  • Included two kittens instead of the single one implied by the prompt
  • The positioning of the animals feels slightly static compared to the action-oriented prompt

Vidu Q2

  • + Dynamic composition with a sense of movement and tumbling
  • + Great variety of colorful butterflies integrated into the scene
  • + Lush and vibrant wildflower meadow with beautiful bokeh
  • Inconsistent animal anatomy, especially the extra leg/tail on the golden retriever
  • Features two retriever puppies instead of the single one requested

Verdict: Grok Imagine Image Pro produces a much cleaner, more professional-looking image with superior lighting and fewer anatomical errors, though it doubles the kittens. Vidu Q2 captures the energetic 'tumbling' aspect of the prompt better but suffers from significant biological artifacts and AI hallucinations. Grok is the preferred choice for its higher visual fidelity and polished fur textures.

Studio Ghibli Anime Style

Editing
Edit instruction

“Transform this photo into a Studio Ghibli–inspired illustration. Use soft pastel colors, hand-painted textures, gentle lighting, dreamy backgrounds, and a warm, nostalgic mood”

Source
Grok Imagine Image Pro
Vidu Q2

AI Judge Analysis

Grok Imagine Image Pro

  • + Captures the characteristic Ghibli watercolor texture and soft edge bleeding perfectly.
  • + Maintains the distinct facial expressions and poses of the original meme characters while translating them to anime.
  • + Features a much softer, cohesive color palette that aligns with the 'warm, nostalgic' request.
  • Slightly less clarity in the fine line work compared to traditional cel-shaded anime.
  • Background feels a bit more abstract/blurry than typical Ghibli scenic art.

Vidu Q2

  • + Strong, clean line work that mimics modern anime production styles.
  • + The background architectural details are more defined and visible.
  • + Accurately replicates the character positions and clothing patterns from the source.
  • Lacks the specific 'hand-painted' watercolor texture associated with the Ghibli aesthetic.
  • Colors are slightly too vibrant and digital, missing the 'pastel' and 'nostalgic' mood of the prompt.
  • The woman on the right has slightly distorted facial proportions.

Verdict: Grok Imagine Image Pro is the clear winner as it successfully captures the specific 'hand-painted watercolor' texture and soft lighting characteristic of Studio Ghibli. While Vidu Q2 produces a high-quality anime illustration, it feels more like a standard digital cel-shaded style and misses the 'dreamy, nostalgic' mood requested in the prompt.

Golden Hour Stroll

Image Editing
Edit instruction

“Add dynamic motion to this photo: make hair blow in the wind, add leaves flying, energetic and lively feel.”

Before After
Grok Imagine Image Pro
Before After
Vidu Q2

AI Judge Analysis

Grok Imagine Image Pro

  • + Excellent preservation of the subject's facial features and the dog's appearance.
  • + Realistic hair movement and strand detail.
  • + Leaves blend naturally with the background lighting.
  • Some leaves appear a bit small and repetitive in shape.
  • The leash handle area has slight artifacts where leaves overlap.

Vidu Q2

  • + Stronger sense of depth with larger, foreground leaves.
  • + Successfully recreates the hair movement while maintaining image quality.
  • + Preserves the overall composition of the original photo well.
  • The color of the leaves is significantly more orange/red, which slightly shifts the seasonal feel from the original.
  • Minor distortion to the subject's left eye compared to the source.

Verdict: Both models followed the instructions very well, effectively adding wind-blown hair and flying leaves while keeping the original identities intact. Grok Imagine Image Pro is the preferred choice because it managed to add the motion effects with near-perfect preservation of the woman's face and the dog, whereas Vidu Q2 slightly altered the expression and eyes.

Vintage Cafe Logo

Text-to-Image

“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”

Grok Imagine Image Pro
Vidu Q2

AI Judge Analysis

Grok Imagine Image Pro

  • + Excellent typography including the grave accent in 'Caffè'
  • + Clean vector emblem style with balanced composition
  • + Accurate rendering of the cloche and banner exactly as requested
  • The swirl of the steam is a bit stylized compared to a realistic vapor

Vidu Q2

  • + Features a nice subtle texture on the background
  • + The cloche illustration has pleasant shading
  • Severe spelling errors in the main text and banner
  • Cluttered composition with redundant and mangled text
  • Included a glass-like dome rather than a traditional metal cloche

Verdict: Grok Imagine Image Pro followed the prompt instructions near-perfectly, delivering a clean, professional vector logo with correct spelling and a logical layout. Vidu Q2 failed significantly on text rendering, producing several misspelled variations of the prompt's text and a cluttered design.

Apollo 11: Journey to Tranquility

Text-to-Image

“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”

Grok Imagine Image Pro
Vidu Q2

AI Judge Analysis

Grok Imagine Image Pro

  • + Excellent typography with perfect spelling of mission steps and astronaut names.
  • + Strict adherence to the requested NASA color palette and clean flat-vector style.
  • + Logical vertical flow that clearly represents the six specific stages requested.
  • The lunar module icon in step 6 is slightly generic compared to real NASA hardware.

Vidu Q2

  • + Features more detailed 3D-shaded icons for the lunar module.
  • + Good use of the pin icon to indicate the landing site.
  • Severe spelling hallucinations including 'ALFONCH' and 'Aeutth Orot'.
  • Failure to follow the numerical sequence of the six requested steps.
  • Cluttered composition with inconsistent icon spacing and confusing labels.

Verdict: Grok Imagine Image Pro successfully followed the complex infographic instructions, delivering a professional, accurately spelled, and well-organized timeline. Vidu Q2 struggled significantly with text rendering and the logical flow of the mission steps, resulting in gibberish text and a confusing layout.

Next steps

Explore each model