Improved version of Alibaba's Qwen image model with better text rendering, finer natural textures, and more realistic human generation.
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
Qwen Image 2512
#30 of 62 in Text-to-Image
Vidu Q2
#42 of 62 in Text-to-Image
Where the votes landed
Qwen Image 2512
0.0%
win rate
Ties
0.0%
Vidu Q2
100.0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
Qwen Image 2512
- + Excellent soft lighting consistent with the prompt
- + High-quality textures on the red book and wooden table
- + Accurate glass reflections and refractive effects
- − The glass cube has a mirrored base which wasn't explicitly requested
- − The plant in the background is very blurry
Vidu Q2
- + Stronger visual coherence of the plant through the glass
- + Dynamic lighting and shadow patterns on the table
- + Includes detailed spine text on the book showing high resolution
- − The sphere is hovering slightly off the bottom surface
- − The lighting is more direct and harsh than the 'soft' request
Verdict: Both models followed the complex spatial instructions perfectly. Qwen Image 2512 captured the 'soft' lighting request better with a more photorealistic render of the cube's materials, whereas Vidu Q2 provided a sharper background and more interesting shadow play but struggled slightly with the sphere's contact point on the surface.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
Qwen Image 2512
- + Excellent skin texture and facial details
- + Accurately represents the 50mm lens and shallow depth of field requested
- + Balanced, cinematic composition that maintains a candid feel
Vidu Q2
- + Successfully captures the 'imperfect framing' prompt with a tight, candid angle
- + Strong portrayal of wet pavement reflections and light rain
- − Anatomical errors in the hands and excessive skin wrinkles make it look stylized
- − Physically impossible bicycle structure with overlapping parts and broken frames
- − Lack of background motion blur compared to Model A
Verdict: Qwen Image 2512 is the superior model as it produces a much more realistic and anatomically correct image with high-quality textures. While Vidu Q2 followed the 'imperfect framing' instruction well, the technical execution of the bicycle and the man's hands was poor, appearing distorted and overly stylized.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
Qwen Image 2512
- + Excellent depiction of skin texture, fine pores, and realistic scars.
- + Superior integration of warm torchlight reflections on the metal armor.
- + Adheres better to the 'close portrait' request with a compelling, detailed facial expression.
- − The bokeh sparks are a bit large and localized compared to a natural distribution.
- − Braid physics near the shoulders look slightly stiff.
Vidu Q2
- + Ornate engraving on the armor is very intricate and well-defined.
- + Cloth texture on the underlayer shows high fidelity and clear material definition.
- + Dynamic lighting creates a strong sense of volume and depth.
- − The facial skin appears a bit too smooth and clean for a 'battle-worn' character.
- − The composition is more of a medium shot than the requested 'close portrait'.
Verdict: Qwen Image 2512 is the winner because it captures the 'battle-worn' essence much more effectively with realistic skin imperfections, grit, and an intense close-up composition. Vidu Q2 offers beautiful armor details and lighting, but the character looks too pristine and the framing is wider than requested.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
Qwen Image 2512
- + Features a clear grid layout of food photos as requested.
- + Organizes sections clearly with distinct headers and colorful markers.
- + Font choice is bold and legible for a menu header.
- − Contains significant spelling errors in headers like 'APPETIIZIZERS' and 'MEANS'.
- − Some food items in the grid look repetitive or lack distinct variety.
Vidu Q2
- + Captures a more playful, modern aesthetic with vibrant photographic accents.
- + Includes the specific 'Pizza' section mentioned in the prompt.
- + Has a high-quality visual gloss and a clean white background.
- − The text is highly garbled and unreadable across the entire image.
- − The grid layout is less structured and more organic than the requested grid.
Verdict: Qwen Image 2512 followed the layout requirements more closely by providing a structured grid and clear sectional divisions, despite some spelling errors. Vidu Q2 captured the 'vibrant' and 'modern' feel well but failed significantly on text legibility and the specific grid structure requested. Qwen Image 2512 is the better overall design for a menu template.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
Qwen Image 2512
- + Excellent typography with a glowing fire texture integrated into the 'MAGIC BURGER' title.
- + High photorealistic detail in the food textures, especially the patty and seeds.
- + Effective use of dynamic particles and flying ingredients to create a sense of motion.
- − The 'LIMITED TIME ONLY' text lacks the glowing effect compared to the main title.
Vidu Q2
- + Strong fiery atmosphere that covers the entire background.
- + Successful integration of all requested text elements, including a bright starburst.
- − The currency symbol is incorrect, showing a distorted 'E' or Hash symbol instead of the requested Euro '€'.
- − The food rendering looks slightly more plastic/saturated and less photorealistic than Image A.
- − The 'LIMITED TIME ONLY' text is placed too close to the main title, causing visual clutter.
Verdict: Qwen Image 2512 is the clear winner due to its superior photorealistic rendering of the burger and more professional graphic design layout. While Vidu Q2 captures the fiery energy well, it fails on the specific currency request and has a less convincing 'exploded' perspective for the ingredients.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
Qwen Image 2512
- + Excellent text legibility and spelling throughout the entire board.
- + Successfully rendered a realistic chalk texture with natural-looking cursive variation.
- + Followed the prompt instructions carefully, completing the 'Brown Butter' item logically.
- − One minor spelling error in 'Risitto' (should be Risotto).
- − The handwriting looks a bit too clean and calculated for a hand-drawn board.
Vidu Q2
- + Captures a very authentic, messy chalk smudging and layering aesthetic.
- + Good variation in stroke weight and pressure that mimics real chalk well.
- − Severe spelling and literacy issues, with many words becoming nonsensical (e.g., 'Octopd wpiln').
- − The text layout becomes cluttered and overlapping toward the bottom.
- − Failed to maintain coherent prices and item names.
Verdict: Qwen Image 2512 is the clear winner as it produces a professional, legible, and aesthetically pleasing chalkboard menu that follows the prompt almost perfectly. While Vidu Q2 captures a more realistic 'smudged' chalk texture, its inability to maintain correct spelling or word structure makes it unusable for a text-based challenge.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
Qwen Image 2512
- + Excellent photorealism in the space suit and horse textures.
- + High cinematic quality with realistic lighting and earth-horizon background.
- + Clear, sharp details on the astronaut's face and visor.
- − Failed the negative constraint; the astronaut is riding the horse, not the other way around.
- − The horse's anatomy is slightly distorted with a missing or warped hind leg area.
Vidu Q2
- + Vibrant colors and a very creative 'nebula' horse design.
- + Good adherence to the 'surreal' aesthetic requested in the prompt.
- + Ornate space suit design with consistent stylization.
- − Failed the crucial negative constraint; the astronaut is riding the horse as per the common trope.
- − The horse's legs are stylistically melted or indistinct at the hooves.
- − Lower level of realism compared to the cinematic request.
Verdict: Both Qwen Image 2512 and Vidu Q2 failed the logic trap in the prompt which requested the horse to be on top of the astronaut. Qwen Image 2512 is the superior image in terms of technical execution and cinematic lighting, whereas Vidu Q2 leans more into a generic colorful digital art style.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
Qwen Image 2512
- + Excellent central composition through the windshield.
- + The capybara's expression is perfectly calm and professional as requested.
- + Realistic bokeh and light reflections on the windshield.
- − The capybara's paws look slightly human-like or primate-like.
- − Only the top half of the taxi is visible.
Vidu Q2
- + Dynamic side-profile composition that shows more of the taxi's interior.
- + The capybara's anatomy (paws and head shape) is very accurate.
- + Captures the 'bored' expression of the businesswoman effectively.
- − Significant architectural artifact where the taxi's door frame and roof meet behind the driver.
- − The passenger is holding her phone in a slightly awkward way.
Verdict: Both models followed the prompt exceptionally well, capturing the surreal humor of the scene with high photo-realism. Qwen Image 2512 is slightly preferred for its more cinematic, centered composition and the perfect 'professional' look of the capybara, whereas Vidu Q2 suffered from a noticeable structural artifact in the car's interior pillars.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
Qwen Image 2512
- + Excellent typography with only one minor misspelling
- + Highly atmospheric cinematic lighting and composition
- + Strong adherence to the border and twisted tree elements
- − Spells Halloween as 'Hallowern'
- − Missing the dark parchment texture requested in the background
Vidu Q2
- + Successfully incorporated the dark parchment texture
- + Good use of the scroll banner element
- − Multiple significant spelling errors and gibberish text
- − Incorrect date (2025 instead of 2026) and time data
- − Visual composition feels cluttered and less polished than the competitor
Verdict: Qwen Image 2512 is the clear winner due to its superior visual quality, atmospheric lighting, and readable typography, despite a single character misspelling in the title. Vidu Q2 followed the 'parchment' instruction better but failed significantly on text rendering, resulting in several illegible or incorrect words and dates.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
Qwen Image 2512
- + Excellent 3D miniature diorama feel with a tangible raised base
- + Perfect text rendering and layout for the title elements
- + Superior material quality, especially the rice grains and fish textures
- − The square base is slightly rotated rather than being perfectly isometric in alignment
Vidu Q2
- + Clean, high-contrast text rendering
- + Good adherence to the solid light blue background requirement
- + Successful 45-degree isometric projection
- − The textures look overly plastic and lack the requested realistic PBR quality
- − The flag icon is floating awkwardly on a pole rather than being integrated as a UI element
- − Details on the sushi, such as the shrimp and the Gunkan maki, are a bit messy
Verdict: Qwen Image 2512 is the clear winner as it perfectly captures the 'miniature 3D cartoon' aesthetic with high-quality PBR textures that make the sushi look both stylized and appetizing. While Vidu Q2 follows the layout instructions well, its rendering quality is much lower, appearing flat and lacking the refined detail found in Qwen's rice and garnish.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
Qwen Image 2512
- + Excellent anatomical details and fur texture on all animals
- + Perfect lighting with soft god rays and clear dew sparkles
- + Strong composition with a focused, wholesome family-portrait style
- − Static posing fails to capture the 'chasing' and 'tumbling' action requested
- − The scale of the butterflies is slightly too large compared to the animals
Vidu Q2
- + Successfully captures the dynamic 'chasing' and 'tumbling' movement
- + Rich, immersive environment with a high volume of wildflowers and butterflies
- + Excellent use of wide composition to show the full scale of the scene
- − Includes two puppies instead of the requested single golden retriever
- − The rabbit's anatomy is slightly distorted with a cat-like tail
- − Animals are less 'hyper-photorealistic' and look more like digital paintings
Verdict: Qwen Image 2512 produces a much higher quality image in terms of texture, lighting, and realistic details, though it opts for a static pose. Vidu Q2 better captures the energetic action of 'chasing and tumbling' described in the prompt but suffers from counting errors (two puppies) and less realistic animal anatomy.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
Qwen Image 2512
- + Perfect text rendering of the name and establishment date
- + Excellent vintage woodcut/vector aesthetic that fits the 'classic' theme
- + High visual clarity and professional composition
- − The detail level is slightly higher than 'minimalist', bordering on an illustrative badge style
Vidu Q2
- + Captures a cleaner minimalist vector style
- + Follows the color palette accurately
- − Significant spelling errors in the brand name ('Farmiin', 'Fopli20')
- − Poorly rendered steam that looks like a flame or icon artifact inside the dome
- − Awkward layout with repetitive text strings
Verdict: Qwen Image 2512 is the clear winner as it produces perfect typography and a cohesive, professional logo design. Vidu Q2 fails on basic text generation and suffers from internal visual incoherence with the steam and cloche rendering.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
Qwen Image 2512
- + Features highly detailed and accurate illustrations of the Saturn V and Lunar Module.
- + Adheres well to the requested navy, white, and muted red color palette.
- + Successfully captures a professional infographic aesthetic with a clear visual hierarchy.
- − Step numbering is repetitive and logically confused (two step 2s, two step 3s).
- − Text becomes garbled in the middle section (e.g., 'Translaurtcit').
Vidu Q2
- + Closer to a true flat-vector style with consistent line weights across icons.
- + Contains very few distracting gradients, maintaining a clean 'sticker' look.
- + Maintains a consistent layout for the steps.
- − Fails significantly on text rendering, including the main header ('ALFONCH').
- − Icon choices are less accurate to NASA history, using a generic toy-like rocket instead of a Saturn V.
- − Logic of the mission steps is disorganized and contains redundant icons.
Verdict: Qwen Image 2512 is the superior choice because it provides high-quality, recognizable illustrations of the specific Apollo hardware requested, whereas Vidu Q2 uses generic icons that do not match the historical context. While Qwen Image 2512 has issues with step numbering, its visual polish and adherence to the NASA-inspired theme make it a much more effective infographic.
Explore each model
ShengShu Technology's text-to-image and reference-to-image model with support for character consistency and multi-reference image processing