Alibaba's Qwen Image 2.0 model with enhanced text rendering, supporting both Chinese and English prompts with up to 6 images per request
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
Qwen Image 2.0
#34 of 62 in Text-to-Image
Vidu Q2
#42 of 62 in Text-to-Image
Where the votes landed
Qwen Image 2.0
0%
win rate
Ties
0%
Vidu Q2
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
Qwen Image 2.0
- + Photorealistic materials and lighting textures
- + Accurate representation of glass reflections and window light direction
- + High naturalistic detail on the red book and wooden table
- − The blue sphere is levitating rather than sitting on a surface
- − The glass cube has internal vertical dividers that weren't requested
Vidu Q2
- + Perfect adherence to all spatial relationships including the sphere sitting inside
- + Excellent handling of shadows and light play through the glass
- + Complex plant details visible through the glass as requested
- − The lighting is a bit harsh compared to 'soft' window light
- − The book binding has slightly anomalous text-like artifacts
Verdict: Both models followed the prompt exceptionally well, but Vidu Q2 is the winner for its superior physics and spatial logic, placing the sphere on the floor of the cube rather than levitating it. While Qwen Image 2.0 has a very pleasing soft-focus aesthetic, it added extra glass panes inside the cube that were not requested.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent skin texture with realistic age spots and wrinkles
- + Strong composition that shows the character's face clearly
- + Accurate depiction of wet pavement and reflections
- − The hands interacting with the pedal are anatomically confusing
- − Very little rain is visible in the air, though the ground is wet
Vidu Q2
- + Effective 'imperfect framing' that heightens the candid feel
- + Impressive detail on the bicycle's mechanical parts and wear
- + Good integration of the background car and street lights
- − The man's face is largely cut off and obscured
- − The hand anatomy is quite distorted and messy
- − The perspective of the bicycle frame is slightly warped
Verdict: Qwen Image 2.0 provides a much better character portrait with high-quality skin textures and a clear focal point. While Vidu Q2 captures the 'imperfect framing' mentioned in the prompt more literally, it suffers from significant anatomical distortions in the hands and poor character visibility.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent depiction of colored beads in the braids
- + Strong gritty realism with visible dirt and skin texture
- + Dynamic lighting with glowing orange embers and reflections
- − The character looks more like a modern mercenary than a traditional 'paladin'
- − Hand anatomy is slightly awkward on the sword hilt
Vidu Q2
- + Ornate engraving on the plate armor is highly intricate
- + Texture on the cloth underlayer is exceptionally clear and detailed
- + Superior lighting consistency across the face and metal
- − The 'hair braided with small beads' prompt is only partially met with just a few beads
- − The background bokeh is less vibrant compared to Model A
Verdict: Qwen Image 2.0 captures the 'battle-worn' and 'beaded hair' aspects of the prompt more effectively, resulting in a character with more depth and history. However, Vidu Q2 offers superior technical execution in the armor engravings and the specific texture of the textile underlayers, making it the more visually polished image despite missing some character details.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent photo-realistic food quality
- + Clearly legible section headers for Appetizers, Pizza, and Mains
- + Extremely clean and professional grid layout
- − Sub-text for individual dishes is largely gibberish
- − The grid layout doesn't follow the section headers logically (pizzas under 'Mains')
- − Price values are repetitive and unrealistic
Vidu Q2
- + Dynamic and modern use of color accents
- + More complex layout variety resembling a multi-page menu
- + Good use of white space and hierarchy
- − Serious text rendering issues with misspellings in main headers (e.g., 'Apecizen')
- − Food photography is cluttered and looks less appetising than the other model
- − Poor source preservation of logic, leading to a confusing visual mess
Verdict: Qwen Image 2.0 produces a far superior result by prioritizing clarity, high-quality food photography, and a clean professional grid that matches the modern minimalist prompt. Vidu Q2 struggles significantly with text legibility and coherent layout, resulting in a cluttered design that would be unusable as a real menu.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent text legibility and high-quality rendering of the fiery effect.
- + Photorealistic textures on the patty and vegetables.
- + Good use of negative space and smoke for a cinematic feel.
- − The burger is not fully 'exploded'; the bottom half is mostly stacked together.
- − The sauce is dripping down rather than being suspended in mid-air as requested.
Vidu Q2
- + Better 'exploded' composition with suspended sauce droplets and separated layers.
- + The background captures a more intense, fiery atmosphere.
- + Correct adherence to the starburst element for the price.
- − The currency symbol is incorrect, showing a pound-like hybrid instead of an Euro sign.
- − The lighting on the internal burger components is slightly flat compared to the bun top.
- − Perspective on the bottom bun is slightly warped.
Verdict: Qwen Image 2.0 produces a more professional-looking advertisement with superior font rendering and photorealistic lighting, but it fails to 'explode' the burger as much as requested. Vidu Q2 follows the 'suspended in mid-air' instruction better for all ingredients, including sauce, but fails on the specific currency symbol. Qwen Image 2.0 is the overall winner for its commercial-grade polish and perfect text accuracy.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent text accuracy with no spelling errors.
- + Consistent and realistic chalk handwriting texture.
- + Natural-looking lighting and background blur.
- − The 'Brown Butter' item is truncated at the end of the line.
Vidu Q2
- + Strong chalk texture and coloring.
- + Centered composition on the chalkboard.
- − Significant spelling errors throughout the menu items.
- − Garbled and illegible text at the bottom of the board.
- − Changed the price of the first item from $24 to $34.
Verdict: Qwen Image 2.0 followed the prompt with extreme accuracy, rendering all requested text perfectly with a very realistic chalk aesthetic. While it truncated the final cookie item, it is far superior to Vidu Q2, which suffered from numerous spelling errors and completely illegible text at the bottom. Qwen's coherent handwriting and correct spelling make it the clear winner.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent anatomical clarity of the horse and astronaut
- + Clean, high-resolution textures on the spacesuit and horse pelt
- + Cinematic lighting with a clear horizon line of Earth
- − Failed the specific spatial instruction 'horse on top'
- − The floating droplets look slightly disconnected from the scene
Vidu Q2
- + Beautiful surreal aesthetic with galaxy patterns on the horse
- + Dynamic composition with vibrant colors and cosmic dust
- + Highly detailed bridle and saddle equipment
- − Failed the specific spatial instruction 'horse on top'
- − Occasional artifacts in the cosmic background clouds
Verdict: Both models failed the negative constraint to have the horse on top of the astronaut, instead providing the traditional 'astronaut riding horse' configuration. Qwen Image 2.0 provides a sharper, more photorealistic cinematic look, whereas Vidu Q2 leans more into the surreal prompt with its cosmic horse textures and ethereal colors.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent fur texture on the capybara.
- + Lighting feels cohesive with a nighttime city environment.
- + Strong adherence to the professional and calm expression requested.
- − The human passenger appears to be in the front passenger seat rather than the back seat.
- − The hand rendering on the steering wheel is anatomically messy/melted.
Vidu Q2
- + Superior perspective that clearly shows both the driver and the passenger in the back seat.
- + Crisp and realistic rendering of the vehicle interior and city lights.
- + Accurate interpretation of 'sitting in the back seat' versus the front-seat placement in Model A.
- − The capybara's hands look slightly more human-like than actual paws.
- − The human face in the background is a bit soft in detail.
Verdict: Vidu Q2 is the winner because it correctly placed the passenger in the back seat as requested, whereas Qwen Image 2.0 placed her in the front next to the driver. Vidu Q2 also offers a much better composition and clearer interior details, despite the slightly less detailed fur texture on the animal.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
Qwen Image 2.0
- + Perfect text rendering for all requested strings.
- + Excellent composition with clear cinematic lighting.
- + High-quality gothic border detail that matches the prompt exactly.
- − None notable.
Vidu Q2
- + Captures the requested border elements including thorns and webs.
- + Includes multiple bats as requested.
- − Significant spelling errors throughout all text layers.
- − The date and time details are incorrect or illegible (e.g., 2025 instead of 2026, 'Tmm' instead of '7pm').
- − Lower visual polish and flatter lighting compared to the opponent.
Verdict: Qwen Image 2.0 is the clear winner as it followed every instruction perfectly, specifically excelling in text accuracy which is vital for an invitation. Vidu Q2 failed on almost every text string, producing numerous typos and an incorrect year, while also offering a less sophisticated artistic style.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent typography with clean, bold font rendering
- + High photo-realism in textures and materials
- + Accurate 45-degree isometric-style camera angle
- − Missed the 'cartoon' aesthetic, opting for a realistic food photography style instead
- − The diorama base looks like a standard cutting board rather than a stylized base
Vidu Q2
- + Perfectly captures the 'cartoon' and 'miniature 3D diorama' aesthetic requested
- + Creative integration of the flag icon onto the text
- + Smooth PBR-style lighting and soft textures
- − Text rendering is slightly less sharp than Model A
- − The sushi components look a bit cluttered compared to the 'minimal garnish' request
Verdict: While Qwen Image 2.0 produced a high-quality, realistic image with perfect text, it largely ignored the 'cartoon' and '3D miniature' stylization requested in the prompt. Vidu Q2 followed the stylistic instructions much better, creating an output that looks like a 3D rendered diorama with the requested soft, playful textures.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent adherence to the specific animal list, featuring exactly one of each requested species.
- + Strong emotional interaction between the animals, creating a literal 'tumbling together' scene.
- + Highly detailed fur textures and sharp focus on the central characters.
- − The fox kit is positioned in a way that its anatomy looks slightly confusing underneath the other animals.
- − The lighting is a bit harsh on the background grass compared to the soft forest floor aesthetic.
Vidu Q2
- + Beautiful, whimsical composition with vibrant colors and many fluttering butterflies.
- + Captures a great sense of motion and playfulness across the entire frame.
- + Excellent lighting with soft bokeh and sparkling dew effects that enhance the 'golden hour' feel.
- − Failed to follow the specific animal count, including two retriever puppies instead of one.
- − The rabbit has somewhat stylized/elongated ears that look a bit less photorealistic than other elements.
Verdict: Qwen Image 2.0 followed the prompt instructions precisely, including one of each specific animal and creating a touching, intertwined composition. Vidu Q2 produced a more visually energetic and colorful scene but failed the counting task by adding an extra puppy, and its level of photorealism was slightly more 'AI-smooth' than the textured approach of Qwen Image 2.0.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent text rendering with no spelling errors
- + Clean vector-style composition
- + High contrast and strong logo characteristics
- − Steam looks more like a flame icon
- − Banner scroll on the right is a bit clunky
Vidu Q2
- + Elegant warm brown and cream color palette
- + Beautiful subtle paper texture
- + Graceful steam illustration
- − Frequent spelling errors in both names and dates
- − Redundant text elements cluttering the design
- − Inconsistent line weights
Verdict: Qwen Image 2.0 is the clear winner because it successfully renders all requested text accurately and follows the vector emblem style instructions. Vidu Q2 has a more sophisticated color palette and texture, but fails significantly on text legibility and accuracy, which is crucial for a logo prompt.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
Qwen Image 2.0
- + Strong text legibility with mostly correct spelling including 'Launch', 'Earth Orbit', and crew names.
- + Followed the requested vertical infographic layout logically.
- + Accurate representation of the Apollo 11 Lunar Module and color palette.
- − Includes a typo in 'Translunjar'.
- − The Saturn V rocket icon is very small and lacks detail.
Vidu Q2
- + Excellent flat-vector aesthetic with clean, consistent iconography.
- + Captures the NASA-inspired color palette perfectly.
- + Creative use of icons for the trajectory and astronauts.
- − Text is largely gibberish or contains significant misspellings (e.g., 'Alfonch', 'Laup', 'Lunan Orutt').
- − Failed to follow the requested 6-step logical sequence accurately.
Verdict: Qwen Image 2.0 is the superior choice for an infographic because it provides readable, accurate text that conveys the actual mission information requested. While Vidu Q2 has a very polished vector art style, its inability to spell basic words or follow the specific logical steps makes it unusable as an informational poster.
Explore each model
ShengShu Technology's text-to-image and reference-to-image model with support for character consistency and multi-reference image processing