Alibaba's Qwen image model
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
Qwen Image
#37 of 62 in Text-to-Image
Vidu Q2
#42 of 62 in Text-to-Image
Where the votes landed
Qwen Image
0%
win rate
Ties
0%
Vidu Q2
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
Qwen Image
- + Excellent photographic quality with soft, realistic lighting.
- + Clean and balanced composition.
- + Follows all spatial instructions including the sphere placement and book.
- − The plant is blurred in the background and not very visible through the glass as requested.
Vidu Q2
- + Successfully captures the plant being visible through the glass cube.
- + Realistic wood texture and lighting shadows.
- + Accurate object placement.
- − The book and sphere have a slightly more 'rendered' look compared to the photo-realism of Model A.
- − The glass cube has some thickness inconsistencies on the top edge.
Verdict: Both models followed the complex spatial instructions perfectly. Qwen Image produced a more aesthetically pleasing, professional photograph look, while Vidu Q2 better captured the specific detail of the plant being partially visible through the glass medium.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
Qwen Image
- + Excellent full-body composition that captures the environment
- + Realistic motion blur on the background vehicles
- + Effective use of shallow depth of field
- − Anatomy errors including a three-legged bicycle stand and floating pedal
- − Faces and hands are a bit soft and lacking fine texture
Vidu Q2
- + Highly detailed 'imperfect framing' that feels like a true candid close-up
- + Superb skin texture on the hands and face
- + Complex mechanical details on the bicycle chain and frame
- − Distorted hands with anatomical blending issues
- − Background car is static despite the request for motion blur
Verdict: Qwen Image follows the overall scene description better, providing a full-bodied shot with the requested motion blur, though it suffers from structural errors in the bicycle. Vidu Q2 excels in textural realism and lighting, but the lack of motion blur and the presence of significant anatomical merging in the hands makes it slightly less successful as a cohesive image.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
Qwen Image
- + Excellent adherence to the 'beads in hair' prompt with colorful details.
- + Strong atmospheric lighting with visible torchlight and stylized sparks.
- + Intricate engraving details on the armor and clear texture on the underlayer.
- − The facial scars look somewhat like artificial paint or fresh cuts rather than 'faint scars'.
- − The torch in the foreground creates some visual clutter.
Vidu Q2
- + Very realistic skin textures and more natural-looking faint scars.
- + Sophisticated armor engraving and better structural consistency of the chest plate.
- + Cleaner overall composition with a more effective shallow depth of field.
- − The 'braided with small beads' instruction is barely followed, with very few beads visible.
- − The bokeh effects in the background are less dynamic than requested.
Verdict: Qwen Image followed the specific details of the prompt much better, particularly the colorful hair beads and the 'battle-worn' aesthetic. While Vidu Q2 has a more realistic skin finish and higher anatomical quality, it missed the primary decorative elements requested for the hair/beads. Therefore, Qwen Image is preferred for its superior prompt adherence.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
Qwen Image
- + Strong minimalist aesthetic with a very clean white background.
- + Excellent font choice that matches the 'bold sans-serif' prompt perfectly.
- + The image grid is perfectly aligned and creates a high-quality professional look.
- − Text contains frequent gibberish and typos like 'Pizzaurant'.
- − The food variety is limited, with several images looking very similar.
Vidu Q2
- + Features a wider variety of food photos and more comprehensive menu content.
- + Follows the multi-section categorization requested in the prompt.
- + Contains vibrant accents and more complex design elements.
- − The layout is cluttered and fails the 'minimalist' requirement.
- − Text rendering is very poor with extreme distortion and illegible characters.
- − Some food items look messy or unappetizing due to low-quality AI generation.
Verdict: Qwen Image is the clear winner as it successfully captures the 'modern minimalist' aesthetic with a professional layout and high-quality typography. While Vidu Q2 includes more sections, the composition is cluttered and the text quality is significantly worse than Qwen Image's bold, clean presentation.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
Qwen Image
- + Excellent typography with a 3D neon-style effect
- + High resolution with clean, sharp food photography aesthetics
- + Well-composed starburst sticker that fits the ad layout
- − The main burger is mostly assembled rather than fully 'exploded' as requested
- − Some floating sauce blobs look slightly artificial
Vidu Q2
- + Better 'exploded' effect with dramatic spacing between all layers
- + Fiery, glowing text effect matches the prompt requirements perfectly
- + Dynamic sense of motion with sauce splashes and embers
- − The currency symbol is incorrect (rendered as an 'E' or '#' cross instead of '€')
- − The bottom bun texture looks slightly less photorealistic than the top
Verdict: Qwen Image produces a cleaner, more professional-looking commercial image with superior text rendering and currency accuracy. However, Vidu Q2 followed the 'exploded' instruction much better, creating a more dynamic sense of motion and a more literal interpretation of the 'fiery' text effect, despite the error in the currency symbol.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
Qwen Image
- + Excellent text legibility and spelling accuracy for almost all required items.
- + Clean and realistic chalkboard aesthetic with smudges and natural wood frame.
- + Good layout and spacing between menu items.
- − Year is incorrectly rendered as 20026 instead of 2026.
- − Text looks slightly more like a digital font than authentic chalk handwriting.
Vidu Q2
- + Authentic chalk texture with varied stroke thickness and realistic dustiness.
- + Followed the year prompt correctly (2026).
- + Captures the 'handwritten' request more effectively with natural slants and variations.
- − Significant spelling errors and 'AI gibberish' in the menu items (e.g., 'Truffe Musshoom', 'Octopd wpiln').
- − Inconsistent pricing characters and messy footer text.
Verdict: Qwen Image provides a much more professional and legible menu, though it failed on the specific year requested (20026). Vidu Q2 captures the requested chalk texture and artistic style better, but the text is riddled with spelling errors and garbled words, making it less usable overall. Qwen Image is the preferred choice for its clarity and nearly perfect adherence to the food descriptions.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
Qwen Image
- + Clean, cinematic composition with realistic lighting for the astronaut's suit.
- + High resolution with sharp details on the planetary background.
- − Failed the specific prompt constraint of having the horse on top of the astronaut.
- − Anatomical issues with the horse's legs, specifically the rear leg merging into the body.
Vidu Q2
- + Beautiful surreal aesthetic with a cosmic, ethereal horse design.
- + Rich colors and intricate details in the nebulae and starry background.
- − Failed the specific prompt constraint of having the horse on top of the astronaut.
- − Minor artifacting on the reins where they meet the astronaut's hands.
Verdict: Both Qwen Image and Vidu Q2 completely ignored the counter-intuitive instruction to place the horse on top of the astronaut, instead providing the standard 'astronaut riding horse' trope. Vidu Q2 is the preferred image because its surreal, cosmic interpretation much better aligns with the 'surreal' and 'cinematic' keywords compared to the more sterile look of Qwen Image.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
Qwen Image
- + Excellent photorealistic texture on the capybara's fur
- + Accurate depiction of a yellow taxi driver cap and dark jacket
- + Cleaner facial rendering for the human passenger
- − The paws on the steering wheel look more like primate hands than capybara paws
- − The passenger is visually very close to the driver due to the framing
Vidu Q2
- + Better composition showing more of the car interior and the street environment
- + Paws on the steering wheel look more anatomically appropriate for a rodent
- + Captures a wider range of cinematic lighting from the city
- − The passenger's face is slightly distorted and less detailed
- − The capybara's head looks somewhat pasted onto the body
- − Reflection artifacts on the bottom left corner
Verdict: Both models followed the prompt well, but Qwen Image (Model A) produced a more photorealistic image with better textures and a more convincing passenger. While Vidu Q2 (Model B) has a superior composition that emphasizes the space between the front and back seats, the overall technical execution and lighting in Qwen Image are more polished.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
Qwen Image
- + Perfect text rendering of the event details at the bottom
- + Excellent layout with a spooky, cinematic atmosphere
- + Successfully includes every requested element including the thorn border, scroll, and twisted trees
- − One minor typo in the large gothic title text ('Halle Party Invitation')
- − The bats are a bit simplistic in design
Vidu Q2
- + Strong gothic aesthetic for the borders and twisted trees
- + Vibrant jack-o-lantern rendering with good lighting effects
- − Significant spelling errors in both titles and scroll text ('Intovztion', 'invieed', 'fiigts')
- − Incorrect event details, displaying the wrong year and time formatting ('30.70.2025', 'Tmm')
- − The composition feels slightly more cluttered compared to the other model
Verdict: Qwen Image is the clear winner as it provides highly accurate text rendering for the event details and follows the atmospheric layout requirements perfectly. Vidu Q2 fails significantly on prompt adherence regarding specific text strings, including typos in the main title and incorrect dates/times.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
Qwen Image
- + Excellent typography with clean, bold text rendering.
- + Highly effective isometric perspective and diorama composition.
- + Soft, pleasing 3D cartoon textures that perfectly match the requested aesthetic.
- − The 3D flag on the diorama base is a bit large compared to the requested 'small icon' next to text.
Vidu Q2
- + Realistic PBR-style materiality on the seafood and plate surfaces.
- + Good adherence to the request for dark bold text and central alignment.
- + Detailed modeling of disparate sushi types.
- − Composition feels slightly cluttered compared to the 'minimal' request.
- − The diorama base has odd artifacts/lumps on the corners.
- − The flag icon is awkwardly attached to the letter 'N' like a flagpole.
Verdict: Qwen Image is the superior output as it perfectly captures the 45° isometric diorama aesthetic and provides much cleaner typography. While Vidu Q2 has more realistic material definitions on the food itself, it suffers from minor artifacts on the base and a less cohesive overall design compared to the ultra-clean look of Qwen Image.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
Qwen Image
- + Excellent soft fur texture and realistic animal renderings.
- + Clean composition with clear focus on all four requested animals.
- + Beautiful lighting with coherent god rays and dew effects.
- − Static posing that feels more like a portrait than a 'tumbling' action scene.
Vidu Q2
- + Successfully captures the 'tumbling' and 'playfully chasing' action described in the prompt.
- + Lush, expansive meadow environment with a high density of flowers.
- − Anatomical issues including a dog with an extra-long body and a bunny with a cat-like tail.
- − Inconsistent animal counts, showing two puppies instead of one.
- − Visual artifacts and clipping in the grass and butterflies.
Verdict: Qwen Image delivers much higher visual quality and realistic fur textures while perfectly adhering to the count and species of animals requested. Although Vidu Q2 captures the sense of movement and 'tumbling' better, it suffers from significant anatomical errors and ignores the specific animal counts in the prompt.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
Qwen Image
- + Excellent adherence to color scheme
- + Clear cloche illustration with simple steam
- + Minimalist vector style
- − Overlapping text makes portions unreadable
Vidu Q2
- + Elegant banner design
- + Nice subtle texture on background
- − Significant text spelling errors
- − Multiple redundant text elements
- − Cloche handle is malformed
Verdict: Qwen Image delivers a cleaner minimalist aesthetic and correctly follows the prompt's structural requirements, though it suffers from poor letter spacing. Vidu Q2 fails significantly on typography and text placement, resulting in multiple misspellings and a cluttered layout.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
Qwen Image
- + Stronger adherence to the navy-heavy NASA color palette.
- + Good use of flat vector graphics and icon silhouettes.
- + Text rendering is surprisingly legible for names like Armstrong and Aldrin.
- − Included instructions in the text output (Stop at landing).
- − Steps are jumbled and disorganized numerically (e.g., step 2 is next to step 1).
- − Nonsensical text integration for labels like 'Sarth Orbit Wicon)'.
Vidu Q2
- + Layout is very clean and structured as a proper infographic grid.
- + Visual style is consistent with high-quality vector illustrations.
- + The sequence of icons logically progresses from top-left to bottom-right.
- − Text consists almost entirely of gibberish strings.
- − Light gray palette lacks the bold contrast requested by the 'navy' color prompt.
- − Included 5 astronauts instead of the 3 specified in the supporting details.
Verdict: Qwen Image delivers a better thematic color palette and much more accurate text rendering, but the actual infographic flow is chaotic. Vidu Q2 creates a far superior layout that looks like a real poster, but fails on text legibility and color contrast. Qwen Image is the winner for actually following the specific steps and labels more closely, despite the numbering errors.
Explore each model
ShengShu Technology's text-to-image and reference-to-image model with support for character consistency and multi-reference image processing