The Max series of Tongyi Qwen’s image generation model excels across a wide range of generation tasks. Compared with the Plus series, it significantly reduces the “AI-like” feel in generated images, enhancing their realism. It delivers more lifelike material textures for human subjects, finer and more detailed natural textures, and more visually appealing text rendering.
Settled by community votes across 8 shared challenges, with an AI judge weighing in on each.
Qwen Image Max
#35 of 62 in Text-to-Image
Vidu Q2
#42 of 62 in Text-to-Image
Where the votes landed
Qwen Image Max
0%
win rate
Ties
0%
Vidu Q2
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
Qwen Image Max
- + Excellent photorealism and soft lighting
- + Accurate glass reflections and refractions
- + The blue sphere is perfectly centered and visually striking
- − The blue sphere appears to be floating mid-air inside the cube without physical support
Vidu Q2
- + Natural composition with the sphere resting on the bottom surface
- + High level of detail in the textures of the red book and wooden table
- + Accurate shadows cast by the plant and object structure
- − The glass refractive logic is slightly distorted on the right side of the cube
Verdict: Both models followed the prompt instructions perfectly, including all objects and lighting directions. Qwen Image Max produced a cleaner, more artistic image with superior lighting, while Vidu Q2 provided a more grounded and physically realistic placement of the sphere on the bottom of the cube.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
Qwen Image Max
- + Excellent adherence to 'battle-worn' with realistic skin textures, dirt, and deep scars.
- + Highly detailed engraving on the plate armor that matches the 'ornate' description perfectly.
- + Strong bokeh effect with visible sparks and a clear torchlight source that interacts naturally with the subject.
- − The hair braids appear more like dreadlocks, though they do feature the requested beads.
- − The composition is slightly tight, cutting off the top of the head.
Vidu Q2
- + Very clean and sharp rendering of the leather straps and underlayer textures.
- + Clear depiction of the braided hair and small beads as requested in the prompt.
- + Polished, cinematic lighting that highlights the contours of the armor well.
- − The character looks too pristine and 'pretty' to be truly battle-worn, despite a couple of light scratches.
- − The bokeh sparks are less prominent and the overall atmosphere feels more sterile than Image A.
Verdict: Qwen Image Max is the superior choice for this prompt as it captures the 'battle-worn' and 'highly detailed texture' aesthetic much more effectively, providing a gritty and realistic portrayal of a paladin. Vidu Q2 produces a high-quality cinematic image, but the character appears too clean and lacks the heavy engraving and textured wear requested in the prompt.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
Qwen Image Max
- + Excellent typography with a natural fiery glow effect.
- + High-quality, photorealistic textures on the patty and bun.
- + Dynamic composition with effective use of sparks and motion blur.
- − The burger is not fully 'exploded' as requested, with many components still tightly layered.
- − The starburst for the price is more of a light flare than a graphic starburst shape.
Vidu Q2
- + Successfully achieved the 'exploded' view with clear vertical separation of all ingredients.
- + Includes a distinct starburst graphic for the price as requested.
- + Vibrant color palette that emphasizes the fiery theme.
- − The pricing text contains a typo in the currency symbol (looks like a double-crossed E or yen variant).
- − The lighting on the burger feels a bit flat compared to the intensity of the background flames.
- − The top bun looks slightly distorted where it meets the sesame seeds.
Verdict: Qwen Image Max produced a more professional and visually polished advertisement with superior typography and realistic textures, though it failed to properly 'explode' the burger components. Vidu Q2 followed the layout instructions more literally by separating each ingredient, but was let down by a currency symbol error and slightly less realistic lighting. Qwen Image Max is the preferred choice for its higher aesthetic quality and cohesive design.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
Qwen Image Max
- + Excellent texture on the capybara's fur and the jacket
- + Highly realistic lighting inside the cabin from the street lights
- + Creative decision to include a cracked ceiling for added realism of an old taxi
- − The hands on the steering wheel look more like primate hands than capybara paws
- − The passenger is looking away from her phone rather than at it as requested
Vidu Q2
- + Perfectly captures the passenger looking at her phone with a bored expression
- + Authentic 'driver cap' with a badge that fits the theme well
- + Clearer composition showing both the driver and passenger in the same frame
- − The capybara's paws are positioned awkwardly on the steering wheel
- − The perspective of the car interior is slightly warped near the door
Verdict: Both models followed the prompt well, but Vidu Q2 is the winner because it accurately captured the passenger's interaction with her phone and her bored expression, which was central to the prompt's narrative. While Qwen Image Max has superior fur texture and more atmospheric lighting, it failed on the specific detail of the passenger looking at her phone.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
Qwen Image Max
- + Excellent typography with perfect spelling of all requested text.
- + High-quality composition with balanced elements and a clear central jack-o-lantern.
- + Strong aesthetic coherence, effectively blending the vintage parchment and dark sky.
- − The border is a bit repetitive in its thorn pattern.
Vidu Q2
- + Texture on the parchment feels very organic and aged.
- + Border elements feel more dynamic and integrated with the composition.
- − Significant spelling errors in both the title and specific event details.
- − Incorrect information displayed, such as the year being 2025 instead of 2026.
- − Lower resolution and more painterly artifacts compared to the clean finish of Model A.
Verdict: Qwen Image Max is the clear winner as it followed every instruction, including specific text strings, perfectly. Vidu Q2 failed significantly on text rendering, with multiple typos and incorrect event details, which makes it unusable for an invitation despite a nice vintage aesthetic.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
Qwen Image Max
- + Excellent PBR material rendering with realistic light refraction on the glass base
- + High-quality text rendering with a clean, professional aesthetic
- + Precise 45-degree isometric composition that feels solid and balanced
- − The 'JAPAN' text is slightly cut off at the top edge
- − The glass base is a bit large compared to the plate
Vidu Q2
- + Successfully incorporates the flag icon as a physical prop in the scene
- + Includes a warm, woody diorama base that fits the diorama theme well
- + Good color contrast and bright, cheerful lighting
- − The sushi shapes are slightly distorted and less realistic than requested
- − Text is somewhat plain and the layout is less sophisticated
- − A few minor artifact blobs on the wooden base
Verdict: Qwen Image Max produces a much more professional and high-fidelity output with superior material rendering, specifically the glass base and the textures of the salmon. While Vidu Q2 followed the instruction for a flag icon in a creative way, the overall image quality and text typography in Qwen Image Max are significantly cleaner and better aligned with the 'ultra-clean' request.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
Qwen Image Max
- + Excellent depiction of god rays and sunrise lighting
- + Very high detail in the fur texture and flower petals
- + The golden retriever puppies show realistic anatomical details and soft expression
- − Completely missed the baby bunny requirement
- − The fox kit has a slightly more mature, adult-like face shape
Vidu Q2
- + Includes all four requested animals: puppy, kitten, bunny, and fox kit
- + Captures the 'playfully chasing' and 'tumbling' energy better through more dynamic poses
- + Lovely sparkly dew effects and warm color palette
- − The bunny has cat-like paw structures and an unusual tail
- − Technical artifacting on the butterflies
- − The fox has a slightly distorted front leg
Verdict: While Qwen Image Max has higher technical fidelity and more realistic lighting, it failed to include the baby bunny. Vidu Q2 succeeded in including all requested subjects and captured the playful spirit of the prompt more effectively, despite some minor anatomical and artifact issues. Vidu Q2 is preferred for overall prompt adherence.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
Qwen Image Max
- + Excellent typography with perfect spelling of 'Caffè Florian' and 'Est. 1720'.
- + Clean vector illustration style that fits the 'minimalist' and 'logo' prompt requirements.
- + Good use of warm brown and cream tones with a textured heritage paper background.
- − The steam effect is a bit stylized/swirly rather than a subtle vapor.
Vidu Q2
- + Matches the requested cream and light brown color palette.
- + Includes the requested cloche and steam elements.
- − Severe spelling errors in 'Caffe Farmiin' and 'Esttt'.
- − Redundant and garbled text appearing at the bottom of the logo.
- − The cloche handle is poorly formed and looks detached.
Verdict: Qwen Image Max successfully followed all prompt instructions, delivering a professional-grade logo with perfect text rendering and a cohesive vintage aesthetic. Vidu Q2 failed significantly on the text elements, producing garbled words and repetitive formatting that rendered the logo unusable.
Explore each model
ShengShu Technology's text-to-image and reference-to-image model with support for character consistency and multi-reference image processing