Head to head
Esc

Models · slot A

to navigate to pick

Stable Diffusion 3.5 Large Turbo Stability AI Vidu Q2 ShengShu Technology

Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.

Stable Diffusion 3.5 Large Turbo

10.7 arena score

#61 of 62 in Text-to-Image

Skill signature · Text-to-Image

Vidu Q2

19.8 arena score

#42 of 62 in Text-to-Image

Vote tally

Where the votes landed

Stable Diffusion 3.5 Large Turbo

0.0%

win rate

Ties

0.0%

Vidu Q2

100.0%

win rate

0.0% 0.0% ties 100.0%
Shared challenges 13

Challenge by challenge

The strongest take from each model on every shared challenge, with the AI judge's read.

Geometric Composition

Text-to-Image

“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”

Stable Diffusion 3.5 Large Turbo
Vidu Q2
0% wins 0% ties 100% wins

AI Judge Analysis

Stable Diffusion 3.5 Large Turbo

  • + Clean, modern aesthetic with sharp lines
  • + Accurately represents light coming from the left
  • Failed the spatial layout by putting the book inside the cube instead of on top
  • The plant is to the left of the cube rather than behind it

Vidu Q2

  • + Perfect adherence to all spatial instructions in the prompt
  • + Excellent realism in glass reflections and textures
  • + Natural-looking lighting and composition
  • The sphere's surface looks slightly matte compared to the highly reflective glass

Verdict: Vidu Q2 followed every aspect of the prompt's spatial instructions, correctly placing the red book on top of the cube and the plant behind it. Stable Diffusion 3.5 Large Turbo failed the basic spatial logic, placing the red book and the sphere together inside/underneath the cube's frame.

Candid Street Photography

Text-to-Image

“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”

Stable Diffusion 3.5 Large Turbo
Vidu Q2
0% wins 0% ties 100% wins

AI Judge Analysis

Stable Diffusion 3.5 Large Turbo

  • + Clean representation of the red bicycle
  • + Visible rain streaks
  • + Correct subjects and general layout
  • Anatomical errors with the man's hands and head placement
  • Lacks the requested motion blur from passing cars
  • Character and environment look like a 3D render rather than a photo

Vidu Q2

  • + Successfully captures motion blur on the passing vehicle
  • + High realism in skin texture and bike weathering
  • + Excellent lighting and wet pavement reflections
  • Poor framing crops into the subject's face
  • Several anatomical errors in the hands and overlapping bike parts
  • The man appears to have too many fingers on his right hand

Verdict: Vidu Q2 much better adheres to the specific photographic instructions, including the mention of motion blur and natural skin textures, whereas Stable Diffusion 3.5 Large Turbo looks more like a digital illustration. Although Vidu Q2 has poor framing and significant anatomical distortions in the hands, its overall atmospheric realism far exceeds the clean but artificial look of Stable Diffusion.

Fantasy Warrior

Text-to-Image

“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”

Stable Diffusion 3.5 Large Turbo
Vidu Q2

AI Judge Analysis

Stable Diffusion 3.5 Large Turbo

  • + Strong implementation of ornate engraved patterns on the armor.
  • + Distinct braids that reflect the hair description well.
  • + Vibrant warm lighting and bokeh.
  • The blood/dirt textures look more like digital paint than realistic grime.
  • The skin has a plastic, overly-smooth '3D render' appearance.
  • Fails to include the requested beads in the braids.

Vidu Q2

  • + Excellent skin texture with realistic scars, dirt, and lifelike eyes.
  • + Strict adherence to the 'beads in braids' and 'leather straps' prompt details.
  • + Superior photographic quality with believable torchlight reflections.
  • The composition is slightly wider than a 'close portrait'.
  • The engraving on the armor is somewhat less intricate compared to Model A.

Verdict: Vidu Q2 is the clear winner as it provides a much more lifelike and gritty interpretation of the 'battle-worn' prompt with realistic skin textures and scars. While Stable Diffusion 3.5 Large Turbo has beautiful armor engravings, its overall aesthetic is too smooth and artificial, and it missed the specific detail of beads in the hair.

Modern Clean Menu

Text-to-Image

“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”

Stable Diffusion 3.5 Large Turbo
Vidu Q2

AI Judge Analysis

Stable Diffusion 3.5 Large Turbo

  • + Features very clear, bold sans-serif headers.
  • + Food images are vibrant and professional.
  • + Maintains a clean, high-contrast aesthetic.
  • Layout feels more like a flat-lay photography set than a cohesive menu design.
  • Misspells 'Mains' as 'Mians'.
  • Includes highly abstract/unrealistic food items like the square pizza slice.

Vidu Q2

  • + Follows a realistic menu layout with clear columns and sections.
  • + Successful integration of food photography within the text sections.
  • + More creative use of vibrant color accents through graphic elements.
  • Text rendering is poor with many gibberish words and spelling errors (e.g., 'Apecizen').
  • Layout is somewhat cluttered with overlapping text and images.
  • Food photos lack variety, appearing repetitive across sections.

Verdict: Stable Diffusion 3.5 Large Turbo produces a sharper, more minimalist aesthetic but fails to create a functional menu layout, instead presenting a collection of separate items. Vidu Q2 creates a more realistic document design that better integrates the requested grid and sections, though its text generation is significantly weaker. Vidu Q2 is the preferred winner as it successfully interprets the 'menu design' aspect of the prompt, whereas Stable Diffusion 3.5 Large Turbo looks like a product collage.

Magic Burger Explosion: Fiery Photorealism Challenge

Text-to-Image

“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”

Stable Diffusion 3.5 Large Turbo
Vidu Q2

AI Judge Analysis

Stable Diffusion 3.5 Large Turbo

  • + Features a vivid, artistic flame and ember effect in the background.
  • + Good texture on the grill marks and melted cheese.
  • Failed to include any of the requested text elements.
  • The burger is not exploded or suspended; most layers are touching.
  • The bottom of the bun appears to be melting into black charcoal.

Vidu Q2

  • + Excellent adherence to all prompt instructions, including various text fields.
  • + Successful 'exploded' view with clear separation between all burger components.
  • + Text is rendered with the requested fiery, glowing effect.
  • The pricing symbol is a stylized mashup rather than a clean Euro (€) symbol.
  • The lettuce texture is slightly less realistic compared to the buns.

Verdict: Vidu Q2 is clearly superior as it followed all prompt instructions, including the complex request for specific text strings and an 'exploded' layout. Stable Diffusion 3.5 Large Turbo completely failed to generate any of the requested text and produced a standard stacked burger instead of the dynamic, suspended version requested.

Chalkboard Menu

Text-to-Image

“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”

Stable Diffusion 3.5 Large Turbo
Vidu Q2

AI Judge Analysis

Stable Diffusion 3.5 Large Turbo

  • + The visual style of the cafe interior is bright and professional.
  • + Clean composition with balanced spacing and greenery.
  • The text is largely illegible and fails to follow the specific prompt requirements for item names and dates.
  • The 'handwriting' looks more like a digital font than actual chalk.
  • Contains numerous spelling errors and nonsensical words like 'trulale' and 'risctto'.

Vidu Q2

  • + Strong adherence to the requested text, successfully rendering the specific date and most menu items.
  • + Excellent chalk texture and realistic handwriting style with natural variations.
  • + Captures the gritty, authentic feel of a chalkboard menu rather than a digital graphic.
  • Contains some minor spelling errors in the menu items (e.g., 'Musshoom', 'Octopd').
  • The bottom-most text becomes cluttered and descends into garbled characters.

Verdict: Vidu Q2 is the clear winner as it successfully rendered the specific text requested in the prompt, including the complex date and specific food items, while capturing a realistic chalk-on-board texture. Stable Diffusion 3.5 Large Turbo failed nearly all text-rendering aspects of the prompt, producing a clean but generic image with largely nonsensical words.

The Reversed Rodeo

Text-to-Image

“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”

Stable Diffusion 3.5 Large Turbo
Vidu Q2

AI Judge Analysis

Stable Diffusion 3.5 Large Turbo

  • + Features a clean, cinematic composition with a striking horizon line.
  • + The astronaut's suit is well-defined and realistic.
  • + Adheres to the specific horse-riding-astronaut prompt logic.
  • The horse's front legs have anatomical issues and strange hoof shapes.
  • The mane and tail textures are a bit stringy and less integrated.

Vidu Q2

  • + Excellent use of color and 'surreal' aesthetic with the celestial horse skin.
  • + Highly detailed background with vibrant nebulae and stars.
  • + The horse's anatomy and movement feel more dynamic and natural.
  • The astronaut's right leg is positioned strangely behind the horse.
  • Minor clipping issues with the reins and the astronaut's hands.

Verdict: Both models successfully interpreted the prompt, but Vidu Q2 delivered a much more creative and visually stunning image by incorporating the 'surreal' instruction into the horse's appearance. While Stable Diffusion 3.5 Large Turbo provides a grounded, cinematic look, Vidu Q2 feels more like a finished piece of digital art with superior lighting and atmosphere.

The Capybara Taxi Driver

Text-to-Image

“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”

Stable Diffusion 3.5 Large Turbo
Vidu Q2

AI Judge Analysis

Stable Diffusion 3.5 Large Turbo

  • + Excellent close-up detail on the capybara's fur and whiskers
  • + Captures the jacket and steering wheel placement well
  • The passenger is mostly out of frame and blurred, losing the 'bored phone use' prompt requirement
  • The taxi cap looks more like a standard baseball cap

Vidu Q2

  • + Complete adherence to every part of the prompt including the passenger looking at her phone
  • + Realistic professional taxi driver hat design
  • + Better composition showing both characters clearly inside the vehicle
  • The capybara's hands look a bit too much like human-monkey hybrids rather than paws
  • Slightly less 'photorealistic' fur texture compared to Model A

Verdict: Vidu Q2 is the clear winner as it successfully incorporated every element of the prompt, including the specific bored expression of the businesswoman on her phone in the back seat. Stable Diffusion 3.5 Large Turbo produced a high-quality close-up of the capybara, but failed to show the interaction between the driver and the passenger as requested.

The Halloween Invitation

Text-to-Image

“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”

Stable Diffusion 3.5 Large Turbo
Vidu Q2

AI Judge Analysis

Stable Diffusion 3.5 Large Turbo

  • + Clean vector-style illustration
  • + Good composition and framing with the thorny border
  • Failed to include most of the requested text
  • Missing the scroll banner element
  • Lacks the 'aged parchment' texture requested

Vidu Q2

  • + Excellent adherence to all prompt elements including the scroll banner and parchment texture
  • + High level of detail in the thorns and moody background
  • + Capture the 'vintage gothic' aesthetic perfectly
  • Significant spelling errors in the title and banner text
  • Date was rendered as 30.70.2025 instead of 30.10.2026

Verdict: Vidu Q2 is the clear winner despite spelling errors because it followed nearly every visual and textual prompt instruction, including complex elements like the scroll banner and specific event details for the bottom. Stable Diffusion 3.5 Large Turbo produced a generic graphic that ignored the majority of the specific text requirements and the parchment aesthetic.

Isometric Miniature Diorama Scenes

Text-to-Image

“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”

Stable Diffusion 3.5 Large Turbo
Vidu Q2

AI Judge Analysis

Stable Diffusion 3.5 Large Turbo

  • + Excellent 3D rendering with soft PBR-style textures
  • + Strong isometric perspective and dioramic composition
  • Misspelled words: 'SUSHI' appears as 'SIIHI' and text is not top-center
  • The flag icon is a generic red/white design rather than the Japanese flag

Vidu Q2

  • + Perfect text rendering for both 'JAPAN' and 'SUSHI'
  • + Accurate Japanese flag icon included in the composition
  • + Correctly positioned text at the top-center as requested
  • The diorama base has odd small artifacts or blobs on the corners
  • Lighting is slightly flatter compared to Model A's depth

Verdict: While Stable Diffusion 3.5 Large Turbo produces a more visually pleasing 3D model with better textures, it failed significantly on the text rendering and flag accuracy. Vidu Q2 followed every aspect of the complex prompt perfectly, including text placement, spelling, and iconography, making it the superior output for this specific challenge.

Adorable Baby Animals in Sunny Meadow

Text-to-Image

“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”

Stable Diffusion 3.5 Large Turbo
Vidu Q2

AI Judge Analysis

Stable Diffusion 3.5 Large Turbo

  • + Excellent light scattering and backlight effects on the fur
  • + Clear and detailed facial features for the animals shown
  • Failed to include the requested baby bunny
  • Missing the fox kit (only kitten and puppy shown)
  • Visual style feels more like digital art than hyper-photorealistic

Vidu Q2

  • + Perfect adherence to the prompt, including all four specific animals
  • + Captures the 'tumbling' and 'chasing' action much better
  • + Successfully renders god rays and the golden sunrise atmosphere
  • Some minor anatomical blurring on the running legs
  • Slightly less sharpness in the distant background foliage

Verdict: Vidu Q2 is the clear winner as it followed the complex prompt precisely, including the puppy, kitten, bunny, and fox kit, whereas Stable Diffusion 3.5 Large Turbo omitted two of the four animals. Vidu Q2 also captured a much more dynamic scene that better reflects the 'chasing' and 'tumbling' keywords with realistic photographic lighting.

Vintage Cafe Logo

Text-to-Image

“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”

Stable Diffusion 3.5 Large Turbo
Vidu Q2

AI Judge Analysis

Stable Diffusion 3.5 Large Turbo

  • + Excellent typography rendering with only a minor spelling artifact on the 'e'
  • + Strong vector illustration style with clean lines and high contrast
  • + Good interpretation of the cloche dome combined with a mug element
  • Included a slight misspelling of the name ('Caffeé' rather than 'Caffè')
  • The handle on the right makes it look more like a mug than a pure cloche

Vidu Q2

  • + Colors perfectly match the 'warm brown and cream tones' requested
  • + Captures the 'subtle texture' on the background effectively
  • Significant text failures and misspellings in multiple locations
  • Weak composition with repetitive text elements stacked vertically
  • Cloche shape is slightly asymmetrical and lacks the requested vintage crispness

Verdict: Stable Diffusion 3.5 Large Turbo produced a significantly more professional and usable logo with clear vector-style lines and mostly accurate spelling. Vidu Q2 failed significantly on text rendering, producing unintelligible words and redundant labels, while having a less coherent layout.

Apollo 11: Journey to Tranquility

Text-to-Image

“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”

Stable Diffusion 3.5 Large Turbo
Vidu Q2

AI Judge Analysis

Stable Diffusion 3.5 Large Turbo

  • + Strong aesthetic alignment with vintage NASA poster styles
  • + Balanced and sophisticated layout with a clear focal point
  • Text is largely illegible and fails to label the steps
  • Missing the specific sequential step-by-step icons requested

Vidu Q2

  • + Better adherence to the structure of the sequential steps
  • + Cleaner, modern vector iconography and clearer text rendering
  • The color palette is slightly too bright compared to the 'navy/muted' request
  • Layout feels more like a grid of stickers than a unified infographic poster

Verdict: Stable Diffusion 3.5 Large Turbo produces a more artistic and visually interesting poster, but it fails to follow the logical step-by-step instructions. Vidu Q2 follows the sequential prompt more accurately and utilizes a cleaner vector style, making it a more functional infographic despite the generic layout.

Next steps

Explore each model