Alibaba's multimodal generation model from the Wan AI suite, supporting text-to-video, image-to-video, reference-to-video with audio, and text-to-image, in both Chinese and English

What Wan 2.6 is good at

Rankings based on human votes in the Lumenfall Arena, where real users pick their favorite without knowing which model made which image. Excels at #1 Prompt Adherence (Text-to-Video), #1 Human Fidelity (Text-to-Video), #3 Physics & Realism (Text-to-Video), and #3 Motion Quality (Text-to-Video) , placing in the top 33% of all competing models.

Wan 2.6 Strengths

Wan 2.6 outperforms most competing models here

Prompt Adherence

Text-to-Video
Ranked #1 of 7 models · 60.7% win rate Leaderboard
#1
of 7 models · 60.7% wins

Human Fidelity

Text-to-Video
Ranked #1 of 7 models · 60.7% win rate Leaderboard
#1
of 7 models · 60.7% wins

Physics & Realism

Text-to-Video
Ranked #3 of 7 models · 60.7% win rate Leaderboard
#3
of 7 models · 60.7% wins
#1 is Seedance 2.0

Motion Quality

Text-to-Video
Ranked #3 of 7 models · 60.7% win rate Leaderboard
#3
of 7 models · 60.7% wins
#1 is Seedance 2.0