Simultaneous generation of synchronized audio
Since audio matched to the video's movements and scenes can be output simultaneously, it saves the hassle of adding sound separately.
A state-of-the-art video generation model supporting high-precision visual expression and audio synchronization. It can generate high-quality 1080P videos of up to 15 seconds from text, images, and existing videos.
Each model has its own Playground, pricing, and field docs. Filter by capability to open the one you need.

WAN
The top-tier model that generates movie-quality video and synchronized audio from video to video.

WAN
A next-generation video generation model with movie quality that supports multi-cuts up to 15 seconds and realistic audio synchronization

WAN
A top-tier video generation model that creates cinematic-quality video from images, with lip-sync capabilities.

WAN
The latest video model that creates cinematic-grade videos from images, with audio synchronization and outstanding consistency
Model id and a short note on what each capability is for.
| Model | ID | Description |
|---|---|---|
| WAN Video 2.6(video to video) | wan-video-2.6-video-to-video | The top-tier model that generates movie-quality video and synchronized audio from video to video. |
| WAN Video 2.6(text to video) | wan-video-2.6-text-to-video | A next-generation video generation model with movie quality that supports multi-cuts up to 15 seconds and realistic audio synchronization |
| WAN Video 2.6(refs image to video) | wan-video-2.6-refs-image-to-video | A top-tier video generation model that creates cinematic-quality video from images, with lip-sync capabilities. |
| WAN Video 2.6(image to video) | wan-video-2.6-image-to-video | The latest video model that creates cinematic-grade videos from images, with audio synchronization and outstanding consistency |
Simultaneously generates high-quality video and synchronized audio, supporting professional video production with a variety of input formats.
The nostalgic JRPG style of the SFC era. Dot-art characters roam the vibrant fantasy streets of the city.
Prompt
A classic 16-bit JRPG pixel art style. A warm and bright plaza in the royal city at dusk. A party of three - a swordsman, a magician, and a priest - walk from the foreground to the background of the screen. Other NPCs (v…
Generated with WAN Video 2.6
Generated with WAN Video 3.0 Prime
Generated with WAN Video 3.0
On the rooftop of the futuristic city, a cyber girl passionately performs street dance in the neon rain, all captured by a dynamic camera.
Prompt
On the rooftop of the futuristic city, a cyber girl in a glowing jacket is dancing street dance. Neon rainwater reflects on the ground, with high-speed passing hover cars and giant holographic advertisements in the backg…
Generated with WAN Video 2.6
Generated with WAN Video 3.0 Prime
Generated with WAN Video 3.0
A powerful mage stands atop the shattered stone altar, raising his staff and unleashing an overwhelming burst of flame encased in Serpentine Blaze. This cinematic scene depicts a dramatic arc shot and a close-up shot of the moment of determination.
Prompt
A cinematic scene. Standing atop a shattered stone altar stands a powerful sorcerer. Above him, crimson clouds swirl ominously and violently. The camera rises from behind him, illuminating the ominous glowing runes at hi…
Generated with WAN Video 2.6
Generated with WAN Video 3.0 Prime
Generated with WAN Video 3.0
Typical jobs this series is used for.
Directly generate 1080p videos with audio from text or product images, streamlining the production of social media ads and promotional videos.
Using the video-to-video conversion feature and up to three reference assets, you can restyle your video into diverse artistic styles while preserving the original motion.
By using up to 3 reference images, you can create story videos that maintain consistency in characters and worldview.
Recreating smooth camera work and natural physical behavior for up to 15 seconds, it is useful for creating storyboards and previz in video production.
Check production limits and starting credit cost before you pick another model for the same job.
| Model | Max clip duration (single run) | Multimodal reference inputs | Max resolution | Native audio | Sousaku starting price |
|---|---|---|---|---|---|
| WAN Video 2.6 | Up to 15 seconds | Up to 3 videos | Up to 1080P | ✅ | 4 Credits/s |
| WAN Video 3.0 Prime | Up to 30 seconds | Up to 10 assets (audio and videos) | Up to 1080P | ✅ | 6 Credits/s |
| WAN Video 3.0 | Up to 30 seconds | Up to 10 assets (audio and videos) | Up to 1080P | ✅ | 6 Credits/s |
From sign-up to your first successful call. The same steps work for every model.
Sign up and verify so you can issue an API key and use Playground.
Open a card above to see that model's details, pricing, and field docs.
Confirm inputs and parameters in the form before you automate.
The field docs on the detail page list required parameters. Use the same names in create_task.
Call the generate API with your key, poll until the task succeeds, then download the result.
Track credit use on the model's pricing page and in the dashboard.
Other series on Sousaku. One API key, one task flow.