Audio synchronization and simultaneous video generation
Simultaneously generates audio matching the motion of the video, allowing one-stop production of immersive videos with sound.
A next-generation AI video model capable of generating high-quality videos with audio from text or images. Achieving cinematic camera work and high consistency, it supports professional video production.
Each model has its own Playground, pricing, and field docs. Filter by capability to open the one you need.

WAN
Next-generation AI video generation model with cinematic quality and advanced control capabilities

WAN
A next-generation integrated AI video generation model that achieves cinematic-level image quality and high consistency

WAN
A next-generation AI model that creates cinematic dynamic videos from still images.
Model id and a short note on what each capability is for.
| Model | ID | Description |
|---|---|---|
| WAN Video 3.0(text to video) | wan-video-3.0-text-to-video | Next-generation AI video generation model with cinematic quality and advanced control capabilities |
| WAN Video 3.0(refs image to video) | wan-video-3.0-refs-image-to-video | A next-generation integrated AI video generation model that achieves cinematic-level image quality and high consistency |
| WAN Video 3.0(image to video) | wan-video-3.0-image-to-video | A next-generation AI model that creates cinematic dynamic videos from still images. |
Equipped with advanced features that enable diverse visual expressions, such as video generation from text or multiple images and audio synchronization.
A powerful mage stands atop the shattered stone altar, raising his staff and unleashing an overwhelming burst of flame encased in Serpentine Blaze. This cinematic scene depicts a dramatic arc shot and a close-up shot of the moment of determination.
Prompt
A cinematic scene. Standing atop a shattered stone altar stands a powerful sorcerer. Above him, crimson clouds swirl ominously and violently. The camera rises from behind him, illuminating the ominous glowing runes at hi…
Generated with WAN Video 3.0
Generated with WAN Video 3.0 Prime
Generated with WAN Video 2.7
On the rooftop of the futuristic city, a cyber girl passionately performs street dance in the neon rain, all captured by a dynamic camera.
Prompt
On the rooftop of the futuristic city, a cyber girl in a glowing jacket is dancing street dance. Neon rainwater reflects on the ground, with high-speed passing hover cars and giant holographic advertisements in the backg…
Generated with WAN Video 3.0
Generated with WAN Video 3.0 Prime
Generated with WAN Video 2.7
The nostalgic JRPG style of the SFC era. Dot-art characters roam the vibrant fantasy streets of the city.
Prompt
A classic 16-bit JRPG pixel art style. A warm and bright plaza in the royal city at dusk. A party of three - a swordsman, a magician, and a priest - walk from the foreground to the background of the screen. Other NPCs (v…
Generated with WAN Video 3.0
Generated with WAN Video 3.0 Prime
Generated with WAN Video 2.7
Typical jobs this series is used for.
Generates cinematic videos of up to 30 seconds from text prompts, reflecting movie-like camera work and lighting.
You can create videos with smooth motion using up to 10 reference images, while maintaining consistent characters and worldviews.
Generate high-quality videos with audio from illustrations or photos, allowing you to quickly create social media videos and ad creatives.
Outputs commercial-ready 1080P high-resolution video, streamlining the production of brand advertisements and product promotional videos.
Check production limits and starting credit cost before you pick another model for the same job.
| Model | Max clip duration (single run) | Multimodal reference inputs | Max resolution | Native audio | Sousaku starting price |
|---|---|---|---|---|---|
| WAN Video 3.0 | Up to 30 seconds | Up to 10 assets (audio and videos) | Up to 1080P | ✅ | 6 Credits/s |
| WAN Video 3.0 Prime | Up to 30 seconds | Up to 10 assets (audio and videos) | Up to 1080P | ✅ | 6 Credits/s |
| WAN Video 2.7 | Up to 15 seconds | Up to 5 images | Up to 1080P | ✅ | 4 Credits/s |
From sign-up to your first successful call. The same steps work for every model.
Sign up and verify so you can issue an API key and use Playground.
Open a card above to see that model's details, pricing, and field docs.
Confirm inputs and parameters in the form before you automate.
The field docs on the detail page list required parameters. Use the same names in create_task.
Call the generate API with your key, poll until the task succeeds, then download the result.
Track credit use on the model's pricing page and in the dashboard.
Other series on Sousaku. One API key, one task flow.