Diverse generation modes and fast inference
Supports generation from text, images, and multiple reference materials, delivering high-speed video output while maintaining high quality.
A high-quality video generation model capable of rapidly generating videos with audio from text, images, and up to 10 reference assets, supporting up to 1080P and 30 seconds.
Each model has its own Playground, pricing, and field docs. Filter by capability to open the one you need.

WAN
Next-generation multimodal video generation model that creates cinematic videos up to 30 seconds at amazing speed

WAN
Next-generation model that transforms still images and audio into movie-quality videos with overwhelming generation speed

WAN
From still images to dynamic and lively videos at high speed! Next-generation video generation model compatible with up to 30 seconds
Model id and a short note on what each capability is for.
| Model | ID | Description |
|---|---|---|
| WAN Video 3.0 Prime(text to video) | wan-video-3.0-prime-text-to-video | Next-generation multimodal video generation model that creates cinematic videos up to 30 seconds at amazing speed |
| WAN Video 3.0 Prime(refs image to video) | wan-video-3.0-prime-refs-image-to-video | Next-generation model that transforms still images and audio into movie-quality videos with overwhelming generation speed |
| WAN Video 3.0 Prime(image to video) | wan-video-3.0-prime-image-to-video | From still images to dynamic and lively videos at high speed! Next-generation video generation model compatible with up to 30 seconds |
Supports fast inference and diverse input formats, seamlessly enabling high-quality video production with audio.
On the rooftop of the futuristic city, a cyber girl passionately performs street dance in the neon rain, all captured by a dynamic camera.
Prompt
On the rooftop of the futuristic city, a cyber girl in a glowing jacket is dancing street dance. Neon rainwater reflects on the ground, with high-speed passing hover cars and giant holographic advertisements in the backg…
Generated with WAN Video 3.0 Prime
Generated with WAN Video 3.0
Generated with WAN Video 2.7
A powerful mage stands atop the shattered stone altar, raising his staff and unleashing an overwhelming burst of flame encased in Serpentine Blaze. This cinematic scene depicts a dramatic arc shot and a close-up shot of the moment of determination.
Prompt
A cinematic scene. Standing atop a shattered stone altar stands a powerful sorcerer. Above him, crimson clouds swirl ominously and violently. The camera rises from behind him, illuminating the ominous glowing runes at hi…
Generated with WAN Video 3.0 Prime
Generated with WAN Video 3.0
Generated with WAN Video 2.7
The nostalgic JRPG style of the SFC era. Dot-art characters roam the vibrant fantasy streets of the city.
Prompt
A classic 16-bit JRPG pixel art style. A warm and bright plaza in the royal city at dusk. A party of three - a swordsman, a magician, and a priest - walk from the foreground to the background of the screen. Other NPCs (v…
Generated with WAN Video 3.0 Prime
Generated with WAN Video 3.0
Generated with WAN Video 2.7
Typical jobs this series is used for.
Specify camera work and lighting from text prompts to generate cinematic sequences with audio in one go.
Add dynamic movement and audio based on your own images while preserving the subject's features and composition.
You can create long-form videos while maintaining the consistency of characters' appearances and the worldview using up to 10 reference materials.
Enables single generation of commercially usable 1080p videos up to 30 seconds, streamlining the production of social media ads and promotional videos.
Check production limits and starting credit cost before you pick another model for the same job.
| Model | Max clip duration (single run) | Multimodal reference inputs | Max resolution | Native audio | Sousaku starting price |
|---|---|---|---|---|---|
| WAN Video 3.0 Prime | Up to 30 seconds | Up to 10 assets (audio and videos) | Up to 1080P | ✅ | 6 Credits/s |
| WAN Video 3.0 | Up to 30 seconds | Up to 10 assets (audio and videos) | Up to 1080P | ✅ | 6 Credits/s |
| WAN Video 2.7 | Up to 15 seconds | Up to 5 images | Up to 1080P | ✅ | 4 Credits/s |
From sign-up to your first successful call. The same steps work for every model.
Sign up and verify so you can issue an API key and use Playground.
Open a card above to see that model's details, pricing, and field docs.
Confirm inputs and parameters in the form before you automate.
The field docs on the detail page list required parameters. Use the same names in create_task.
Call the generate API with your key, poll until the task succeeds, then download the result.
Track credit use on the model's pricing page and in the dashboard.
Other series on Sousaku. One API key, one task flow.