Supports simultaneous audio generation
Because audio can be generated simultaneously with video, you can easily create immersive short videos and promotional clips.
A professional-grade model capable of generating high-quality videos with audio from text or images. It supports up to 1080p resolution and smooth video rendering of up to 16 seconds.
Each model has its own Playground, pricing, and field docs. Filter by capability to open the one you need.
Model id and a short note on what each capability is for.
| Model | ID | Description |
|---|---|---|
| Vidu Q3 Pro(text to video) | vidu-video-viduq3-pro-text-to-video | A professional-grade video generation model that pursues cinematic-level overwhelming visual beauty and camera work |
| Vidu Q3 Pro(image to video) | vidu-video-viduq3-pro-image-to-video | Movie-grade footage from a single image. A professional video generation model for creators. |
Supporting high-definition visual expression and audio generation, it assists in the creation of diverse video creatives.
On the rooftop of the futuristic city, a cyber girl passionately performs street dance in the neon rain, all captured by a dynamic camera.
Prompt
On the rooftop of the futuristic city, a cyber girl in a glowing jacket is dancing street dance. Neon rainwater reflects on the ground, with high-speed passing hover cars and giant holographic advertisements in the backg…
Generated with Vidu Q3 Pro
Generated with Vidu Q2
Generated with HappyHorse 1.1
An extremely realistic white wolf races across the mountain peak on a snowy night. The moonlight and the dancing snow dynamically chase after it.
Prompt
An extremely realistic wild white wolf is galloping through the snowy mountains at night. Its fur flutters in the wind, snow particles fly up, and the moonlight illuminates the lines of its muscles and the white breath i…
Generated with Vidu Q3 Pro
Generated with Vidu Q2
Generated with HappyHorse 1.1
A powerful mage stands atop the shattered stone altar, raising his staff and unleashing an overwhelming burst of flame encased in Serpentine Blaze. This cinematic scene depicts a dramatic arc shot and a close-up shot of the moment of determination.
Prompt
A cinematic scene. Standing atop a shattered stone altar stands a powerful sorcerer. Above him, crimson clouds swirl ominously and violently. The camera rises from behind him, illuminating the ominous glowing runes at hi…
Generated with Vidu Q3 Pro
Generated with Vidu Q2
Generated with HappyHorse 1.1
Typical jobs this series is used for.
Generates videos with audio from text prompts or still images, streamlining the creation of assets for advertising and web promotions.
With natural camera work and smooth motion of people and objects, you can create expressive, cinematic scenes.
By inputting key visuals or illustrations, it converts them into dynamic videos while preserving the worldview and composition of the original image.
Check production limits and starting credit cost before you pick another model for the same job.
| Model | Max clip duration (single run) | Multimodal reference inputs | Max resolution | Native audio | Sousaku starting price |
|---|---|---|---|---|---|
| Vidu Q3 Pro | Up to 16 seconds | ❌ | Up to 1080P | ✅ | 2 Credits/s |
| Vidu Q3 | Up to 16 seconds | Up to 3 images | Up to 1080P | ✅ | 2 Credits/s |
| Vidu Q2 Pro | Up to 10 seconds | Up to 3 images | Up to 1080P | ❌ | 2 Credits/s |
From sign-up to your first successful call. The same steps work for every model.
Sign up and verify so you can issue an API key and use Playground.
Open a card above to see that model's details, pricing, and field docs.
Confirm inputs and parameters in the form before you automate.
The field docs on the detail page list required parameters. Use the same names in create_task.
Call the generate API with your key, poll until the task succeeds, then download the result.
Track credit use on the model's pricing page and in the dashboard.
Other series on Sousaku. One API key, one task flow.