Video generation from multiple images
Combines up to 3 reference images to maintain the subject and composition, reproducing natural camera work and movement according to your intent.
A video generation model with an excellent balance between image quality and generation speed. It can efficiently create 1080P videos with audio of up to 16 seconds from up to 3 reference images.
Each model has its own Playground, pricing, and field docs. Filter by capability to open the one you need.
Model id and a short note on what each capability is for.
| Model | ID | Description |
|---|---|---|
| Vidu Q3(refs image to video) | vidu-video-viduq3-refs-image-to-video | Combining speed and high quality. A next-generation video generation model that breathes life into still images. |
Balancing speed and quality, it enables flexible video production supporting audio generation and multiple reference image inputs.
On the rooftop of the futuristic city, a cyber girl passionately performs street dance in the neon rain, all captured by a dynamic camera.
Prompt
On the rooftop of the futuristic city, a cyber girl in a glowing jacket is dancing street dance. Neon rainwater reflects on the ground, with high-speed passing hover cars and giant holographic advertisements in the backg…
Generated with Vidu Q3 Pro
Generated with Vidu Q2
Generated with Kling O1
An extremely realistic white wolf races across the mountain peak on a snowy night. The moonlight and the dancing snow dynamically chase after it.
Prompt
An extremely realistic wild white wolf is galloping through the snowy mountains at night. Its fur flutters in the wind, snow particles fly up, and the moonlight illuminates the lines of its muscles and the white breath i…
Generated with Vidu Q3 Pro
Generated with Vidu Q2
Generated with Kling O1
A powerful mage stands atop the shattered stone altar, raising his staff and unleashing an overwhelming burst of flame encased in Serpentine Blaze. This cinematic scene depicts a dramatic arc shot and a close-up shot of the moment of determination.
Prompt
A cinematic scene. Standing atop a shattered stone altar stands a powerful sorcerer. Above him, crimson clouds swirl ominously and violently. The camera rises from behind him, illuminating the ominous glowing runes at hi…
Generated with Vidu Q3 Pro
Generated with Vidu Q2
Generated with Kling O1
Typical jobs this series is used for.
Quickly generate smooth videos with audio from reference images and text, allowing you to efficiently create video content for social media posts.
Utilizes up to three key visuals to construct scene progressions as intended, supporting the visualization of planning and concept validation.
Leveraging up to 1080p resolution and commercial-grade quality, you can create videos in a wide variety of styles, including anime and photorealistic.
Check production limits and starting credit cost before you pick another model for the same job.
| Model | Max clip duration (single run) | Multimodal reference inputs | Max resolution | Native audio | Sousaku starting price |
|---|---|---|---|---|---|
| Vidu Q3 | Up to 16 seconds | Up to 3 images | Up to 1080P | ✅ | 2 Credits/s |
| Vidu Q3 Pro | Up to 16 seconds | ❌ | Up to 1080P | ✅ | 2 Credits/s |
| Vidu Q2 Pro | Up to 10 seconds | Up to 3 images | Up to 1080P | ❌ | 2 Credits/s |
From sign-up to your first successful call. The same steps work for every model.
Sign up and verify so you can issue an API key and use Playground.
Open a card above to see that model's details, pricing, and field docs.
Confirm inputs and parameters in the form before you automate.
The field docs on the detail page list required parameters. Use the same names in create_task.
Call the generate API with your key, poll until the task succeeds, then download the result.
Track credit use on the model's pricing page and in the dashboard.
Other series on Sousaku. One API key, one task flow.