Simultaneous Generation of Voice and Ambient Sound
It can output realistic audio and ambient sounds simultaneously with video generation, allowing for the creation of highly complete video content on its own.
A state-of-the-art video generation model developed by Google. It can generate high-quality video with audio from text or images, making it ideal for creative production and promotional videos.
Each model has its own Playground, pricing, and field docs. Filter by capability to open the one you need.

Google DeepMind's state-of-the-art AI video generation model. It produces cinematic-level visuals and synchronized sound.

Google's top-tier video generation model, crafting cinematic-level visuals and sound.

Google’s top-tier AI video generation model. Turn still images into movie-quality videos.
Model id and a short note on what each capability is for.
| Model | ID | Description |
|---|---|---|
| Veo 3.1(text to video) | google-video-veo-3.1-text-to-video | Google DeepMind's state-of-the-art AI video generation model. It produces cinematic-level visuals and synchronized sound. |
| Veo 3.1(refs image to video) | google-video-veo-3.1-refs-image-to-video | Google's top-tier video generation model, crafting cinematic-level visuals and sound. |
| Veo 3.1(image to video) | google-video-veo-3.1-image-to-video | Google’s top-tier AI video generation model. Turn still images into movie-quality videos. |
Featuring audio generation and multi-image reference capabilities, it powerfully supports high-quality, consistent video production.
A powerful mage stands atop the shattered stone altar, raising his staff and unleashing an overwhelming burst of flame encased in Serpentine Blaze. This cinematic scene depicts a dramatic arc shot and a close-up shot of the moment of determination.
Prompt
A cinematic scene. Standing atop a shattered stone altar stands a powerful sorcerer. Above him, crimson clouds swirl ominously and violently. The camera rises from behind him, illuminating the ominous glowing runes at hi…
Generated with Veo 3.1
Generated with Gemini Omni Flash
Generated with Veo 3.1 Fast
A moment of high-speed photography brimming with dynamism. The milky-white milk dynamically splashes out from a coconut splitting in the air, and the sunlight of the tropics shines through the water droplets. It conveys a sense of luxury and sophistication.
Prompt
A high-speed cinematic shot of a coconut splitting in the air. The white, creamy milk splashes dynamically based on fluid physics. Sunlight penetrating the water droplets, palm leaves swaying in the breeze in the backgro…
Generated with Veo 3.1
Generated with Gemini Omni Flash
Generated with Veo 3.1 Fast
On the rooftop of the futuristic city, a cyber girl passionately performs street dance in the neon rain, all captured by a dynamic camera.
Prompt
On the rooftop of the futuristic city, a cyber girl in a glowing jacket is dancing street dance. Neon rainwater reflects on the ground, with high-speed passing hover cars and giant holographic advertisements in the backg…
Generated with Veo 3.1
Generated with Gemini Omni Flash
Generated with Veo 3.1 Fast
Typical jobs this series is used for.
Generate videos with audio from text or images, and quickly create engaging promotional videos for social media.
With a maximum resolution of up to 1080P and realistic texture rendering, you can efficiently produce high-quality video ads ready for commercial use.
You can create narrative scene videos while maintaining character and composition consistency using up to 3 reference images.
Leverage cinematic camera work and physics expressions to quickly build previsualizations for dynamically reviewing video production plots and storyboards.
Check production limits and starting credit cost before you pick another model for the same job.
| Model | Max clip duration (single run) | Multimodal reference inputs | Max resolution | Native audio | Sousaku starting price |
|---|---|---|---|---|---|
| Veo 3.1 | Up to 8 seconds | Up to 3 images | Up to 1080P | ✅ | 6 Credits/s |
| Gemini Omni Flash | Up to 10 seconds | Up to 10 videos | Up to 720P | ✅ | 4 Credits/s |
| Veo 3.1 Lite | Up to 8 seconds | ❌ | Up to 1080P | ✅ | 2 Credits/s |
From sign-up to your first successful call. The same steps work for every model.
Sign up and verify so you can issue an API key and use Playground.
Open a card above to see that model's details, pricing, and field docs.
Confirm inputs and parameters in the form before you automate.
The field docs on the detail page list required parameters. Use the same names in create_task.
Call the generate API with your key, poll until the task succeeds, then download the result.
Track credit use on the model's pricing page and in the dashboard.
Other series on Sousaku. One API key, one task flow.