Simultaneous audio generation
Simultaneously generates ambient sounds and background music tailored to video scenes and movements, providing highly polished video assets.
A video generation model capable of rapidly generating high-quality video and audio. It can create up to 8 seconds of 1080p video from text or multiple reference images.
Each model has its own Playground, pricing, and field docs. Filter by capability to open the one you need.

High-speed generation of cinematic-level quality. Google's state-of-the-art video model that instantly turns ideas into reality.

Google’s cutting-edge high-speed video generation model that instantly turns ideas into visuals.

Google's most advanced video generation, at an incredible speed. A high-speed model that turns ideas into shape instantly.
Model id and a short note on what each capability is for.
| Model | ID | Description |
|---|---|---|
| Veo 3.1 Fast(text to video) | google-video-veo-3.1-fast-text-to-video | High-speed generation of cinematic-level quality. Google's state-of-the-art video model that instantly turns ideas into reality. |
| Veo 3.1 Fast(refs image to video) | google-video-veo-3.1-fast-refs-image-to-video | Google’s cutting-edge high-speed video generation model that instantly turns ideas into visuals. |
| Veo 3.1 Fast(image to video) | google-video-veo-3.1-fast-image-to-video | Google's most advanced video generation, at an incredible speed. A high-speed model that turns ideas into shape instantly. |
A suite of features that combines speedy video export with audio integration to accelerate production workflows.
A powerful mage stands atop the shattered stone altar, raising his staff and unleashing an overwhelming burst of flame encased in Serpentine Blaze. This cinematic scene depicts a dramatic arc shot and a close-up shot of the moment of determination.
Prompt
A cinematic scene. Standing atop a shattered stone altar stands a powerful sorcerer. Above him, crimson clouds swirl ominously and violently. The camera rises from behind him, illuminating the ominous glowing runes at hi…
Generated with Veo 3.1 Fast
Generated with Gemini Omni Flash
Generated with Veo 3.1
On the rooftop of the futuristic city, a cyber girl passionately performs street dance in the neon rain, all captured by a dynamic camera.
Prompt
On the rooftop of the futuristic city, a cyber girl in a glowing jacket is dancing street dance. Neon rainwater reflects on the ground, with high-speed passing hover cars and giant holographic advertisements in the backg…
Generated with Veo 3.1 Fast
Generated with Gemini Omni Flash
Generated with Veo 3.1
An extremely realistic white wolf races across the mountain peak on a snowy night. The moonlight and the dancing snow dynamically chase after it.
Prompt
An extremely realistic wild white wolf is galloping through the snowy mountains at night. Its fur flutters in the wind, snow particles fly up, and the moonlight illuminates the lines of its muscles and the white breath i…
Generated with Veo 3.1 Fast
Generated with Gemini Omni Flash
Generated with Veo 3.1
Typical jobs this series is used for.
Quickly generate videos with audio from text, streamlining content creation and validation for social media.
Based on product images or key visuals, you can generate commercially usable promotional videos with motion and sound effects.
Combine up to 3 reference images to maintain visual consistency and quickly build storyboards and visual prototypes for world-building.
Input illustrations or photographic graphics and transform them into videos with smooth camera work and natural motion.
Check production limits and starting credit cost before you pick another model for the same job.
| Model | Max clip duration (single run) | Multimodal reference inputs | Max resolution | Native audio | Sousaku starting price |
|---|---|---|---|---|---|
| Veo 3.1 Fast | Up to 8 seconds | Up to 3 images | Up to 1080P | ✅ | 6 Credits/s |
| Gemini Omni Flash | Up to 10 seconds | Up to 10 videos | Up to 720P | ✅ | 4 Credits/s |
| Veo 3.1 | Up to 8 seconds | Up to 3 images | Up to 1080P | ✅ | 6 Credits/s |
From sign-up to your first successful call. The same steps work for every model.
Sign up and verify so you can issue an API key and use Playground.
Open a card above to see that model's details, pricing, and field docs.
Confirm inputs and parameters in the form before you automate.
The field docs on the detail page list required parameters. Use the same names in create_task.
Call the generate API with your key, poll until the task succeeds, then download the result.
Track credit use on the model's pricing page and in the dashboard.
Other series on Sousaku. One API key, one task flow.