Gemini Omni Flash

A multimodal video model that can quickly generate videos with audio from text or images. It achieves high-quality visual representation of up to 10 seconds.

Models in this series

Each model has its own Playground, pricing, and field docs. Filter by capability to open the one you need.

The Gemini Omni Flash series at a glance

Model id and a short note on what each capability is for.

ModelIDDescription
Gemini Omni Flash(text to video)google-video-omni-flash-text-to-videoThe latest model that seamlessly combines text and audio to quickly generate consistent, high-quality videos.
Gemini Omni Flash(refs image to video)google-video-omni-flash-refs-image-to-videoGoogle's fast multimodal video generation and editing model, which highly integrates audio and video
Gemini Omni Flash(image to video)google-video-omni-flash-image-to-videoA multimodal model that generates and edits high-quality videos from images and audio at overwhelming speed.

Main Features and Functions

Integrates advanced multimodal processing and voice generation to support efficient video production.

Simultaneous audio generation

Simultaneous audio generation

Automatically generates scene-synchronized audio alongside video generation, creating immersive content in a single generation.

Support for multiple reference images

Support for multiple reference images

You can input up to 10 reference images, enabling video generation that maintains the subject's features and style.

Realistic physical behavior representation

Realistic physical behavior representation

Simulates natural gravity and object movement, outputting realistic and natural motion video.

Commercial use supported

Commercial use supported

It supports commercially usable tags, so you can use it with confidence for promotional and business content creation.

The white wolf racing through the moonlit snowy mountains.

An extremely realistic white wolf races across the mountain peak on a snowy night. The moonlight and the dancing snow dynamically chase after it.

Prompt

An extremely realistic wild white wolf is galloping through the snowy mountains at night. Its fur flutters in the wind, snow particles fly up, and the moonlight illuminates the lines of its muscles and the white breath i…

Generated with Gemini Omni Flash

Generated with Veo 3.1

Generated with Veo 3.1 Fast

Coconuts Moment

A moment of high-speed photography brimming with dynamism. The milky-white milk dynamically splashes out from a coconut splitting in the air, and the sunlight of the tropics shines through the water droplets. It conveys a sense of luxury and sophistication.

Prompt

A high-speed cinematic shot of a coconut splitting in the air. The white, creamy milk splashes dynamically based on fluid physics. Sunlight penetrating the water droplets, palm leaves swaying in the breeze in the backgro…

Generated with Gemini Omni Flash

Generated with Veo 3.1

Generated with Veo 3.1 Fast

The hustle and bustle of the royal capital (in the style of a JRPG)

The nostalgic JRPG style of the SFC era. Dot-art characters roam the vibrant fantasy streets of the city.

Prompt

A classic 16-bit JRPG pixel art style. A warm and bright plaza in the royal city at dusk. A party of three - a swordsman, a magician, and a priest - walk from the foreground to the background of the screen. Other NPCs (v…

Generated with Gemini Omni Flash

Generated with Veo 3.1

Generated with Veo 3.1 Fast

Featured use cases

Typical jobs this series is used for.

Short-form video content creation

Batch-generate videos with synchronized audio from text or image prompts, streamlining post creation for social media and short-form videos.

Video generation from multiple images

You can create story-driven videos that maintain character and worldview consistency using up to 10 reference images.

Ad creative verification

Generates commercially usable video and audio at high speed, making it ideal for rapid prototyping and validation of diverse ad variations.

Compare with other models

Check production limits and starting credit cost before you pick another model for the same job.

ModelMax clip duration (single run)Multimodal reference inputsMax resolutionNative audioSousaku starting price
Gemini Omni FlashUp to 10 secondsUp to 10 videosUp to 720P4 Credits/s
Veo 3.1Up to 8 secondsUp to 3 imagesUp to 1080P6 Credits/s
Veo 3.1 LiteUp to 8 secondsUp to 1080P2 Credits/s

Quick start guide

From sign-up to your first successful call. The same steps work for every model.

Step 1

Create your account

Sign up and verify so you can issue an API key and use Playground.

Step 2

Choose a model in this series

Open a card above to see that model's details, pricing, and field docs.

Step 3

Try it in Playground

Confirm inputs and parameters in the form before you automate.

Step 4

Copy the model id and fields

The field docs on the detail page list required parameters. Use the same names in create_task.

Step 5

Integrate and poll

Call the generate API with your key, poll until the task succeeds, then download the result.

Step 6

Check usage

Track credit use on the model's pricing page and in the dashboard.

FAQ

Explore other series

Other series on Sousaku. One API key, one task flow.