Kling O1

In addition to high-precision video generation from text and images, it supports advanced editing using natural language. With excellent consistency and physical rendering, it strongly supports professional video production.

Text Friendly
Commercial Use

Input

Result

Idle

Ready to Generate

Configure your inputs and click run to generate an image preview.

Kling O1 Model Introduction

Model Overview

Kling O1 is the next-generation, top-of-the-line multimodal video model in the Kling series, integrating video "generation" and "editing" workflows into a single platform. It adopts an advanced MVL (Multimodal Visual Language) architecture, enabling not only video generation from text, images, and subject references, but also seamless editing via specifying start and end frames, natural language-based operations, and style conversion—all in one stop.

Thanks to a significant improvement in physics simulation performance, it maintains a high level of dynamic camera work and character consistency. It can create extremely smooth and natural video with flawless movement even in complex scenes. It greatly expands the creative possibilities for creators and designers who pursue high quality, such as in movie previsualization, storyboard production, advertisements, short dramas, virtual shooting, and more, while revolutionizing the production process.

Pricing

ResolutionCredits Consumed
720p(credits/s)4
1080p(credits/s)8

Technical Specifications

ParameterSpecification
Core CapabilityText to video
Resolution720p,1080p
Aspect Ratio16:9,9:16,1:1
Duration5,10

Explore similar models