Drag and drop media files from your computer, paste from clipboard (Ctrl/Cmd+V), or provide a URL. Accepted: .jpg, .jpeg, .png, .webp
Ready to Generate
Configure your inputs and click run to generate an image preview.
The Vidu Q2 Pro is the next-generation flagship model of the Vidu series, delivering overwhelmingly expressive character rendering and cinematic visual beauty. It sublimates images into dynamic, natural videos while maintaining the features of the reference image with extremely high accuracy. It has achieved remarkable advancements particularly in the expression of character performances, enabling detailed depictions that far surpass previous models, including subtle changes in facial expressions, natural blinks, eye movements, and precise lip-sync. The stability of camera work has also been greatly improved, allowing for seamless cinematic visual effects. It supports audio and commercial use, making it a powerful tool that meets all the demands of professional video production scenarios such as advertising creatives, previsualizations for films and animations, and social media content.
| Resolution | Credits Consumed |
|---|---|
| 480p(credits/s) | 2 |
| 720p(credits/s) | 2 |
| 1080p(credits/s) | 6 |
| Parameter | Specification |
|---|---|
| Core Capability | Reference image to video |
| Resolution | 480p,720p,1080p |
| Aspect Ratio | 16:9,9:16,1:1 |
| Duration | 1,2,3,4,5,6,7,8,9,10 |

BytePlus
An innovative multimodal generative model that creates cinematic videos up to 30 seconds long from reference images.

BytePlus
A lightweight video generation model that quickly generates high-quality videos from reference images.

WAN
Next-generation model that transforms still images and audio into movie-quality videos with overwhelming generation speed

Google's fast multimodal video generation and editing model, which highly integrates audio and video

WAN
A next-generation integrated AI video generation model that achieves cinematic-level image quality and high consistency

Google's top-tier video generation model, crafting cinematic-level visuals and sound.