Tool profile · Video

Vidu

Generate controllable video from text, images, reference subjects, and keyframes

A multimodal creation platform spanning video, images, audio, templates, digital humans, and APIs.

Video generationReference videoKeyframesDigital humansAPI

01 / Overview

About

Vidu supports text-to-video, image-to-video, reference-to-video, start/end-frame, and multi-keyframe generation. The same platform and API also cover images, audio, lip sync, digital humans, subject replacement, video extension, and upscaling. It can be used in the web studio or integrated through APIs and MCP.

02 / Core features

Core features

01

Multiple video inputs

Generate from text, images, reference subjects, start/end frames, and multiple keyframes.

02

Images and audio

Create or edit images and generate sound effects, speech, or cloned voices.

03

Video post-processing

Use lip sync, subject replacement, video extension, and upscaling up to 8K.

04

APIs, templates, and digital humans

Connect capabilities to production workflows through APIs, MCP, templates, and digital-human tools.

03 / Use cases

Use cases

01

Character and scene consistency

Guide connected shots with reference subjects or multiple keyframes.

02

Advertising and short drama production

Combine templates, digital humans, voice, and video generation for repeatable asset creation.

03

Product integration

Bring generation into an application through APIs or MCP.

04 / Platforms and languages

Platforms and languages

Platforms and devices

WebAPIMCP

Languages

ChineseEnglishand multilingual prompt and voice workflows

Models, allowances, and pricing may differ between the web product and API platform.

05 / Pricing and quotas

Pricing and quotas

API credits

$0.005 each / Pay as needed

Each generation task consumes a different number of credits; taxes depend on location.

06 / Similar products

Similar products

可灵

Multimodal video and image generation platform

AI image and video generation platform

Runway

Professional creation, editing, and generation workflows

video

LibTV

Multiple video models in a canvas-based workflow

video

07 / Common questions

Frequently asked questions

01Is Vidu limited to video generation?

No. The official function list also includes images, audio, lip sync, digital humans, replacement, extension, and upscaling.

02How much does one API generation cost?

Credits must be calculated by model, resolution, duration, and task type; the single-credit price is not enough.

03Can it keep characters consistent?

Reference-to-video and multi-keyframe modes guide subjects, scenes, and frame relationships, but the selected model and material still need testing.