Multiple video inputs
Generate from text, images, reference subjects, start/end frames, and multiple keyframes.
Tool profile · Video
A multimodal creation platform spanning video, images, audio, templates, digital humans, and APIs.
01 / Overview
02 / Core features
Generate from text, images, reference subjects, start/end frames, and multiple keyframes.
Create or edit images and generate sound effects, speech, or cloned voices.
Use lip sync, subject replacement, video extension, and upscaling up to 8K.
Connect capabilities to production workflows through APIs, MCP, templates, and digital-human tools.
03 / Use cases
Guide connected shots with reference subjects or multiple keyframes.
Combine templates, digital humans, voice, and video generation for repeatable asset creation.
Bring generation into an application through APIs or MCP.
04 / Platforms and languages
Models, allowances, and pricing may differ between the web product and API platform.
05 / Pricing and quotas
Each generation task consumes a different number of credits; taxes depend on location.
07 / Common questions
No. The official function list also includes images, audio, lip sync, digital humans, replacement, extension, and upscaling.
Credits must be calculated by model, resolution, duration, and task type; the single-credit price is not enough.
Reference-to-video and multi-keyframe modes guide subjects, scenes, and frame relationships, but the selected model and material still need testing.
08 / Sources
Video, image, audio, template, and post-processing capabilities.
2026-08-08↗SourceVidu API official pricingCredit price and consumption rules by model, resolution, and duration.
2026-08-08↗SourceVidu API official introductionAPI platform, integration paths, and usage boundaries.
2026-08-08↗