Audio-to-video lipsync. Upload a 2–10 second focal video and a clean vocal track (≤5 MB). Kling aligns mouth shapes and facial muscles to the audio while preserving the original footage.
Approx. Price
$0.450 per video
Model Type
video-to-video
Settings
Generation controls available for this model.
Output Format
N/A
Default Duration
N/A
No configurable settings are exposed for this model yet.
Benchmarks
Benchmarks
No benchmark data is available yet for this model.
Examples
Loading examples…
Related video models
Compare Kling Lipsync A2V with similar models from the same provider or model family.
Kling Lipsync T2V
kling-lipsync-t2vText-to-video lipsync. Upload a 2–10 second focal video and provide a script. Kling synthesizes a matching voiceover and animates lips/micro-expressions to the dialogue.
Kling 3.0 Turbo Pro
kling-v3-turbo-proKling 3.0 Turbo Pro generates high quality 1080p video from text or a first-frame image with stronger prompt consistency, stable motion, multi-shot storyboards, and 3-15 second durations.
Kling 3.0 Turbo Standard
kling-v3-turbo-standardKling 3.0 Turbo Standard generates fast, cost-efficient 720p video from text or a first-frame image. Supports stable motion, multi-shot storyboard prompts, and 3-15 second durations.
Kling O3 4K
kling-o3-4kNative 4K Kling O3 model for text-to-video, image-to-video, and reference-to-video generation. Supports start/end frame control, up to 7 reference images, and 3-15 second durations.
Kling V3 4K
kling-v3-4kNative 4K Kling V3 generation for text-to-video and image-to-video. Supports first/last-frame image control, 3-15 second durations, 16:9/9:16/1:1 output, and optional native audio.
Kling 3.0 Pro Motion Control
kling-v30-pro-motion-controlTransfer motion from a reference video to animate a still character image with Kling 3.0 Pro. Requires an image and a motion clip, supports optional prompts, and can retain the original video audio.