LatentSync

latentsync

LatentSync

latentsync

State-of-the-art audio-to-video lip synchronization using latent diffusion. Upload a talking-head video (480p+) and target audio to generate perfectly synchronized lip movements while preserving identity, pose, and background.

Added Dec 6, 2025

Starting Price

From $0.150 per video

Final price depends on the selected settings and is shown before generation.

Model Type

video-to-video

Settings

Generation controls available for this model.

Output Format

N/A

Default Duration

N/A

No configurable settings are exposed for this model yet.

Benchmarks

No benchmark data is available yet for this model.

Examples

Loading examples…

Compare LatentSync with similar models from the same provider or model family.

BytePlus Video Enhancement Pro

bytedance-video-enhancement-pro

Apply stronger BytePlus video restoration and enhancement for footage with real people. Pro improves skin texture and fine detail while supporting upscaling to 2K or 4K, frame interpolation up to 120 FPS, denoising, deblocking, deblurring, sharpening, and color and contrast enhancement.

BytePlus Video Enhancement Standard

bytedance-video-enhancement-standard

Post-process videos with BytePlus restoration and enhancement. Upscale to 2K or 4K, interpolate up to 120 FPS, repair noise, blocking, blur, scratches, and jitter, and improve color and contrast across AI video, UGC, short drama, and archival footage.

Seedance Upscaler

bytedance-seedance-upscaler

Enhance existing videos with ByteDance’s Seedance super-resolution for cleaner 1080p, 2K, or 4K output with strong temporal consistency. Supports clips up to 10 minutes.

Avatar Omni Human 1.5

bytedance-avatar-omni-human-1.5

Animate a portrait using ByteDance's cognitive avatar model. Upload a static image and an audio track for expressive lip-sync and emotion.

MiniMax H3 Max Extend

minimax/h3-max/extend-video

Continue an existing video with a prompt describing what happens next. Add 5–15 seconds of new footage at 480p through 2K, returning the full extended video or just the continuation. Source videos must be MP4 or MOV, 1.625–60 seconds, at most 50 MB, with an aspect ratio between 0.4 and 2.5.

MiniMax H3 Max Lip Sync

minimax/h3-max/lip-sync/image-to-video

Animate a portrait or character image from supplied audio with transcription-guided lip sync. Audio must be at least 5 seconds; longer inputs are clipped to the first 15 seconds. Supports talking and singing clips from 480p through 2K.