State-of-the-art audio-to-video lip synchronization using latent diffusion. Upload a talking-head video (480p+) and target audio to generate perfectly synchronized lip movements while preserving identity, pose, and background.
Added Dec 6, 2025
Approx. Price
$0.150 per video
Model Type
video-to-video
Settings
Generation controls available for this model.
Output Format
N/A
Default Duration
N/A
No configurable settings are exposed for this model yet.
Benchmarks
Benchmarks
No benchmark data is available yet for this model.
Examples
Loading examples…
Related video models
Compare LatentSync with similar models from the same provider or model family.
LTX-2.5 Fast
lightricks/ltx-2.5/fastSpeed-optimized audiovisual generation from text, an image, or a 2-20 second audio clip. Creates synchronized video and audio in one pass, with output up to 4K and optional start/end-frame control.
LTX-2.5 Pro
lightricks/ltx-2.5/proHigh-fidelity audiovisual generation from text, an image, or a 2-20 second audio clip. Creates polished synchronized video and audio in one pass, with 720p/1080p output and optional start/end-frame control.
Wan 3.0 Image-to-Video
alibaba/wan-3.0/image-to-videoAnimate a first-frame image into a cinematic video with optional last-frame guidance, synchronized audio, deep-thinking controls, and 2–30 second output.
Wan 3.0 Reference-to-Video
alibaba/wan-3.0/reference-to-videoReference-guided video generation using images, videos, and audio for subject consistency, motion, timing, and scene continuity, with 2–30 second output.
Wan 3.0 Text-to-Video
alibaba/wan-3.0/text-to-videoCinematic text-to-video generation with synchronized audio, deep-thinking prompt interpretation, 2–30 second duration, and 480p, 720p, or 1080p output.
Wan 2.2 Animate 2
wan-22-animate-2Next-generation character motion transfer with strong identity preservation and prompt-controlled backgrounds. Requires a reference image and driver video; supports 480p/720p up to 120s.