State-of-the-art audio-to-video lip synchronization using latent diffusion. Upload a talking-head video (480p+) and target audio to generate perfectly synchronized lip movements while preserving identity, pose, and background.
Added Dec 6, 2025
Approx. Price
$0.150 per video
Model Type
video-to-video
Settings
Generation controls available for this model.
Output Format
N/A
Default Duration
N/A
No configurable settings are exposed for this model yet.
Benchmarks
Benchmarks
No benchmark data is available yet for this model.
Examples
Loading examples…
Related video models
Compare LatentSync with similar models from the same provider or model family.
InfiniteTalk
wavespeed-ai/infinitetalkAnimate one or two people from an image and audio tracks.
P-Video Edit
pruna-ai/p-video/editInstruction-based editing for videos up to 15 seconds, with optional reference-image guidance, Draft and Full quality modes, prompt enhancement, and source-audio preservation.
MiniMax H3 Max Turbo
minimax/h3-max-turboMiniMax H3 Max Turbo brings H3 Max prompt understanding and aesthetics to video generation at roughly twice the speed and half the cost while targeting 97% of its quality. Create 5–15 second 480p, 768p, or 1080p clips from text or a starting image, with optional first/last-frame transitions.
MiniMax H3 Spicy Image-to-Video
wavespeed-ai/minimax-h3/image-to-video-spicyOpen-weights MiniMax H3 image-to-video generation with expressive unrestricted motion, native stereo audio, optional last-frame control, 3–15 second clips, and 480p or 768p output.
Gemini Omni Flash 1.1
google/gemini-omni-flash/v1.1Google’s multimodal video model for text-to-video, image animation with optional end frames, multimodal reference generation, and instruction-based video editing. Generates synchronized native audio at resolutions from 360p through 4K.
MiniMax H3 Max
minimax/h3-maxGenerate from text, first/last frames, or image/video/audio references.