Audio-driven talking or singing avatar generation from a single image with lip-synced motion and consistent identity. Supports 480p/720p output up to 2 minutes.
Added Dec 24, 2025
Approx. Price
$0.150 per video
Model Type
image-to-video
Settings
Generation controls available for this model.
Output Format
Default Duration
N/A
Prompt
Default
N/A
Optional expression/style prompt
Resolution
Default
480p
Options (2)
480p, 720p
Video resolution
Benchmarks
Benchmarks
No benchmark data is available yet for this model.
Examples
Loading examples…
Related video models
Compare LongCat Avatar with similar models from the same provider or model family.
LongCat Avatar 1.5
wavespeed-ai/longcat-avatar-1.5Upgraded audio-driven talking or singing avatar generation from a single image with sharper lip sync and faster generation. Supports 480p/720p output up to 30 seconds.
LongCat Avatar 1.5 Multi
wavespeed-ai/longcat-avatar-1.5/multiAudio-driven two-person avatar generation from a single image and left/right audio tracks. Supports simultaneous or sequential dialogue, 480p/720p output, and up to 30 seconds of audio.
P-Video Avatar
pruna-ai/p-video/avatarImage-and-audio avatar video generation for speech-driven talking-head clips, with 720p and 1080p output.
Kling V2 Avatar (Pro)
kling-v2-avatar-proCreates social-ready talking avatars from one portrait and your audio with sharper detail, stable motion, and strong identity consistency. Optional prompt to nudge camera feel, expression, or mood.
Kling V2 Avatar (Standard)
kling-v2-avatar-standardTurns a single portrait and one audio track into a realistic talking avatar with accurate lip sync, expressive facial motion, and consistent identity. Optional prompt can guide mood or energy.
Avatar Omni Human 1.5
bytedance-avatar-omni-human-1.5Animate a portrait using ByteDance's cognitive avatar model. Upload a static image and an audio track for expressive lip-sync and emotion.