Animate one or two people from an image and audio tracks.
Added Sep 8, 2026
Approx. Price
$0.090 per video
Model Type
image-to-video
Settings
Generation controls available for this model.
Output Format
Default Duration
N/A
Audio URL (one person)
Default
N/A
Left person audio URL (two people)
Default
N/A
Mask image URL (one person, optional)
Default
N/A
People
Default
single
Options (2)
One person, Two people
Resolution
Default
480p
Options (2)
480p, 720p
Right person audio URL (two people)
Default
N/A
Seed
Default
-1
Speaking Order
Default
meanwhile
Options (3)
Together, Left, then right, Right, then left
Benchmarks
Benchmarks
No benchmark data is available yet for this model.
Related video models
Compare InfiniteTalk with similar models from the same provider or model family.
MiniMax H3 Spicy Image-to-Video
wavespeed-ai/minimax-h3/image-to-video-spicyOpen-weights MiniMax H3 image-to-video generation with expressive unrestricted motion, native stereo audio, optional last-frame control, 3–15 second clips, and 480p or 768p output.
LTX-2.3 Spicy Image-to-Video
wavespeed-ai/ltx-2.3-spicy/image-to-videoLTX-2.3 Spicy image-to-video turns a reference image and prompt into expressive clips with native audio, style-tuned motion, 480p/720p/1080p output, and 3-20 second durations.
LTX-2.3 Spicy LoRA Image-to-Video
wavespeed-ai/ltx-2.3-spicy/image-to-video-loraLTX-2.3 Spicy LoRA image-to-video adds selectable LoRA presets and per-LoRA strength overrides on top of the same reference-image animation workflow.
LongCat Avatar 1.5
wavespeed-ai/longcat-avatar-1.5Upgraded audio-driven talking or singing avatar generation from a single image with sharper lip sync and faster generation. Supports 480p/720p output up to 30 seconds.
LongCat Avatar 1.5 Multi
wavespeed-ai/longcat-avatar-1.5/multiAudio-driven two-person avatar generation from a single image and left/right audio tracks. Supports simultaneous or sequential dialogue, 480p/720p output, and up to 30 seconds of audio.
Music Video Generator
wavespeed-ai/music-video-generatorGenerate a lip-synced music video from an audio track, with optional reference portraits (1-3 images). Supports cinematic scene transitions up to 10 minutes at 480p or 720p.