MiniMax H3 Max Lip Sync
Animate a portrait or character image from supplied audio with transcription-guided lip sync. Audio must be at least 5 seconds; longer inputs are clipped to the first 15 seconds. Supports talking and singing clips from 480p through 2K.
Added Sep 18, 2026
Details
- Starting Price
- From $0.250 per video
- Model type
- image-to-video
Final price depends on the selected settings and is shown before generation.
Settings
Generation controls available for this model.
Output Format
up to 1,080p
Default Duration
N/A
Resolution
Select
Default
768p
Options (4)
480p, 768p, 1080p, 2K
Output resolution
Seed
Number
Default
-1
Use -1 for a random output
Transcription-guided lip sync
Toggle
Default
Yes
Use audio transcription to improve lip-sync accuracy
Benchmarks
No public benchmark scores for this model yet.
Examples
Loading examples…