Expressive image-to-video avatar generation with natural facial performance, realistic body motion, accurate A/V sync, and optional driving audio for lip-sync mode.
Added Apr 3, 2026
Approx. Price
$0.250 per video
Model Type
image-to-video
Settings
Generation controls available for this model.
Output Format
Default Duration
5
30 duration options
Driving Audio URL
Default
N/A
Optional audio URL for lip-sync mode. If omitted, audio is generated from the prompt.
Duration
Default
5
Options (30)
1 second, 2 seconds, 3 seconds, 4 seconds +26 more
Length of the generated video in seconds.
Guidance Scale
Default
5
Optional classifier-free guidance scale (0-20).
Inference Steps
Default
8
Optional denoising steps (1-50).
Resolution
Default
256p
Options (4)
256p, 540p, 720p, 1080p
Output resolution.
Safety Checker
Default
Yes
Run prompt and image safety checks before generation.
Seed
Optional seed for reproducible output.
Benchmarks
Benchmarks
No benchmark data is available yet for this model.
Examples
Loading examples…
Related video models
Compare DaVinci MagiHuman with similar models from the same provider or model family.
LTX-2.5 Fast
lightricks/ltx-2.5/fastSpeed-optimized audiovisual generation from text, an image, or a 2-20 second audio clip. Creates synchronized video and audio in one pass, with output up to 4K and optional start/end-frame control.
LTX-2.5 Pro
lightricks/ltx-2.5/proHigh-fidelity audiovisual generation from text, an image, or a 2-20 second audio clip. Creates polished synchronized video and audio in one pass, with 720p/1080p output and optional start/end-frame control.
FLUX.3
flux-3Generate up to 20-second videos with native audio from a prompt, a start image, start/end frames, multiple keyframes, or a source clip. FLUX.3 chooses the matching workflow automatically from what you attach.
Pixelcut Video Background Remover
pixelcut/video-background-removalRemove video backgrounds with frame-by-frame AI segmentation and temporally consistent edges. Supports transparent output, preset solid backgrounds, custom RGB backgrounds, and common video formats.
Bernini R Video
bernini-r-videoBernini R video generation and editing. A single NanoGPT model id routes prompt-only requests to text-to-video, up to 5 image inputs to reference-to-video, video inputs to edit-video, and image plus video inputs to reference-edit-video.
Luma Ray 3.2
luma/agent/ray/v3.2Cinematic text-to-video and image-to-video generation with strong motion control, optional reference images, seamless loops, 540p/720p/1080p output, and 5 or 10 second clips.