Hunyuan Video 1.5 generates 5 or 8 second clips from text or an input image. Supports 480p/720p and landscape or portrait runs routed automatically based on whether an image is attached.
Added Nov 24, 2025
Approx. Price
$0.150 per video
Model Type
both
Settings
Generation controls available for this model.
Output Format
Default Duration
5
2 duration options
Duration
Default
5
Options (2)
5 seconds, 8 seconds
Clip length in seconds
Orientation
Default
landscape
Options (2)
Landscape (16:9), Portrait (9:16)
Landscape or portrait (text-to-video only)
Resolution
Default
720p
Options (2)
480p, 720p
Choose 480p or 720p output
Benchmarks
Benchmarks
Human preference benchmarks sourced from LMArena.
Text to Video
#34 / 44
Arena Score
1169.3
Votes
4,290
Confidence Interval
1153.0 - 1185.6
Image to Video
#36 / 44
Arena Score
1195.8
Votes
5,478
Confidence Interval
1180.3 - 1211.2
Published 2026-08-13 · Matched as hunyuan-video-1.5
LMArena DatasetExamples
Loading examples…
Related video models
Compare Hunyuan Video 1.5 with similar models from the same provider or model family.
Hunyuan Video
hunyuan-videoHunyuan Video text-to-video generator creates high-quality 720p videos with customizable resolution, aspect ratio, and frame count. Features pro mode for enhanced quality.
LTX-2.5 Fast
lightricks/ltx-2.5/fastSpeed-optimized audiovisual generation from text, an image, or a 2-20 second audio clip. Creates synchronized video and audio in one pass, with output up to 4K and optional start/end-frame control.
LTX-2.5 Pro
lightricks/ltx-2.5/proHigh-fidelity audiovisual generation from text, an image, or a 2-20 second audio clip. Creates polished synchronized video and audio in one pass, with 720p/1080p output and optional start/end-frame control.
Wan 3.0 Image-to-Video
alibaba/wan-3.0/image-to-videoAnimate a first-frame image into a cinematic video with optional last-frame guidance, synchronized audio, deep-thinking controls, and 2–30 second output.
Wan 3.0 Reference-to-Video
alibaba/wan-3.0/reference-to-videoReference-guided video generation using images, videos, and audio for subject consistency, motion, timing, and scene continuity, with 2–30 second output.
Wan 3.0 Text-to-Video
alibaba/wan-3.0/text-to-videoCinematic text-to-video generation with synchronized audio, deep-thinking prompt interpretation, 2–30 second duration, and 480p, 720p, or 1080p output.