Vidu Q3 text-to-video and image-to-video with high visual fidelity, multiple styles, 540p/720p/1080p output, 1-16s duration, and optional audio plus background music.
Added Jan 31, 2026
Approx. Price
$0.350 per video
Model Type
both
Settings
Generation controls available for this model.
Output Format
Default Duration
5
16 duration options
Aspect Ratio (T2V only)
Default
4:3
Options (5)
Landscape (16:9), Standard (4:3), Square (1:1), Portrait (3:4) +1 more
Applies to text-to-video
Background Music
Default
Yes
Add background music
Duration
Default
5
Options (16)
1 second, 2 seconds, 3 seconds, 4 seconds +12 more
Video length in seconds (1-16)
Generate Audio
Default
Yes
Generate synchronized audio
Motion
Default
auto
Options (4)
Auto, Small, Medium, Large
Movement intensity
Resolution
Default
720p
Options (3)
540p, 720p, 1080p
Output resolution
Style (T2V only)
Default
general
Options (2)
General, Anime
Visual style for text-to-video
Benchmarks
Benchmarks
No benchmark data is available yet for this model.
Related video models
Compare Vidu Q3 with similar models from the same provider or model family.
Vidu Q3 Pro
vidu-q3-proVidu Q3 Pro text-to-video, image-to-video, and start/end-frame video generation with high visual fidelity, 540p/720p/1080p output, 1-16s duration, and optional audio plus background music.
Vidu Q1
vidu-videoVidu Q1 video generation model. Creates high-quality 5-second videos. Supports both text-to-video and image-to-video generation with customizable visual styles (general or anime), movement amplitude control, and fixed 16:9 output.
LTX-2.5 Fast
lightricks/ltx-2.5/fastSpeed-optimized audiovisual generation from text, an image, or a 2-20 second audio clip. Creates synchronized video and audio in one pass, with output up to 4K and optional start/end-frame control.
LTX-2.5 Pro
lightricks/ltx-2.5/proHigh-fidelity audiovisual generation from text, an image, or a 2-20 second audio clip. Creates polished synchronized video and audio in one pass, with 720p/1080p output and optional start/end-frame control.
Wan 3.0 Image-to-Video
alibaba/wan-3.0/image-to-videoAnimate a first-frame image into a cinematic video with optional last-frame guidance, synchronized audio, deep-thinking controls, and 2–30 second output.
Wan 3.0 Reference-to-Video
alibaba/wan-3.0/reference-to-videoReference-guided video generation using images, videos, and audio for subject consistency, motion, timing, and scene continuity, with 2–30 second output.