Vidu Q3 text-to-video and image-to-video with high visual fidelity, multiple styles, 540p/720p/1080p output, 1-16s duration, and optional audio plus background music.
Added Jan 31, 2026
Starting Price
From $0.350 per video
Final price depends on the selected settings and is shown before generation.
Model Type
both
Settings
Generation controls available for this model.
Output Format
Default Duration
5
16 duration options
Aspect Ratio (T2V only)
Default
4:3
Options (5)
Landscape (16:9), Standard (4:3), Square (1:1), Portrait (3:4) +1 more
Applies to text-to-video
Background Music
Default
Yes
Add background music
Duration
Default
5
Options (16)
1 second, 2 seconds, 3 seconds, 4 seconds +12 more
Video length in seconds (1-16)
Generate Audio
Default
Yes
Generate synchronized audio
Motion
Default
auto
Options (4)
Auto, Small, Medium, Large
Movement intensity
Resolution
Default
720p
Options (3)
540p, 720p, 1080p
Output resolution
Style (T2V only)
Default
general
Options (2)
General, Anime
Visual style for text-to-video
Benchmarks
Benchmarks
No benchmark data is available yet for this model.
Examples
Loading examples…
Related video models
Compare Vidu Q3 with similar models from the same provider or model family.
Vidu Q3 Pro
vidu-q3-proVidu Q3 Pro text-to-video, image-to-video, and start/end-frame video generation with high visual fidelity, 540p/720p/1080p output, 1-16s duration, and optional audio plus background music.
Vidu Q1
vidu-videoVidu Q1 video generation model. Creates high-quality 5-second videos. Supports both text-to-video and image-to-video generation with customizable visual styles (general or anime), movement amplitude control, and fixed 16:9 output.
MiniMax H3 Max Extend
minimax/h3-max/extend-videoContinue an existing video with a prompt describing what happens next. Add 5–15 seconds of new footage at 480p through 2K, returning the full extended video or just the continuation. Source videos must be MP4 or MOV, 1.625–60 seconds, at most 50 MB, with an aspect ratio between 0.4 and 2.5.
MiniMax H3 Max Lip Sync
minimax/h3-max/lip-sync/image-to-videoAnimate a portrait or character image from supplied audio with transcription-guided lip sync. Audio must be at least 5 seconds; longer inputs are clipped to the first 15 seconds. Supports talking and singing clips from 480p through 2K.
Wan 3.0 Prime Video Edit
alibaba/wan-3.0-prime/video-editAccelerated Wan 3.0 editing for an existing video with optional image or audio references. Uses the first 15 seconds of the source and supports 2–15 second output at 480p, 720p, or 1080p.
Wan 3.0 Prime Video Extend
alibaba/wan-3.0-prime/video-extendAccelerated video extension that appends 2–30 seconds with optional target last-frame guidance. Source audio is preserved, and clips longer than 120 seconds use their final 120 seconds as context.