Vidu Q1

vidu-video

Vidu Q1

vidu-video

Vidu Q1 video generation model. Creates high-quality 5-second videos. Supports both text-to-video and image-to-video generation with customizable visual styles (general or anime), movement amplitude control, and fixed 16:9 output.

Added Jul 10, 2025

Price

$0.150 per video

Model Type

both

Settings

Generation controls available for this model.

Output Format

N/A

Default Duration

5

Duration

Default

5

Movement Amplitude

Select

Default

auto

Options (4)

Auto, Small, Medium, Large

Amount of movement in the generated video

Size

Select

Default

16:9

Options (1)

16:9 (1920x1080 / Landscape) - Only supported resolution

Video resolution - Vidu only supports 1920x1080 (16:9)

Style

Select

Default

general

Options (2)

General, Anime

Visual style for the video (only available for text-to-video)

Benchmarks

Human preference benchmarks sourced from Artificial Analysis.

Image to Video

#67 / 76

ELO

1032.0

Appearances

2,421

95% CI

-13/13

Release Date 2025-04 · Matched as Vidu Q1

Artificial Analysis API

Examples

Loading examples…

Compare Vidu Q1 with similar models from the same provider or model family.

Vidu Q3 Pro

vidu-q3-pro

Vidu Q3 Pro text-to-video, image-to-video, and start/end-frame video generation with high visual fidelity, 540p/720p/1080p output, 1-16s duration, and optional audio plus background music.

Vidu Q3

vidu-q3

Vidu Q3 text-to-video and image-to-video with high visual fidelity, multiple styles, 540p/720p/1080p output, 1-16s duration, and optional audio plus background music.

MiniMax H3 Max Extend

minimax/h3-max/extend-video

Continue an existing video with a prompt describing what happens next. Add 5–15 seconds of new footage at 480p through 2K, returning the full extended video or just the continuation. Source videos must be MP4 or MOV, 1.625–60 seconds, at most 50 MB, with an aspect ratio between 0.4 and 2.5.

MiniMax H3 Max Lip Sync

minimax/h3-max/lip-sync/image-to-video

Animate a portrait or character image from supplied audio with transcription-guided lip sync. Audio must be at least 5 seconds; longer inputs are clipped to the first 15 seconds. Supports talking and singing clips from 480p through 2K.

Wan 3.0 Prime Video Edit

alibaba/wan-3.0-prime/video-edit

Accelerated Wan 3.0 editing for an existing video with optional image or audio references. Uses the first 15 seconds of the source and supports 2–15 second output at 480p, 720p, or 1080p.

Wan 3.0 Prime Video Extend

alibaba/wan-3.0-prime/video-extend

Accelerated video extension that appends 2–30 seconds with optional target last-frame guidance. Source audio is preserved, and clips longer than 120 seconds use their final 120 seconds as context.