Video-to-video generation. Requires input video + prompt. Optional reference image. Only the first 5 seconds of the input are used by the model; we charge per-second based on your input video length.
Approx. Price
$1.25 per video
Model Type
video-to-video
Settings
Generation controls available for this model.
Output Format
Default Duration
5
Aspect Ratio
Default
auto
Options (6)
Auto (match input), 16:9, 4:3, 1:1 +2 more
Default auto uses the input video ratio
Duration
Default
5
Benchmarks
Benchmarks
Human preference benchmarks sourced from LMArena.
Video Edit
#8 / 8
Arena Score
1182.7
Votes
9,817
Confidence Interval
1174.3 - 1191.2
Published 2026-08-13 · Matched as runway-gen4-aleph
LMArena DatasetExamples
Loading examples…
Related video models
Compare Runway Gen-4 Aleph (V2V) with similar models from the same provider or model family.
Runway Gen-4.5
runway-gen-45Runway Gen-4.5 video generation via Runware. Supports text-to-video and image-to-video with multiple aspect ratios, 5/8/10 second durations, and 24 FPS output.
Wan 2.2 (V2V)
wan-wavespeed-video-editEdit an existing video using a natural language prompt. Examples: "Change the color of the clothes to yellow", "Change the woman to a handsome boy". Supports 480p or 720p output, up to 120 seconds.
LTX-2.5 Fast
lightricks/ltx-2.5/fastSpeed-optimized audiovisual generation from text, an image, or a 2-20 second audio clip. Creates synchronized video and audio in one pass, with output up to 4K and optional start/end-frame control.
LTX-2.5 Pro
lightricks/ltx-2.5/proHigh-fidelity audiovisual generation from text, an image, or a 2-20 second audio clip. Creates polished synchronized video and audio in one pass, with 720p/1080p output and optional start/end-frame control.
Wan 3.0 Image-to-Video
alibaba/wan-3.0/image-to-videoAnimate a first-frame image into a cinematic video with optional last-frame guidance, synchronized audio, deep-thinking controls, and 2–30 second output.
Wan 3.0 Reference-to-Video
alibaba/wan-3.0/reference-to-videoReference-guided video generation using images, videos, and audio for subject consistency, motion, timing, and scene continuity, with 2–30 second output.