Text-to-video and image-to-video with ultra-smooth motion, cinematic visuals, and precise prompt control. Supports 5s and 10s outputs and multiple aspect ratios.
Added Oct 7, 2025
Approx. Price
$0.350 per video
Model Type
both
Settings
Generation controls available for this model.
Output Format
Default Duration
5
2 duration options
Aspect Ratio (T2V)
Default
16:9
Options (3)
Landscape (16:9), Portrait (9:16), Square (1:1)
Applies to text-to-video
Duration
Default
5
Options (2)
5 seconds, 10 seconds
5 or 10 seconds
Benchmarks
Benchmarks
No benchmark data is available yet for this model.
Examples
Loading examples…
Related video models
Compare Kling 2.5 Turbo Pro with similar models from the same provider or model family.
Kling 3.0 Turbo Pro
kling-v3-turbo-proKling 3.0 Turbo Pro generates high quality 1080p video from text or a first-frame image with stronger prompt consistency, stable motion, multi-shot storyboards, and 3-15 second durations.
Kling 2.5 Turbo Standard
kling-v25-turbo-stdImage-to-video only version of Kling 2.5 Turbo delivering cinematic motion at 720p with 5s and 10s clips. Optimized for fast, affordable production with 25% lower pricing than Kling 2.1 Standard.
Kling 3.0 Turbo Standard
kling-v3-turbo-standardKling 3.0 Turbo Standard generates fast, cost-efficient 720p video from text or a first-frame image. Supports stable motion, multi-shot storyboard prompts, and 3-15 second durations.
Kling 3.0 Pro Motion Control
kling-v30-pro-motion-controlTransfer motion from a reference video to animate a still character image with Kling 3.0 Pro. Requires an image and a motion clip, supports optional prompts, and can retain the original video audio.
Kling 3.0 Pro
kling-v30-proKling 3.0 Pro delivers top-tier text-to-video and image-to-video generation with smooth motion, cinematic visuals, and strong prompt adherence. Upload an image to switch to image-to-video, with optional native audio. Warning: Kling 3.0 Pro is currently timing out frequently due to high demand.
Kling V2 Avatar (Pro)
kling-v2-avatar-proCreates social-ready talking avatars from one portrait and your audio with sharper detail, stable motion, and strong identity consistency. Optional prompt to nudge camera feel, expression, or mood.