Cost-effective Kling 2.6 Standard for text-to-video and image-to-video. Smooth motion, cinematic visuals, strong prompt adherence, and 5s or 10s durations with multiple aspect ratios.
Added Feb 4, 2026
Approx. Price
$0.250 per video
Model Type
both
Settings
Generation controls available for this model.
Output Format
Default Duration
5
2 duration options
Aspect Ratio (T2V)
Default
16:9
Options (3)
Landscape (16:9), Portrait (9:16), Square (1:1)
Applies to text-to-video
Duration
Default
5
Options (2)
5 seconds, 10 seconds
5 or 10 seconds
Benchmarks
Benchmarks
No benchmark data is available yet for this model.
Related video models
Compare Kling 2.6 Standard with similar models from the same provider or model family.
Kling 3.0 Standard Motion Control
kling-v30-std-motion-controlTransfer motion from a reference video to animate a still character image. Requires an image and a motion clip, supports optional prompts, and can retain the original video audio.
Kling 3.0 Standard
kling-v30-stdKling 3.0 Standard delivers high-quality text-to-video and image-to-video with smooth motion, cinematic visuals, and strong prompt adherence. Upload an image to switch to image-to-video, with optional native audio.
Kling 2.5 Turbo Standard
kling-v25-turbo-stdImage-to-video only version of Kling 2.5 Turbo delivering cinematic motion at 720p with 5s and 10s clips. Optimized for fast, affordable production with 25% lower pricing than Kling 2.1 Standard.
Kling 2.6 Std Motion Control
kling-v26-std-motion-controlTransfer motion from a reference video onto a character image. Requires a subject image and a motion clip to animate the character with smooth, realistic movement. Supports up to 30 seconds with optional prompts and original audio retention.
Kling 3.0 Turbo Standard
kling-v3-turbo-standardKling 3.0 Turbo Standard generates fast, cost-efficient 720p video from text or a first-frame image. Supports stable motion, multi-shot storyboard prompts, and 3-15 second durations.
Kling Video O1 Standard
kling-video-o1-standardKuaishou's unified multi-modal video model (Standard tier) optimized for cost efficiency. Supports text-only input for text-to-video, image input for image-to-video, reference images/video for reference-based generation, or video-only input for natural language video editing.