Kling 2.1 Standard image-to-video model. Creates high-quality videos from images with text prompts. Requires an input image.
Added May 29, 2025
Approx. Price
$0.140 per video
Model Type
image-to-video
Settings
Generation controls available for this model.
Output Format
Default Duration
5
2 duration options
Aspect Ratio
Default
16:9
Options (3)
Landscape (16:9), Portrait (9:16), Square (1:1)
Choose between landscape (16:9), portrait (9:16), or square (1:1) orientation
CFG Scale
Default
0.5
Controls how closely the generation follows the prompt (0.0-1.0)
Duration
Default
5
Options (2)
5 seconds, 10 seconds (double price)
Length of the generated video in seconds
Negative Prompt
Default
blur, distort, and low quality
What to avoid in the video (default: blur, distort, and low quality)
Benchmarks
Benchmarks
Human preference benchmarks sourced from Artificial Analysis.
Image to Video
#48 / 72
ELO
1170.0
Appearances
3,313
95% CI
-10/10
Release Date 2025-05 · Matched as Kling 2.1 Standard
Artificial Analysis APIExamples
Loading examples…
Related video models
Compare Kling 2.1 Standard with similar models from the same provider or model family.
Kling 3.0 Turbo Standard
kling-v3-turbo-standardKling 3.0 Turbo Standard generates fast, cost-efficient 720p video from text or a first-frame image. Supports stable motion, multi-shot storyboard prompts, and 3-15 second durations.
Kling 3.0 Standard Motion Control
kling-v30-std-motion-controlTransfer motion from a reference video to animate a still character image. Requires an image and a motion clip, supports optional prompts, and can retain the original video audio.
Kling 2.6 Standard
kling-v26-stdCost-effective Kling 2.6 Standard for text-to-video and image-to-video. Smooth motion, cinematic visuals, strong prompt adherence, and 5s or 10s durations with multiple aspect ratios.
Kling 3.0 Standard
kling-v30-stdKling 3.0 Standard delivers high-quality text-to-video and image-to-video with smooth motion, cinematic visuals, and strong prompt adherence. Upload an image to switch to image-to-video, with optional native audio.
Kling Video O1 Standard
kling-video-o1-standardKuaishou's unified multi-modal video model (Standard tier) optimized for cost efficiency. Supports text-only input for text-to-video, image input for image-to-video, reference images/video for reference-based generation, or video-only input for natural language video editing.
Kling V2 Avatar (Standard)
kling-v2-avatar-standardTurns a single portrait and one audio track into a realistic talking avatar with accurate lip sync, expressive facial motion, and consistent identity. Optional prompt can guide mood or energy.