Kimi K3
Kimi K3 is Moonshot AI's open-weight, always-thinking multimodal model for long-context reasoning, coding, tool use, and native image/video understanding.
- Reasoning
- Vision
- Video Input
- Tool Calling
- Structured Output
Added Jul 16, 2026
Model weightsPricing
Auto routing · per 1M tokens- Input
- $1.80
- Output
- $9.00
- Cache read
- $0.18
Specifications
- Context window
- 1M
- Max output
- 1M
- Avg output (7d)
- 1.1K tokens
- Longer than 74% of models
Benchmarks
Benchmarks
Sourced from Artificial Analysis.
Intelligence Index
43.6
Coding Index
76.2
Agentic Index
50.0
Agentic work
AutomationBench-AA
Workflow automation with guardrail penalties
58.3%
Better than 82% of models compared
AutomationBench-AA Tasks Completed
Fully completed workflows without guardrail violations
24.8%
Better than 30% of models compared
Harvey LAB-AA
Legal agentic work criterion pass rate
94.6%
Better than 96% of models compared
AA-Briefcase
Agentic knowledge work (Elo)
1501 Elo
Better than 86% of models compared
GDPval-AA v2
Economically valuable tasks (Elo)
1538 Elo
Better than 85% of models compared
Document reasoning
GDP.pdf
Professional PDF reasoning: all-pass rate
22.0%
Better than 76% of models compared
AA-LCR v1.1
Long context reasoning with updated grading
88.7%
Better than 99% of models compared
MLCR-AA
Medical long-context reasoning
38.3%
Better than 84% of models compared
Reasoning
HLE
Humanity's Last Exam
46.9%
Better than 94% of models compared
CritPt
Research-level physics reasoning
23.4%
Coding
Terminal-Bench v4.0
Practical coding and terminal tasks
12.6%
Better than 66% of models compared
SciCode
Python programming for scientific computing
59.5%
Better than 95% of models compared
Knowledge
AA-Omniscience Accuracy
Proportion of correctly answered questions
47.6%
AA-Omniscience Hallucination Rate
Rate of incorrect answers among non-correct responses
53.2%
Legacy benchmarks
GPQA Diamond (legacy)
Graduate-level scientific reasoning
93.5%
Better than 97% of models compared
AA-LCR (unversioned / legacy)
Long context reasoning evaluation
88.7%
Better than 99% of models compared
GDPval-AA (unversioned / legacy)
Economically valuable tasks
51.9%
Last updated Oct 5, 2026
Artificial AnalysisProviders
Choose explicit providers for this model. Auto routing remains available as the default option.
Loading provider options…