Devstral 2 123B is a 123 billion parameter model from Mistral AI optimized for coding and development tasks. Features advanced reasoning capabilities for software engineering workflows.
Added Dec 9, 2025
Model weightsContext Window
262.1K
Max Output
65.5K
Avg output tokens (7d)
495 tokens
Input Price (Auto)
$0.40/1M
Output Price (Auto)
$1.40/1M
Cache Read (Auto)
$0.20/1M
Benchmarks
Benchmarks
Performance metrics and benchmarks
Sourced from Artificial Analysis.
Intelligence Index
9.4
Coding Index
31.3
Agentic Index
4.9
Agentic work
AutomationBench-AA
Workflow automation with guardrail penalties
3.1%
Better than 28% of models compared
AA-Briefcase
Agentic knowledge work (Elo)
483 Elo
Better than 26% of models compared
GDPval-AA v2
Economically valuable tasks (Elo)
689 Elo
Better than 32% of models compared
Document reasoning
GDP.pdf
Professional PDF reasoning: all-pass rate
2.4%
Better than 20% of models compared
AA-LCR v1.1
Long context reasoning with updated grading
32.3%
Better than 34% of models compared
Reasoning
HLE
Humanity's Last Exam
3.6%
Better than 7% of models compared
IFBench
Instruction-following benchmark
38.1%
Better than 34% of models compared
CritPt
Research-level physics reasoning
0.0%
Coding
Terminal-Bench v4.0
Practical coding and terminal tasks
0.0%
Better than 17% of models compared
SciCode
Python programming for scientific computing
32.8%
Better than 13% of models compared
LiveCodeBench
Contamination-free coding benchmark
44.8%
Better than 51% of models compared
Math
AIME 2025
American Invitational Mathematics Examination 2025
36.7%
Better than 37% of models compared
Knowledge
MMLU-Pro
Professional and academic subject knowledge
76.2%
Better than 53% of models compared
AA-Omniscience Accuracy
Proportion of correctly answered questions
20.8%
AA-Omniscience Hallucination Rate
Rate of incorrect answers among non-correct responses
85.3%
Legacy benchmarks
GPQA Diamond (legacy)
Graduate-level scientific reasoning
59.4%
Better than 34% of models compared
Terminal-Bench Hard (legacy)
Agentic coding and terminal use
18.9%
Better than 59% of models compared
T²-Bench Telecom (legacy)
Conversational AI agents in dual-control scenarios
24.9%
Better than 28% of models compared
AA-LCR (unversioned / legacy)
Long context reasoning evaluation
32.3%
Better than 34% of models compared
GDPval-AA (unversioned / legacy)
Economically valuable tasks
9.4%
Last updated Sep 13, 2026
Artificial AnalysisProviders
Auto routing is available for this model. Explicit provider selection is not available.
Loading provider options…
Related text models
Compare Devstral 2 123B with similar models from the same provider or model family.
Mistral Large 3 675B
mistralai/mistral-large-3-675b-instruct-2512Mistral Large 3 675B is Mistral AI's flagship language model featuring advanced rope scaling and Eagle speculative decoding. Delivers exceptional performance across reasoning, coding, and multilingual tasks.
Ministral 3 14B
mistralai/ministral-14b-instruct-2512Ministral 3 14B is a balanced model in the Ministral 3 family, designed for edge deployment. A powerful, efficient language model with vision capabilities, fine-tuned for instruction tasks. Features multilingual support, strong system prompt adherence, and native function calling. Apache 2.0 licensed.
Mistral Small 3 24B (2501)
mistralai/mistral-small-24b-instruct-2501Mistral Small 3 24B (2501) hosted by IONOS in Berlin, Germany. Zero data retention.
Mixtral 8x22B
mistralai/mixtral-8x22b-instruct-v0.1Mixtral 8x22B is a powerful sparse Mixture of Experts (MoE) model with 141B total parameters and 39B active per token. Features a 64K context window, exceptional math performance, and cost-efficient inference. Supports English, French, German, Spanish, and Italian. Apache 2.0 licensed.
Ministral 14B
mistralai/ministral-14b-2512Ministral 14B is a powerful 14B parameter model from Mistral AI with vision capabilities, offering frontier performance in a compact size.
Ministral 3B
mistralai/ministral-3b-2512Ministral 3B is a tiny, efficient 3B parameter model from Mistral AI with vision capabilities, designed for edge deployment.