Mistral Code Agent Latest is Mistral's direct API alias for devstral-2512, an agentic coding model built for autonomous software engineering, tool use, and long-running code tasks.
Added Jun 2, 2026
Context Window
262.1K
Max Output
32.8K
Avg output tokens (7d)
616 tokens
Input Price (Auto)
$0.40/1M
Output Price (Auto)
$2.00/1M
Cache Read (Auto)
$0.20/1M
Capabilities
Benchmarks
Benchmarks
Performance metrics and benchmarks
Sourced from Artificial Analysis.
Intelligence Index
9.4
Coding Index
31.3
Agentic Index
4.9
Agentic work
AutomationBench-AA
Workflow automation with guardrail penalties
3.1%
Better than 28% of models compared
AA-Briefcase
Agentic knowledge work (Elo)
483 Elo
Better than 26% of models compared
GDPval-AA v2
Economically valuable tasks (Elo)
689 Elo
Better than 32% of models compared
Document reasoning
GDP.pdf
Professional PDF reasoning: all-pass rate
2.4%
Better than 21% of models compared
AA-LCR v1.1
Long context reasoning with updated grading
32.3%
Better than 34% of models compared
Reasoning
HLE
Humanity's Last Exam
3.6%
Better than 7% of models compared
IFBench
Instruction-following benchmark
38.1%
Better than 34% of models compared
CritPt
Research-level physics reasoning
0.0%
Coding
Terminal-Bench v4.0
Practical coding and terminal tasks
0.0%
Better than 18% of models compared
SciCode
Python programming for scientific computing
32.8%
Better than 13% of models compared
LiveCodeBench
Contamination-free coding benchmark
44.8%
Better than 51% of models compared
Math
AIME 2025
American Invitational Mathematics Examination 2025
36.7%
Better than 37% of models compared
Knowledge
MMLU-Pro
Professional and academic subject knowledge
76.2%
Better than 53% of models compared
AA-Omniscience Accuracy
Proportion of correctly answered questions
20.8%
AA-Omniscience Hallucination Rate
Rate of incorrect answers among non-correct responses
85.3%
Legacy benchmarks
GPQA Diamond (legacy)
Graduate-level scientific reasoning
59.4%
Better than 34% of models compared
Terminal-Bench Hard (legacy)
Agentic coding and terminal use
18.9%
Better than 59% of models compared
T²-Bench Telecom (legacy)
Conversational AI agents in dual-control scenarios
24.9%
Better than 28% of models compared
AA-LCR (unversioned / legacy)
Long context reasoning evaluation
32.3%
Better than 34% of models compared
GDPval-AA (unversioned / legacy)
Economically valuable tasks
9.4%
Last updated Sep 11, 2026
Artificial AnalysisProviders
Auto routing is available for this model. Explicit provider selection is not available.
Loading provider options…
Related text models
Compare Mistral Code Agent Latest with similar models from the same provider or model family.
Mistral Code Latest
mistral-code-latestMistral Code Latest is Mistral's direct API alias for codestral-2508, a low-latency coding model for code generation, completion, fill-in-the-middle workflows, function calling, and structured output.
Mistral Medium 3.5 Thinking
mistral/mistral-medium-3.5:thinkingMistral Medium 3.5 with reasoning enabled by default (reasoning_effort=high), for complex coding, agentic, and multi-step reasoning prompts.
Mistral Medium 3.5
mistral/mistral-medium-3.5Mistral Medium 3.5 is a 128B dense open-weights flagship model for instruction-following, reasoning, coding, long-horizon agentic work, tool use, structured output, and multimodal prompts. It supports a 256k context window and configurable reasoning effort.
Mistral Small 4 119B Thinking
mistralai/mistral-small-4-119b-2603:thinkingMistral Small 4 with reasoning enabled (reasoning_effort=high). A hybrid MoE model with deep step-by-step reasoning for complex prompts, coding, and multi-step problem solving.
Mistral Small 4 119B
mistralai/mistral-small-4-119b-2603Mistral Small 4 is a hybrid MoE model that unifies instruct, reasoning, and coding behavior in a single multimodal model. It supports text and image input, native function calling, JSON output, and per-request reasoning effort controls.
Mistral Large 3 675B
mistralai/mistral-large-3-675b-instruct-2512Mistral Large 3 675B is Mistral AI's flagship language model featuring advanced rope scaling and Eagle speculative decoding. Delivers exceptional performance across reasoning, coding, and multilingual tasks.