The reasoning version of Thinking Machines' 276B-parameter open-weights multimodal Mixture-of-Experts model with 12B active parameters. It reasons over text and images, and is designed for agentic coding, tool use, instruction following, and long workflows with controllable thinking effort.
Added Jul 31, 2026
Model weightsContext Window
524.3K
Max Output
32.8K
Input Price (Auto)
$0.45/1M
Output Price (Auto)
$1.20/1M
Cache Read (Auto)
$0.100/1M
Capabilities
Benchmarks
Benchmarks
Performance metrics and benchmarks
Sourced from Artificial Analysis.
Intelligence Index
26.1
Coding Index
52.9
Agentic Index
25.0
Agentic work
AutomationBench-AA
Workflow automation with guardrail penalties
4.8%
Better than 32% of models compared
AA-Briefcase
Agentic knowledge work (Elo)
912 Elo
Better than 48% of models compared
GDPval-AA v2
Economically valuable tasks (Elo)
1191 Elo
Better than 64% of models compared
Document reasoning
GDP.pdf
Professional PDF reasoning: all-pass rate
11.2%
Better than 51% of models compared
AA-LCR v1.1
Long context reasoning with updated grading
75.7%
Better than 79% of models compared
Reasoning
HLE
Humanity's Last Exam
33.3%
Better than 84% of models compared
CritPt
Research-level physics reasoning
8.3%
Coding
Terminal-Bench v4.0
Practical coding and terminal tasks
1.0%
Better than 47% of models compared
SciCode
Python programming for scientific computing
49.7%
Better than 56% of models compared
Knowledge
AA-Omniscience Accuracy
Proportion of correctly answered questions
33.2%
AA-Omniscience Hallucination Rate
Rate of incorrect answers among non-correct responses
63.0%
Legacy benchmarks
GPQA Diamond (legacy)
Graduate-level scientific reasoning
89.5%
Better than 87% of models compared
AA-LCR (unversioned / legacy)
Long context reasoning evaluation
75.7%
Better than 79% of models compared
GDPval-AA (unversioned / legacy)
Economically valuable tasks
34.6%
Last updated Sep 9, 2026
Artificial AnalysisProviders
Choose explicit providers for this model. Auto routing remains available as the default option.
Loading provider options…
Related text models
Compare Inkling Small Thinking with similar models from the same provider or model family.
Inkling Small
thinkingmachines/Inkling-SmallThe direct-answer version of Thinking Machines' 276B-parameter open-weights multimodal Mixture-of-Experts model with 12B active parameters. It accepts text and images, and is designed for coding, tool use, instruction following, and general conversational work.
Inkling
thinkingmachines/inklingThe non-thinking version of Thinking Machines' 975B-parameter open-weights Mixture-of-Experts generalist with 41B active parameters. It gives faster direct answers across text, images, and audio, and is built for agentic coding, tool use, detailed instruction following, and long-context work.
Inkling Thinking
thinkingmachines/inkling:thinkingThe thinking version of Thinking Machines' 975B-parameter open-weights Mixture-of-Experts generalist with 41B active parameters. It reasons natively over text, images, and audio, and is built for agentic coding, tool use, detailed instruction following, long-context work, and controllable thinking effort.
Mistral Small 4 119B Thinking
mistralai/mistral-small-4-119b-2603:thinkingMistral Small 4 with reasoning enabled (reasoning_effort=high). A hybrid MoE model with deep step-by-step reasoning for complex prompts, coding, and multi-step problem solving.
Mistral Small 4 119B
mistralai/mistral-small-4-119b-2603Mistral Small 4 is a hybrid MoE model that unifies instruct, reasoning, and coding behavior in a single multimodal model. It supports text and image input, native function calling, JSON output, and per-request reasoning effort controls.
Mistral Devstral Small 2505
mistralai/Devstral-Small-2505OpenHands+Devstral is 100% local 100% open, and is SOTA for the category on SWE-Bench Verified: 46.8% accuracy.