The thinking version of Thinking Machines' 975B-parameter open-weights Mixture-of-Experts generalist with 41B active parameters. It reasons natively over text, images, and audio, and is built for agentic coding, tool use, detailed instruction following, long-context work, and controllable thinking effort.
Added Jul 15, 2026
Model weightsContext Window
1.0M
Max Output
32.8K
Avg output tokens (7d)
944 tokens
Input Price (Auto)
$1.00/1M
Output Price (Auto)
$4.05/1M
Cache Read (Auto)
$0.17/1M
Capabilities
Benchmarks
Benchmarks
Performance metrics and benchmarks
Sourced from Artificial Analysis.
Intelligence Index
25.5
Coding Index
52.1
Agentic Index
24.3
Agentic work
AutomationBench-AA
Workflow automation with guardrail penalties
5.0%
Better than 33% of models compared
AutomationBench-AA Tasks Completed
Fully completed workflows without guardrail violations
0.3%
Better than 7% of models compared
AA-Briefcase
Agentic knowledge work (Elo)
836 Elo
Better than 43% of models compared
GDPval-AA v2
Economically valuable tasks (Elo)
1165 Elo
Better than 60% of models compared
Document reasoning
GDP.pdf
Professional PDF reasoning: all-pass rate
12.8%
Better than 58% of models compared
AA-LCR v1.1
Long context reasoning with updated grading
77.3%
Better than 81% of models compared
Reasoning
HLE
Humanity's Last Exam
31.9%
Better than 83% of models compared
CritPt
Research-level physics reasoning
5.4%
Coding
Terminal-Bench v4.0
Practical coding and terminal tasks
1.0%
Better than 47% of models compared
SciCode
Python programming for scientific computing
47.0%
Better than 50% of models compared
Knowledge
AA-Omniscience Accuracy
Proportion of correctly answered questions
41.5%
AA-Omniscience Hallucination Rate
Rate of incorrect answers among non-correct responses
67.7%
Legacy benchmarks
GPQA Diamond (legacy)
Graduate-level scientific reasoning
87.2%
Better than 83% of models compared
AA-LCR (unversioned / legacy)
Long context reasoning evaluation
77.3%
Better than 81% of models compared
GDPval-AA (unversioned / legacy)
Economically valuable tasks
33.3%
Last updated Sep 14, 2026
Artificial AnalysisProviders
Choose explicit providers for this model. Auto routing remains available as the default option.
Loading provider options…
Related text models
Compare Inkling Thinking with similar models from the same provider or model family.
Inkling
thinkingmachines/inklingThe non-thinking version of Thinking Machines' 975B-parameter open-weights Mixture-of-Experts generalist with 41B active parameters. It gives faster direct answers across text, images, and audio, and is built for agentic coding, tool use, detailed instruction following, and long-context work.
Inkling Small
thinkingmachines/Inkling-SmallThe direct-answer version of Thinking Machines' 276B-parameter open-weights multimodal Mixture-of-Experts model with 12B active parameters. It accepts text and images, and is designed for coding, tool use, instruction following, and general conversational work.
Inkling Small Thinking
thinkingmachines/Inkling-Small:thinkingThe reasoning version of Thinking Machines' 276B-parameter open-weights multimodal Mixture-of-Experts model with 12B active parameters. It reasons over text and images, and is designed for agentic coding, tool use, instruction following, and long workflows with controllable thinking effort.
Schematron V2 Small
inference-net/schematron-v2-smallInference.net's 3B-parameter HTML-to-JSON extraction model, focused on accuracy for complex schemas and long web pages. It turns HTML into typed, structured data for web scraping and product catalog ingestion, with a 128K-token context window. Supply HTML in the user message and extraction instructions in a JSON schema via response_format; it does not follow ordinary chat or system prompts.
Schematron V2 Turbo
inference-net/schematron-v2-turboInference.net's 3B-parameter HTML-to-JSON extraction model, optimized for throughput and low cost on high-volume workloads. It turns HTML into typed, structured data for web scraping and product catalog ingestion, with a 128K-token context window. Supply HTML in the user message and extraction instructions in a JSON schema via response_format; it does not follow ordinary chat or system prompts.
DeepSeek V4.1 Flash TEE
TEE/deepseek-v4.1-flashDeepSeek V4.1 Flash supports text and image input, reasoning, tool calling, and structured output with a 1M-token context window. This route runs through Tinfoil attested inference inside a Trusted Execution Environment.