The direct-answer version of Thinking Machines' 276B-parameter open-weights multimodal Mixture-of-Experts model with 12B active parameters. It accepts text and images, and is designed for coding, tool use, instruction following, and general conversational work.
Added Jul 31, 2026
Model weightsContext Window
524.3K
Max Output
32.8K
Avg output tokens (7d)
653 tokens
Input Price (Auto)
$0.45/1M
Output Price (Auto)
$1.20/1M
Cache Read (Auto)
$0.100/1M
Capabilities
Benchmarks
Benchmarks
Performance metrics and benchmarks
Sourced from Artificial Analysis.
Intelligence Index
27.8
Coding Index
52.9
Agentic work
AutomationBench-AA
Workflow automation with guardrail penalties
4.8%
Better than 27% of models compared
AA-Briefcase
Agentic knowledge work (Elo)
912 Elo
Better than 42% of models compared
Document reasoning
GDP.pdf
Professional PDF reasoning: all-pass rate
11.2%
Better than 44% of models compared
AA-LCR v1.1
Long context reasoning with updated grading
75.7%
Better than 75% of models compared
Reasoning
HLE
Humanity's Last Exam
33.3%
Better than 81% of models compared
CritPt
Research-level physics reasoning
8.3%
Coding
Terminal-Bench v4.0
Practical coding and terminal tasks
1.0%
Better than 42% of models compared
SciCode
Python programming for scientific computing
49.7%
Better than 52% of models compared
Knowledge
AA-Omniscience Accuracy
Proportion of correctly answered questions
33.2%
AA-Omniscience Hallucination Rate
Rate of incorrect answers among non-correct responses
63.0%
Legacy benchmarks
GPQA Diamond (legacy)
Graduate-level scientific reasoning
89.5%
Better than 88% of models compared
AA-LCR (unversioned / legacy)
Long context reasoning evaluation
75.7%
Better than 75% of models compared
Last updated Sep 29, 2026
Artificial AnalysisProviders
Choose explicit providers for this model. Auto routing remains available as the default option.
Loading provider options…
Related text models
Compare Inkling Small with similar models from the same provider or model family.
Inkling Small Thinking
thinkingmachines/Inkling-Small:thinkingThe reasoning version of Thinking Machines' 276B-parameter open-weights multimodal Mixture-of-Experts model with 12B active parameters. It reasons over text and images, and is designed for agentic coding, tool use, instruction following, and long workflows with controllable thinking effort.
Inkling
thinkingmachines/inklingThe non-thinking version of Thinking Machines' 975B-parameter open-weights Mixture-of-Experts generalist with 41B active parameters. It gives faster direct answers across text, images, and audio, and is built for agentic coding, tool use, detailed instruction following, and long-context work.
Inkling Thinking
thinkingmachines/inkling:thinkingThe thinking version of Thinking Machines' 975B-parameter open-weights Mixture-of-Experts generalist with 41B active parameters. It reasons natively over text, images, and audio, and is built for agentic coding, tool use, detailed instruction following, long-context work, and controllable thinking effort.
Schematron V2 Small
inference-net/schematron-v2-smallInference.net's 3B-parameter HTML-to-JSON extraction model, focused on accuracy for complex schemas and long web pages. It turns HTML into typed, structured data for web scraping and product catalog ingestion, with a 128K-token context window. Supply HTML in the user message and extraction instructions in a JSON schema via response_format; it does not follow ordinary chat or system prompts.
Mistral Small 3 24B (2501)
mistralai/mistral-small-24b-instruct-2501Mistral Small 3 24B (2501) hosted by IONOS in Berlin, Germany. Zero data retention.
Mistral Small 4 119B Thinking
mistralai/mistral-small-4-119b-2603:thinkingMistral Small 4 with reasoning enabled (reasoning_effort=high). A hybrid MoE model with deep step-by-step reasoning for complex prompts, coding, and multi-step problem solving.