Browse all Thinkingmachines text models
Provider logo

Inkling Thinking

thinkingmachines/inkling:thinking
Provider logo

Inkling Thinking

thinkingmachines/inkling:thinking

The thinking version of Thinking Machines' 975B-parameter open-weights Mixture-of-Experts generalist with 41B active parameters. It reasons natively over text, images, and audio, and is built for agentic coding, tool use, detailed instruction following, long-context work, and controllable thinking effort.

Added Jul 15, 2026

Model weights

Context Window

1.0M

Max Output

32.8K

Avg output tokens (7d)

944 tokens

67%

Input Price (Auto)

$1.00/1M

Output Price (Auto)

$4.05/1M

Cache Read (Auto)

$0.17/1M

Capabilities

Benchmarks

Performance metrics and benchmarks

Sourced from Artificial Analysis.

Intelligence Index

25.5

Better than 79% of models compared

Coding Index

52.1

Better than 62% of models compared

Agentic Index

24.3

Better than 65% of models compared

Agentic work

AutomationBench-AA

Workflow automation with guardrail penalties

5.0%

Better than 33% of models compared

AutomationBench-AA Tasks Completed

Fully completed workflows without guardrail violations

0.3%

Better than 7% of models compared

AA-Briefcase

Agentic knowledge work (Elo)

836 Elo

Better than 43% of models compared

GDPval-AA v2

Economically valuable tasks (Elo)

1165 Elo

Better than 60% of models compared

Document reasoning

GDP.pdf

Professional PDF reasoning: all-pass rate

12.8%

Better than 58% of models compared

AA-LCR v1.1

Long context reasoning with updated grading

77.3%

Better than 81% of models compared

Reasoning

HLE

Humanity's Last Exam

31.9%

Better than 83% of models compared

CritPt

Research-level physics reasoning

5.4%

Coding

Terminal-Bench v4.0

Practical coding and terminal tasks

1.0%

Better than 47% of models compared

SciCode

Python programming for scientific computing

47.0%

Better than 50% of models compared

Knowledge

AA-Omniscience Accuracy

Proportion of correctly answered questions

41.5%

AA-Omniscience Hallucination Rate

Rate of incorrect answers among non-correct responses

67.7%

Legacy benchmarks

GPQA Diamond (legacy)

Graduate-level scientific reasoning

87.2%

Better than 83% of models compared

AA-LCR (unversioned / legacy)

Long context reasoning evaluation

77.3%

Better than 81% of models compared

GDPval-AA (unversioned / legacy)

Economically valuable tasks

33.3%

Last updated Sep 14, 2026

Artificial Analysis

Providers

Choose explicit providers for this model. Auto routing remains available as the default option.

Loading provider options…

Compare Inkling Thinking with similar models from the same provider or model family.

Inkling

thinkingmachines/inkling

The non-thinking version of Thinking Machines' 975B-parameter open-weights Mixture-of-Experts generalist with 41B active parameters. It gives faster direct answers across text, images, and audio, and is built for agentic coding, tool use, detailed instruction following, and long-context work.

Inkling Small

thinkingmachines/Inkling-Small

The direct-answer version of Thinking Machines' 276B-parameter open-weights multimodal Mixture-of-Experts model with 12B active parameters. It accepts text and images, and is designed for coding, tool use, instruction following, and general conversational work.

Inkling Small Thinking

thinkingmachines/Inkling-Small:thinking

The reasoning version of Thinking Machines' 276B-parameter open-weights multimodal Mixture-of-Experts model with 12B active parameters. It reasons over text and images, and is designed for agentic coding, tool use, instruction following, and long workflows with controllable thinking effort.

Schematron V2 Small

inference-net/schematron-v2-small

Inference.net's 3B-parameter HTML-to-JSON extraction model, focused on accuracy for complex schemas and long web pages. It turns HTML into typed, structured data for web scraping and product catalog ingestion, with a 128K-token context window. Supply HTML in the user message and extraction instructions in a JSON schema via response_format; it does not follow ordinary chat or system prompts.

Schematron V2 Turbo

inference-net/schematron-v2-turbo

Inference.net's 3B-parameter HTML-to-JSON extraction model, optimized for throughput and low cost on high-volume workloads. It turns HTML into typed, structured data for web scraping and product catalog ingestion, with a 128K-token context window. Supply HTML in the user message and extraction instructions in a JSON schema via response_format; it does not follow ordinary chat or system prompts.

DeepSeek V4.1 Flash TEE

TEE/deepseek-v4.1-flash

DeepSeek V4.1 Flash supports text and image input, reasoning, tool calling, and structured output with a 1M-token context window. This route runs through Tinfoil attested inference inside a Trusted Execution Environment.