Ling-3.0-flash Thinking enables visible reasoning on inclusionAI's token-efficient 124B-parameter Mixture-of-Experts model for harder coding, tool use, planning, and production-scale agent workflows.
Added Jul 23, 2026
Model weightsContext Window
262.1K
Max Output
32.8K
Input Price (Auto)
$0.075/1M
Output Price (Auto)
$0.22/1M
Cache Read (Auto)
$0.015/1M
Capabilities
Benchmarks
Benchmarks
Performance metrics and benchmarks
Sourced from Artificial Analysis.
Intelligence Index
24.9
Coding Index
50.6
Agentic work
AutomationBench-AA
Workflow automation with guardrail penalties
3.2%
Better than 26% of models compared
AA-Briefcase
Agentic knowledge work (Elo)
796 Elo
Better than 37% of models compared
Document reasoning
GDP.pdf
Professional PDF reasoning: all-pass rate
5.4%
Better than 28% of models compared
AA-LCR v1.1
Long context reasoning with updated grading
73.0%
Better than 71% of models compared
Reasoning
HLE
Humanity's Last Exam
23.7%
Better than 74% of models compared
CritPt
Research-level physics reasoning
1.7%
Coding
Terminal-Bench v4.0
Practical coding and terminal tasks
0.0%
Better than 17% of models compared
SciCode
Python programming for scientific computing
42.0%
Better than 33% of models compared
Knowledge
AA-Omniscience Accuracy
Proportion of correctly answered questions
18.2%
AA-Omniscience Hallucination Rate
Rate of incorrect answers among non-correct responses
44.1%
Legacy benchmarks
GPQA Diamond (legacy)
Graduate-level scientific reasoning
85.5%
Better than 78% of models compared
AA-LCR (unversioned / legacy)
Long context reasoning evaluation
73.0%
Better than 71% of models compared
Last updated Sep 22, 2026
Artificial AnalysisProviders
Provider information for this model’s automatic routing. These routes cannot be selected individually.
Loading provider options…
Related text models
Compare Ling 3.0 Flash Thinking with similar models from the same provider or model family.
Ling 3.0 Flash VL
inclusionai/ling-3.0-flash-vlLing 3.0 Flash VL is inclusionAI's native multimodal Mixture-of-Experts model with 124B total parameters and 5.5B active parameters per token. It combines image and video understanding with reasoning and tool use for document analysis, charts, visual verification, and interface-based agent tasks. Thinking is enabled by default and can be turned off in settings.
Ling 3.0 Flash
inclusionai/ling-3.0-flashLing-3.0-flash is a 124B-parameter Mixture-of-Experts model with approximately 5.1B parameters active per token. It prioritizes token efficiency and production-scale agentic inference, helping coding and tool-using agents complete more work within constrained latency and serving budgets.
MiMo V2.6 Flash
xiaomi/mimo-v2.6-flashMiMo V2.6 Flash is Xiaomi's native omnimodal 309B-parameter mixture-of-experts model, activating 15B parameters per token. It balances intelligence, efficiency, and cost for coding, general agents, visual tasks, and cybersecurity, with text, image, video, and audio understanding and a 1M-token context window.
GLM 5.3 Flash Cybersecurity
z-ai/glm-5.3-flash-cybersecurityGLM 5.3 Flash Cybersecurity is a cybersecurity-focused variant based on the uncensored model, with provider moderation for illegal activities. It supports always-on reasoning, image understanding, tool calling, and a 1,048,576-token context window.
Qwen3.8 Omni Flash
qwen/qwen3.8-omni-flashQwen3.8 Omni Flash is a fast omni-modal reasoning model for understanding text, images, audio, and video. It is especially suited to meeting summaries, transcripts and subtitles, speaker-aware audiovisual analysis, and long-form content review. It returns text and supports a nearly one-million-token context window, tool calling, and structured output.
DeepSeek V4.1 Flash TEE
TEE/deepseek-v4.1-flashDeepSeek V4.1 Flash supports text and image input, reasoning, tool calling, and structured output with a 1M-token context window. This route runs through Tinfoil attested inference inside a Trusted Execution Environment.