Ling-2.6-flash is an instant (instruct) model from inclusionAI with 104B total parameters and 7.4B active parameters, designed for real-world agents that require fast responses, strong execution, and high token efficiency. It delivers performance comparable to state-of-the-art models at a similar scale while significantly reducing token usage across coding, document processing, and lightweight agent workflows.
Added Apr 21, 2026
Context Window
262.1K
Max Output
32.8K
Input Price (Auto)
$0.10/1M
Output Price (Auto)
$0.30/1M
Cache Read (Auto)
$0.020/1M
Capabilities
Benchmarks
Benchmarks
Performance metrics and benchmarks
Sourced from Artificial Analysis.
Intelligence Index
14.2
Coding Index
25.3
Agentic Index
2.3
Reasoning
GPQA Diamond
Graduate-level scientific reasoning
59.3%
Better than 36% of models compared
HLE
Humanity's Last Exam
6.3%
Better than 43% of models compared
IFBench
Instruction-following benchmark
57.4%
Better than 69% of models compared
T²-Bench Telecom
Conversational AI agents in dual-control scenarios
86.0%
Better than 78% of models compared
AA-LCR
Long context reasoning evaluation
28.0%
Better than 35% of models compared
GDPval-AA
Economically valuable tasks
2.2%
CritPt
Research-level physics reasoning
0.0%
Coding
SciCode
Python programming for scientific computing
27.1%
Better than 33% of models compared
Terminal-Bench Hard
Agentic coding and terminal use
21.2%
Better than 61% of models compared
Knowledge
AA-Omniscience Accuracy
Proportion of correctly answered questions
15.5%
AA-Omniscience Hallucination Rate
Rate of incorrect answers among non-correct responses
96.7%
Last updated Aug 16, 2026
Artificial AnalysisProviders
Auto routing is available for this model. Explicit provider selection is not available.
Loading provider options…
Related text models
Compare Ling 2.6 Flash with similar models from the same provider or model family.
Ling 3.0 Flash
inclusionai/ling-3.0-flashLing-3.0-flash is a 124B-parameter Mixture-of-Experts model with approximately 5.1B parameters active per token. It prioritizes token efficiency and production-scale agentic inference, helping coding and tool-using agents complete more work within constrained latency and serving budgets.
Ling 3.0 Flash Thinking
inclusionai/ling-3.0-flash:thinkingLing-3.0-flash Thinking enables visible reasoning on inclusionAI's token-efficient 124B-parameter Mixture-of-Experts model for harder coding, tool use, planning, and production-scale agent workflows.
Ling 2.6 1T
inclusionai/ling-2.6-1tLing-2.6-1T is an inclusionAI instruction model optimized for large-scale agentic and coding workloads with long-context support and structured output capabilities.
Ring 2.6 1T
inclusionai/ring-2.6-1tRing-2.6-1T is an inclusionAI thinking model for real-world agent workflows, coding agents, tool use, and long-horizon task execution.
Gemini 3.7 Flash
google/gemini-3.7-flashGoogle's frontier-performance Flash model for multimodal and agentic workloads, including coding, tool use, image understanding, PDF and document extraction, audio, and video. Google reports 65.3% on DeepSWE v1.1, up from 49.0% for Gemini 3.6 Flash, and 34% on GDP.pdf, up from 14%.
DeepSeek V4 Flash Latest
deepseek/deepseek-v4-flash-latestCompatibility alias that routes to the newest dated DeepSeek V4 Flash release. Currently uses Fireworks first for DeepSeek V4 Flash 0731, with automatic failover when needed. ⚠️ Privacy and logging guarantees are limited.