Browse all Inclusionai text models
Provider logo

Ling 3.0 Flash VL

inclusionai/ling-3.0-flash-vl
Back
Provider logo

Ling 3.0 Flash VL

inclusionai/ling-3.0-flash-vl
Back

Ling 3.0 Flash VL is inclusionAI's native multimodal Mixture-of-Experts model with 124B total parameters and 5.5B active parameters per token. It combines image and video understanding with reasoning and tool use for document analysis, charts, visual verification, and interface-based agent tasks. Thinking is enabled by default and can be turned off in settings.

Added Sep 9, 2026

Model weights

Context Window

262.1K

Max Output

32.8K

Avg output tokens (7d)

1.7K tokens

89%

Input Price (Auto)

$0.060/1M

Output Price (Auto)

$0.18/1M

Cache Read (Auto)

$0.012/1M

Capabilities

Benchmarks

Performance metrics and benchmarks

Sourced from Artificial Analysis.

Intelligence Index

24.6

Better than 74% of models compared

Coding Index

57.0

Better than 70% of models compared

Agentic Index

28.7

Better than 71% of models compared

Agentic work

AutomationBench-AA

Workflow automation with guardrail penalties

15.7%

Better than 38% of models compared

AA-Briefcase

Agentic knowledge work (Elo)

984 Elo

Better than 48% of models compared

GDPval-AA v2

Economically valuable tasks (Elo)

1150 Elo

Better than 58% of models compared

Document reasoning

GDP.pdf

Professional PDF reasoning: all-pass rate

8.4%

Better than 34% of models compared

AA-LCR v1.1

Long context reasoning with updated grading

78.3%

Better than 81% of models compared

Reasoning

HLE

Humanity's Last Exam

22.0%

Better than 71% of models compared

CritPt

Research-level physics reasoning

2.0%

Coding

Terminal-Bench v4.0

Practical coding and terminal tasks

0.0%

Better than 16% of models compared

SciCode

Python programming for scientific computing

44.2%

Better than 37% of models compared

Knowledge

AA-Omniscience Accuracy

Proportion of correctly answered questions

14.3%

AA-Omniscience Hallucination Rate

Rate of incorrect answers among non-correct responses

22.0%

Legacy benchmarks

GPQA Diamond (legacy)

Graduate-level scientific reasoning

86.2%

Better than 81% of models compared

AA-LCR (unversioned / legacy)

Long context reasoning evaluation

78.3%

Better than 81% of models compared

GDPval-AA (unversioned / legacy)

Economically valuable tasks

32.5%

Last updated Sep 29, 2026

Artificial Analysis

Providers

Provider information for this model’s automatic routing. These routes cannot be selected individually.

Loading provider options…

Compare Ling 3.0 Flash VL with similar models from the same provider or model family.

Ling 3.0 Flash

inclusionai/ling-3.0-flash

Ling-3.0-flash is a 124B-parameter Mixture-of-Experts model with approximately 5.1B parameters active per token. It prioritizes token efficiency and production-scale agentic inference, helping coding and tool-using agents complete more work within constrained latency and serving budgets.

Ling 3.0 Flash Thinking

inclusionai/ling-3.0-flash:thinking

Ling-3.0-flash Thinking enables visible reasoning on inclusionAI's token-efficient 124B-parameter Mixture-of-Experts model for harder coding, tool use, planning, and production-scale agent workflows.

MiMo V2.6 Flash Uncensored Thinking

xiaomi/mimo-v2.6-flash-uncensored:thinking

MiMo V2.6 Flash Uncensored with maximum thinking enabled. Supports a 1M-token context window and tool calling.

MiMo V2.6 Flash Abliterated

xiaomi/mimo-v2.6-flash-abliterated

A MiMo V2.6 Flash finetune with refusal-direction ablation, image input, a 1M-token context window, separate reasoning output, and tool calling.

MiMo V2.6 Flash Uncensored

xiaomi/mimo-v2.6-flash-uncensored

A lower-refusal MiMo V2.6 Flash finetune with image input, a 1M-token context window, separate reasoning output, and tool calling.

MiMo V2.6 Flash

xiaomi/mimo-v2.6-flash

MiMo V2.6 Flash is Xiaomi's native omnimodal 309B-parameter mixture-of-experts model, activating 15B parameters per token. It balances intelligence, efficiency, and cost for coding, general agents, visual tasks, and cybersecurity, with text, image, video, and audio understanding and a 1M-token context window.