Browse all Inclusionai text models
Provider logo

Ling 3.0 Flash VL

inclusionai/ling-3.0-flash-vl
Provider logo

Ling 3.0 Flash VL

inclusionai/ling-3.0-flash-vl

Ling 3.0 Flash VL is inclusionAI's native multimodal Mixture-of-Experts model with 124B total parameters and 5.5B active parameters per token. It combines image and video understanding with reasoning and tool use for document analysis, charts, visual verification, and interface-based agent tasks. Thinking is enabled by default and can be turned off in settings.

Added Sep 9, 2026

Context Window

262.1K

Max Output

32.8K

Input Price (Auto)

$0.060/1M

Output Price (Auto)

$0.18/1M

Cache Read (Auto)

$0.012/1M

Capabilities

Benchmarks

Performance metrics and benchmarks

No benchmark data is available yet for this model.

Providers

Provider information for this model’s automatic routing. These routes cannot be selected individually.

Loading provider options…

Compare Ling 3.0 Flash VL with similar models from the same provider or model family.

Ling 3.0 Flash

inclusionai/ling-3.0-flash

Ling-3.0-flash is a 124B-parameter Mixture-of-Experts model with approximately 5.1B parameters active per token. It prioritizes token efficiency and production-scale agentic inference, helping coding and tool-using agents complete more work within constrained latency and serving budgets.

Ling 3.0 Flash Thinking

inclusionai/ling-3.0-flash:thinking

Ling-3.0-flash Thinking enables visible reasoning on inclusionAI's token-efficient 124B-parameter Mixture-of-Experts model for harder coding, tool use, planning, and production-scale agent workflows.

Agnes 3.0 Flash

agnes-3.0-flash

Agnes 3.0 Flash is a low-cost model for coding, tool use, and multi-turn agent tasks. It supports text and image input, optional thinking, and a 512K-token context window.

DeepSeek V4.1 Flash

deepseek/deepseek-v4.1-flash

DeepSeek V4.1 Flash supports text and image input, reasoning, tool calling, and structured output with a 1M-token context window. This is a rate-limited beta with limited capacity, intended for testing rather than production use. Assume prompts and responses are logged by the provider and may be used for model training or service improvement. Do not send sensitive or confidential data.

DeepSeek V4.1 Flash Thinking

deepseek/deepseek-v4.1-flash:thinking

DeepSeek V4.1 Flash supports text and image input, reasoning, tool calling, and structured output with a 1M-token context window. This is a rate-limited beta with limited capacity, intended for testing rather than production use. Assume prompts and responses are logged by the provider and may be used for model training or service improvement. Do not send sensitive or confidential data.

DeepSeek V4 Flash Vision Exp Uncensored

deepseek/deepseek-v4-flash-vision-exp-uncensored

An uncensored variant of the experimental vision-enabled DeepSeek V4 Flash model for chat, image understanding, reasoning, coding, and tool use, with a 524K context window.