Private AI
Ling-3.0-flash is a 124B-parameter Mixture-of-Experts model with approximately 5.1B parameters active per token. It prioritizes token efficiency and production-scale agentic inference, helping coding and tool-using agents complete more work within constrained latency and serving budgets.
Added Jul 23, 2026
Context Window
262.1K
Max Output
32.8K
Input Price (Auto)
$0.060/1M
Output Price (Auto)
$0.18/1M
Cache Read (Auto)
$0.012/1M
Capabilities
Performance metrics and benchmarks
No benchmark data is available yet for this model.
Auto routing is available for this model. Explicit provider selection is not available.
Loading provider options…