Agnes 3.0 Flash is a low-cost model for coding, tool use, and multi-turn agent tasks. It supports text and image input, optional thinking, and a 512K-token context window.
Added Sep 9, 2026
Context Window
524.3K
Max Output
65.5K
Avg output tokens (7d)
1.1K tokens
Input Price (Auto)
$0.050/1M
Output Price (Auto)
$0.15/1M
Cache Read (Auto)
$0.0050/1M
Capabilities
Benchmarks
Benchmarks
Performance metrics and benchmarks
No benchmark data is available yet for this model.
Providers
Auto routing is available for this model. Explicit provider selection is not available.
Loading provider options…
Related text models
Compare Agnes 3.0 Flash with similar models from the same provider or model family.
MiMo V2.6 Flash Uncensored Thinking
xiaomi/mimo-v2.6-flash-uncensored:thinkingMiMo V2.6 Flash Uncensored with maximum thinking enabled. Supports a 1M-token context window and tool calling.
MiMo V2.6 Flash Abliterated
xiaomi/mimo-v2.6-flash-abliteratedA MiMo V2.6 Flash finetune with refusal-direction ablation, image input, a 1M-token context window, separate reasoning output, and tool calling.
MiMo V2.6 Flash Uncensored
xiaomi/mimo-v2.6-flash-uncensoredA lower-refusal MiMo V2.6 Flash finetune with image input, a 1M-token context window, separate reasoning output, and tool calling.
MiMo V2.6 Flash
xiaomi/mimo-v2.6-flashMiMo V2.6 Flash is Xiaomi's native omnimodal 309B-parameter mixture-of-experts model, activating 15B parameters per token. It balances intelligence, efficiency, and cost for coding, general agents, visual tasks, and cybersecurity, with text, image, video, and audio understanding and a 1M-token context window.
GLM 5.3 Flash Cybersecurity
z-ai/glm-5.3-flash-cybersecurityGLM 5.3 Flash Cybersecurity is a cybersecurity-focused variant based on the uncensored model, with provider moderation for illegal activities. It supports always-on reasoning, image understanding, tool calling, and a 1,048,576-token context window.
Qwen3.8 Omni Flash
qwen/qwen3.8-omni-flashQwen3.8 Omni Flash is a fast omni-modal reasoning model for understanding text, images, audio, and video. It is especially suited to meeting summaries, transcripts and subtitles, speaker-aware audiovisual analysis, and long-form content review. It returns text and supports a nearly one-million-token context window, tool calling, and structured output.