Vision-Language Model with thinking paradigm and reinforcement learning. Achieves state-of-the-art performance among 10B-parameter VLMs. Supports 64k context length, handles arbitrary aspect ratios and up to 4K image resolution. Bilingual Chinese/English.
Added Jul 9, 2025
Context Window
64.0K
Max Output
8.2K
Input Price (Auto)
$0.30/1M
Output Price (Auto)
$0.30/1M
Cache Read (Auto)
$0.15/1M
Benchmarks
Benchmarks
Performance metrics and benchmarks
No benchmark data is available yet for this model.
Providers
Auto routing is available for this model. Explicit provider selection is not available.
Loading provider options…
Related text models
Compare GLM 4.1V Thinking Flash with similar models from the same provider or model family.
GLM 4.7 Flash
zai-org/glm-4.7-flashGLM-4.7-Flash is a lightweight 30B model optimized for coding and agentic tasks. Balances high performance with efficiency.
GLM 4.7 Flash Original
zai-org/glm-4.7-flash-originalGLM-4.7-Flash is a lightweight 30B model optimized for coding and agentic tasks. Balances high performance with efficiency, perfect for local deployment. Routed directly via Z-AI (Zhipu) subscription.
GLM 4.7 Flash Original Thinking
zai-org/glm-4.7-flash-original:thinkingGLM-4.7-Flash with extended thinking capabilities for complex reasoning. Lightweight 30B model optimized for coding and agentic tasks.
GLM 4.7 Flash Thinking
zai-org/glm-4.7-flash:thinkingGLM-4.7-Flash with extended thinking capabilities for complex reasoning. Lightweight 30B model optimized for coding and agentic tasks.
GLM 4.6V Flash
zai-org/glm-4.6v-flash-originalGLM-4.6V-Flash (9B), a lightweight model optimized for local deployment and low-latency applications. Scales context window to 128k tokens and achieves SoTA performance in visual understanding among similar-scale models.
GLM-4 Flash
glm-4-flashExtremely cheap model with 128K context window