GLM 4.1V Thinking Flash

Vision-Language Model with thinking paradigm and reinforcement learning. Achieves state-of-the-art performance among 10B-parameter VLMs. Supports 64k context length, handles arbitrary aspect ratios and up to 4K image resolution. Bilingual Chinese/English.

Added Jul 9, 2025

Pricing

Auto routing · per 1M tokens
Input
$0.30
Output
$0.30
Compare provider prices

Specifications

Context window
64K
Max output
8.2K
Parameters
9B
Avg output (7d)
690 tokens
Longer than 55% of models

Benchmarks

No public benchmark scores for this model yet.

Providers

Auto routing is available for this model. Explicit provider selection is not available.

Loading provider options…