Z.ai's native multimodal agent model for vision-based coding and agent workflows. This is the standard non-thinking variant for image, video, and text inputs, tuned for perceive-plan-execute loops, complex coding, and tool-driven task execution.
Added Apr 1, 2026
Context Window
202.8K
Max Output
131.1K
Avg output tokens (7d)
476 tokens
Input Price (Auto)
$1.20/1M
Output Price (Auto)
$4.00/1M
Cache Read (Auto)
$0.24/1M
Capabilities
Benchmarks
Benchmarks
Performance metrics and benchmarks
Sourced from LMArena.
Arena Score
1432.7
Overall Rank
#96 / 401
Votes
9,367
Confidence Interval
1426.0 - 1439.4
Category Scores
Coding
#84 / 396
2,614 votes
1489.2
Math
#74 / 384
444 votes
1442.9
Longer Query
#94 / 379
4,207 votes
1442.5
Creative Writing
#98 / 399
1,651 votes
1400.7
Instruction Following
#94 / 401
3,243 votes
1422.0
Hard Prompts
#98 / 401
6,184 votes
1452.0
Additional Categories20
Spanish
#77 / 283
290 votes
1435.6
Chinese
#78 / 372
602 votes
1478.6
English
#78 / 401
3,992 votes
1451.8
Hard Prompts English
#82 / 399
2,642 votes
1466.7
Korean
#87 / 269
209 votes
1381.8
Expert
#88 / 350
1,033 votes
1461.5
Polish
#90 / 222
187 votes
1427.3
Industry Mathematical
#91 / 378
520 votes
1439.3
Industry Software And It Services
#92 / 401
3,728 votes
1472.1
Industry Writing And Literature And Language
#92 / 400
2,388 votes
1411.5
French
#93 / 281
394 votes
1449.7
Industry Business And Management And Financial Operations
#94 / 394
1,951 votes
1438.1
Exclude Ties
#97 / 401
6,897 votes
1426.2
Non English
#99 / 401
5,373 votes
1413.9
Multi Turn
#101 / 399
1,556 votes
1432.9
Industry Life And Physical And Social Science
#102 / 399
1,464 votes
1448.1
Industry Entertainment And Sports And Media
#104 / 399
2,199 votes
1395.6
Russian
#105 / 365
947 votes
1421.6
Industry Legal And Government
#125 / 373
734 votes
1423.2
Industry Medicine And Healthcare
#143 / 369
648 votes
1427.5
Published 2026-09-11 · Matched as glm-5v-turbo
LMArena DatasetProviders
Auto routing is available for this model. Explicit provider selection is not available.
Loading provider options…
Related text models
Compare GLM 5V Turbo with similar models from the same provider or model family.
GLM 5V Turbo Thinking
z-ai/glm-5v-turbo:thinkingThinking-enabled GLM 5V Turbo for image, video, and text inputs. Uses the same multimodal foundation model with more deliberate vision-grounded analysis, planning, and tool use.
GLM 5 Turbo
z-ai/glm-5-turboFast GLM 5 Turbo variant from Z-AI for general chat, coding, and tool use.
GLM 4.6 Turbo
z-ai/GLM-4.6-turboFast variant of GLM 4.6 for general chat, coding, and analysis with improved latency and strong reasoning.
GLM 4.6 Turbo (Thinking)
z-ai/GLM-4.6-turbo:thinkingGLM 4.6 Turbo with thinking mode enabled for enhanced reasoning; shows internal reasoning and supports long context.
GLM 5.3 Flash
z-ai/glm-5.3-flashox-alpha out of stealth! GLM-5.3 Flash is Z.ai's first natively multimodal GLM-5 model, with 320B total parameters and just 18B active parameters for efficient coding, agentic work, and precise 1M-token context. Its hybrid sparse-and-linear attention architecture helps it outperform GLM-5.2 at one-tenth the price while approaching Claude Opus 4.8 on coding and agentic benchmarks.
GLM 5.3
z-ai/glm-5.3GLM-5.3 for long-horizon autonomous coding and engineering workflows. This variant defaults to the model's low reasoning tier for faster responses.