Z.ai's native multimodal agent model for vision-based coding and agent workflows. This is the standard non-thinking variant for image, video, and text inputs, tuned for perceive-plan-execute loops, complex coding, and tool-driven task execution.
Added Apr 1, 2026
Context Window
202.8K
Max Output
131.1K
Avg output tokens (7d)
575 tokens
Input Price (Auto)
$1.20/1M
Output Price (Auto)
$4.00/1M
Cache Read (Auto)
$0.24/1M
Capabilities
Benchmarks
Benchmarks
Performance metrics and benchmarks
Sourced from LMArena.
Arena Score
1432.7
Overall Rank
#97 / 402
Votes
9,361
Confidence Interval
1426.0 - 1439.4
Category Scores
Coding
#85 / 397
2,612 votes
1489.1
Math
#76 / 384
442 votes
1441.3
Longer Query
#95 / 380
4,205 votes
1442.6
Creative Writing
#100 / 400
1,653 votes
1400.6
Instruction Following
#95 / 402
3,244 votes
1422.2
Hard Prompts
#99 / 402
6,180 votes
1452.1
Additional Categories20
Chinese
#79 / 373
602 votes
1478.5
English
#79 / 402
3,990 votes
1451.7
Spanish
#79 / 283
290 votes
1435.7
Hard Prompts English
#82 / 400
2,641 votes
1466.9
Korean
#85 / 269
209 votes
1383.5
Expert
#90 / 352
1,033 votes
1461.6
Polish
#90 / 222
187 votes
1427.1
French
#92 / 281
394 votes
1449.8
Industry Software And It Services
#92 / 402
3,724 votes
1472.2
Industry Writing And Literature And Language
#93 / 401
2,388 votes
1411.5
Industry Business And Management And Financial Operations
#95 / 395
1,947 votes
1438.8
Industry Mathematical
#95 / 379
518 votes
1438.0
Exclude Ties
#98 / 402
6,894 votes
1426.1
Non English
#100 / 402
5,369 votes
1414.0
Industry Life And Physical And Social Science
#103 / 400
1,459 votes
1448.2
Multi Turn
#103 / 400
1,557 votes
1432.7
Industry Entertainment And Sports And Media
#104 / 400
2,196 votes
1396.0
Russian
#106 / 366
944 votes
1421.1
Industry Legal And Government
#127 / 375
731 votes
1423.7
Industry Medicine And Healthcare
#145 / 371
646 votes
1427.7
Published 2026-09-13 · Matched as glm-5v-turbo
LMArena DatasetProviders
Auto routing is available for this model. Explicit provider selection is not available.
Loading provider options…
Related text models
Compare GLM 5V Turbo with similar models from the same provider or model family.
GLM 5V Turbo Thinking
z-ai/glm-5v-turbo:thinkingThinking-enabled GLM 5V Turbo for image, video, and text inputs. Uses the same multimodal foundation model with more deliberate vision-grounded analysis, planning, and tool use.
GLM 5 Turbo
z-ai/glm-5-turboFast GLM 5 Turbo variant from Z-AI for general chat, coding, and tool use.
GLM 4.6 Turbo
z-ai/GLM-4.6-turboFast variant of GLM 4.6 for general chat, coding, and analysis with improved latency and strong reasoning.
GLM 4.6 Turbo (Thinking)
z-ai/GLM-4.6-turbo:thinkingGLM 4.6 Turbo with thinking mode enabled for enhanced reasoning; shows internal reasoning and supports long context.
GLM 5.3 Flash
z-ai/glm-5.3-flashox-alpha out of stealth! GLM-5.3 Flash is Z.ai's first natively multimodal GLM-5 model, with 320B total parameters and just 18B active parameters for efficient coding, agentic work, and precise 1M-token context. Its hybrid sparse-and-linear attention architecture helps it outperform GLM-5.2 at one-tenth the price while approaching Claude Opus 4.8 on coding and agentic benchmarks.
GLM 5.3
z-ai/glm-5.3GLM-5.3 for long-horizon autonomous coding and engineering workflows. This variant defaults to the model's low reasoning tier for faster responses.