Provider logo

GLM 4.6V

z-ai/glm-4.6v
Provider logo

GLM 4.6V

z-ai/glm-4.6v

GLM-4.6V scales its context window to 128k tokens in training, and achieves SoTA performance in visual understanding among models of similar parameter scales. Integrates native Function Calling capabilities, bridging 'visual perception' and 'executable action' for multimodal agents. Quantized at FP8.

Added Dec 11, 2025

Model weights

Context Window

128.0K

Max Output

24.0K

Input Price (Auto)

$0.30/1M

Output Price (Auto)

$0.90/1M

Cache Read (Auto)

$0.15/1M

Capabilities

Benchmarks

Performance metrics and benchmarks

Sourced from Artificial Analysis.

Intelligence Index

10.9

Better than 34% of models compared

Reasoning

GPQA Diamond

Graduate-level scientific reasoning

56.6%

Better than 32% of models compared

HLE

Humanity's Last Exam

3.7%

Better than 9% of models compared

IFBench

Instruction-following benchmark

27.9%

Better than 11% of models compared

T²-Bench Telecom

Conversational AI agents in dual-control scenarios

30.7%

Better than 38% of models compared

AA-LCR

Long context reasoning evaluation

14.3%

Better than 22% of models compared

CritPt

Research-level physics reasoning

0.0%

Coding

SciCode

Python programming for scientific computing

27.2%

Better than 34% of models compared

Terminal-Bench Hard

Agentic coding and terminal use

3.0%

Better than 23% of models compared

LiveCodeBench

Contamination-free coding benchmark

41.1%

Better than 49% of models compared

Math

AIME 2025

American Invitational Mathematics Examination 2025

26.3%

Better than 28% of models compared

Knowledge

MMLU-Pro

Professional and academic subject knowledge

75.2%

Better than 49% of models compared

AA-Omniscience Accuracy

Proportion of correctly answered questions

17.4%

AA-Omniscience Hallucination Rate

Rate of incorrect answers among non-correct responses

66.6%

Last updated Aug 16, 2026

Artificial Analysis

Providers

Auto routing is available for this model. Explicit provider selection is not available.

Loading provider options…

Compare GLM 4.6V with similar models from the same provider or model family.

GLM 5.3 Flash

z-ai/glm-5.3-flash

ox-alpha out of stealth! GLM-5.3 Flash is Z.ai's first natively multimodal GLM-5 model, with 320B total parameters and just 18B active parameters for efficient coding, agentic work, and precise 1M-token context. Its hybrid sparse-and-linear attention architecture helps it outperform GLM-5.2 at one-tenth the price while approaching Claude Opus 4.8 on coding and agentic benchmarks.

GLM 5.3

z-ai/glm-5.3

GLM-5.3 for long-horizon autonomous coding and engineering workflows. This variant defaults to the model's low reasoning tier for faster responses.

GLM 5.3 Thinking

z-ai/glm-5.3:thinking

GLM-5.3 with higher reasoning enabled for harder long-horizon coding, autonomous agent workflows, and complex engineering tasks.

GLM 5.3 Flash Uncensored

z-ai/glm-5.3-flash-uncensored

GLM 5.3 Flash Uncensored is an FP8 uncensored fine-tune of the efficient 320B mixture-of-experts reasoning model with restored vision support, built for unrestricted chat, creative writing, coding, agentic work, tool use, and long-context tasks.

GLM 5.2

z-ai/glm-5.2

GLM-5.2 is Z.AI's flagship model for long-horizon autonomous coding and engineering workflows. It is built to plan, execute, iterate, and optimize complex development tasks over extended runs. This variant keeps thinking disabled for faster direct responses.

GLM 5.2 Thinking

z-ai/glm-5.2:thinking

GLM-5.2 with thinking enabled for harder long-horizon coding, autonomous agent workflows, complex engineering optimization, and real-world development tasks.