Provider logo

GLM 5.3 Flash

z-ai/glm-5.3-flash
Provider logo

GLM 5.3 Flash

z-ai/glm-5.3-flash

ox-alpha out of stealth! GLM-5.3 Flash is Z.ai's first natively multimodal GLM-5 model, with 320B total parameters and just 18B active parameters for efficient coding, agentic work, and precise 1M-token context. Its hybrid sparse-and-linear attention architecture helps it outperform GLM-5.2 at one-tenth the price while approaching Claude Opus 4.8 on coding and agentic benchmarks.

Added Aug 26, 2026

Model weights

Context Window

1.0M

Max Output

131.1K

Input Price (Auto)

$0.075/1M

Output Price (Auto)

$0.25/1M

Cache Read (Auto)

$0.015/1M

Capabilities

Benchmarks

Performance metrics and benchmarks

No benchmark data is available yet for this model.

Providers

Auto routing is available for this model. Explicit provider selection is not available.

Loading provider options…

Compare GLM 5.3 Flash with similar models from the same provider or model family.