GLM-5.3-FlashX is the high-speed variant of Z.ai's GLM-5.3-Flash, a native multimodal model delivering inference speeds of up to 200 tokens/s. Built on the same hybrid sparse and linear attention architecture...
Z.ai: GLM 5.3 FlashX — Z.ai. Context window 1,048,576 tokens. API pricing (per 1M tokens): Input $0.37, Output $1.25. License: Proprietary.
| Provider | Z.ai |
|---|---|
| API model ID | z-ai/glm-5.3-flashx |
| Tier | Frontier |
| Context window | 1,048,576 tokens |
| Max output | 131,072 tokens |
| License | Proprietary |
| API pricing (per 1M tokens) · Input | $0.37 |
| API pricing (per 1M tokens) · Output | $1.25 |