GLM 5.3
Z.ai's flagship model for complex software engineering and long-horizon agent tasks (same base as GLM-5.2, improved post-training); 1M context, 128K output, thinking always on with effort low / high / max.
USD per million tokens, exactly what Z.AI publishes. ElevenRouter adds nothing on tokens.
| Input | $1.40 |
| Output | $4.40 |
| Cache read | $0.260 |
| Cache write | $1.40 |
| Example · 1,000 in + 500 out | $0.00360 |
- Tool calling
- Structured outputs
- JSON mode
- Reasoning
- Vision
- Audio input
- Prompt caching
- Streaming
- Modalities
- text->text
- Tokenizer
- GLM
- Released
- 2026-09-25
- Family
- glm
Measured by the routing engine; up when at least one credential can serve the model.
Real requests over the last 7 days.
- TTFT p50
- —
- TTFT p95
- —
- Tokens / s p50
- —
- Tokens / s p95
- —
Percentiles appear after five successful requests.
Tokens per day across all customers, anonymous and aggregate.
Any OpenAI or Anthropic SDK works by changing the base URL. Supported parameters: max_tokens, temperature, top_p, top_k, frequency_penalty, presence_penalty, repetition_penalty, seed, stop, tools, tool_choice, response_format, reasoning, include_reasoning.
curl https://elevenrouter.com/api/v1/chat/completions \
-H "Authorization: Bearer $ELEVENROUTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{ "model": "z-ai/glm-5.3", "messages": [{ "role": "user", "content": "Hello" }] }'- GLM 5.2Z.AI$1.40 / $4.40
- GLM 5.3 FlashZ.AI$0.150 / $0.500
- Claude Haiku 4.5Anthropic$1.00 / $5.00
- DeepSeek V4 ProDeepSeek$1.32 / $3.96
- Kimi K2.6Moonshot AI$0.950 / $4.00
- Qwen3.8 MaxQwen$2.00 / $6.00