Z.AI

GLM 5.2

z-ai/glm-5.2Operational

Z.ai GLM-5.2: flagship coding model for long-horizon tasks with a usable 1M context and 128K output; thinking with effort high / max.

Input
$1.40
per 1M tokens
Output
$4.40
per 1M tokens
Cache read
$0.260
per 1M tokens
Context
1.0M
tokens
Max output
131K
tokens
Official pricing

USD per million tokens, exactly what Z.AI publishes. ElevenRouter adds nothing on tokens.

Input$1.40
Output$4.40
Cache read$0.260
Cache write$1.40
Example · 1,000 in + 500 out$0.00360
Capabilities
  • Tool calling
  • Structured outputs
  • JSON mode
  • Reasoning
  • Vision
  • Audio input
  • Prompt caching
  • Streaming
Modalities
text->text
Tokenizer
GLM
Released
2026-09-25
Family
glm
Availability

Measured by the routing engine; up when at least one credential can serve the model.

Last 3 days
99.82%
Last 30 days
99.82%
30 days agotoday
Full status page
Observed performance

Real requests over the last 7 days.

TTFT p50
—
TTFT p95
—
Tokens / s p50
—
Tokens / s p95
—

Percentiles appear after five successful requests.

Usage on ElevenRouter

Tokens per day across all customers, anonymous and aggregate.

0 tokens · 30d
0 requests in the last 7 days
Use it

Any OpenAI or Anthropic SDK works by changing the base URL. Supported parameters: max_tokens, temperature, top_p, top_k, frequency_penalty, presence_penalty, repetition_penalty, seed, stop, tools, tool_choice, response_format, reasoning, include_reasoning.

curl https://elevenrouter.com/api/v1/chat/completions \
  -H "Authorization: Bearer $ELEVENROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{ "model": "z-ai/glm-5.2", "messages": [{ "role": "user", "content": "Hello" }] }'
Related models

Questions

How much does GLM 5.2 cost on ElevenRouter?
GLM 5.2 is billed at Z.AI's official list price: $1.40 per million input tokens and $4.40 per million output tokens, with cached input at $0.260. A request with 1,000 input and 500 output tokens costs about $0.00360.
What is the context length of GLM 5.2?
GLM 5.2 accepts up to 1.0M tokens of context and can produce up to 131K output tokens.
Does GLM 5.2 support tool calling and structured outputs?
Yes, tool calling is supported and JSON mode is available; add the response-healing plugin for schema validation. Supported parameters: max_tokens, temperature, top_p, top_k, frequency_penalty, presence_penalty, repetition_penalty, seed, stop, tools, tool_choice, response_format, reasoning, include_reasoning.
How do I call GLM 5.2?
Send an OpenAI-compatible chat completion to https://elevenrouter.com/api/v1/chat/completions with "model": "z-ai/glm-5.2" and your ElevenRouter key, or use the Anthropic Messages endpoint. Any OpenAI or Anthropic SDK works by changing the base URL.
Is GLM 5.2 available right now?
Availability is measured continuously by the routing engine. Current status: operational, 99.82% uptime over the last 30 days. When one credential fails, requests fail over to another automatically.