Z.ai (GLM)

Z.ai (GLM) is GLM models with an unusually cheap flat-rate coding plan. We compare it in LLM APIs. It is 2nd of 8 on $/M output ($2.2 /M tok). The trade-off is being marked No on Zero retention, where Amazon Bedrock records Yes.

What is Z.ai (GLM)?

Z.ai (GLM) — GLM models with an unusually cheap flat-rate coding plan. Zhipu's GLM series, sold internationally as Z.ai, targets coding and agentic workloads directly and releases its flagship weights under MIT. Its most disruptive product is not the token price but the flat-rate coding subscription that undercuts per-token billing for anyone driving an agent hard all day. Widely mirrored by routers, which is how most Western teams reach it. Built by Zhipu AI. Founded 2019. Open source under MIT.

We track it in 1 comparison — LLM APIs — so every claim below is a cell in a table you can open and check rather than an impression. Across those rosters it sits against 17 other tools, and what follows is where it visibly separates from them.

Where it wins.

  • $/M output — $2.2 /M tok, 2nd of 8.
  • $/M input — $0.6 /M tok, 2nd of 8.

Each of those is ranked against the whole roster on its category page, not against a hand-picked subset, so a first place here means first of everything we list.

Where it gives ground.

  • Zero retention — No. Amazon Bedrock records Yes.
  • Context window — 200,000 tokens, 9th of 10. Anthropic records 1,000,000 tokens.

None of these disqualify it on their own. They are the fields to check against your own requirements before you commit, because they are the ones where a competitor genuinely does better.

Provenance. 20 of 27 tracked fields carry a value for Z.ai (GLM), and 1 of those cite a document you can open. Last verified 2026-01-15. Every figure keeps its own provenance — measured by us, claimed by the vendor, inferred, or community-reported — and we would rather print a dash than a guess.

Its nearest neighbour in our data is Qwen. Z.ai (GLM) is ahead on Open weights (Yes against Partial) and $/M output ($2.2 /M tok against $6 /M tok). Qwen takes MCP support (Partial against No) and AU region (Partial against No). That pattern repeats across the rest of the roster — see Z.ai (GLM) alternatives for the other rivals, each compared the same way.

At a glance
Company
Zhipu AI
Founded
2019
Licence
MIT (open source)
Fields we track
20 of 27
Last verified
2026-01-15

Z.ai (GLM) in LLM APIs

Ranked against 18 tools across 27 sourced fields. Open the full LLM APIs table.

The flat-rate coding plan is the genuinely disruptive product here — it makes running an agent hard all day a fixed cost rather than a variable one, which no per-token vendor can match. Model quality is a step below the frontier on hard reasoning; for routine coding turns most users do not notice.

Where it lands in this roster
$/M output
$2.2 /M tok2nd of 8
Inferredverified 2026-01-15
$/M input
$0.6 /M tok2nd of 8
Inferredverified 2026-01-15
Zero retention
No10th of 12
Inferredverified 2026-01-15
Context window
200,000 tokens9th of 10
Inferredverified 2026-01-15
Tool use
Yes1st of 18
Community-reportedverified 2026-01-15
Pinnable versions
Yes1st of 18
Inferredverified 2026-01-15
OpenAI-compat API
Yes1st of 18
Inferredverified 2026-01-15
MCP support
No6th of 18
Inferredverified 2026-01-15

When to use Z.ai (GLM)

Z.ai (GLM) is the right call in these situations, each one drawn from a field we actually record:

  • Driving a coding agent all day on a predictable monthly cost.
  • Cheap high-volume code completion and refactoring.
  • Self-hosting an MIT-licensed coding model.
  • $/M output is your binding constraint. Z.ai (GLM) records $2.2 /M tok, 2nd of 8 in the LLM APIs roster. We define that field as pay-as-you-go list price per million output tokens for the named flagship model, including reasoning or thinking tokens where those are billed as output.
  • $/M input is your binding constraint. Z.ai (GLM) records $0.6 /M tok, 2nd of 8 in the LLM APIs roster. We define that field as pay-as-you-go list price per million input tokens for the named flagship model, uncached, at the standard context tier.

When not to use Z.ai (GLM)

Reach for something else when any of the following is a requirement rather than a nice-to-have:

  • Zero retention. Z.ai (GLM) records No on the LLM APIs table. Amazon Bedrock records Yes on the same field. If that is a hard requirement rather than a preference, start elsewhere.
  • Context window. Z.ai (GLM) records 200,000 tokens, 9th of 10 in the LLM APIs roster. Anthropic records 1,000,000 tokens on the same field. If that is a hard requirement rather than a preference, start elsewhere.

We publish this block because a comparison that only lists what a tool is good at is marketing. Every figure above sits on the same page as its source, and the field definitions are on the category tables if you want to check how we measured them.

Tools compared alongside Z.ai (GLM)

Everything below shares at least one comparison category with Z.ai (GLM), ordered by how much overlap there is. For the reasoning on each — which fields it wins, which it loses — see Z.ai (GLM) alternatives.
Alibaba's model family — huge open-weight range, closed flagship
Wafer-scale inference — the fastest tokens per second available
Enterprise-focused models built for RAG and private deployment

Recent Z.ai (GLM) changes

We have not logged a dated change for Z.ai (GLM) yet. The timeline fills in as pricing moves, features ship and things get deprecated.

Sources and gaps

What we don't know. 7 of the 27 fields we track for Z.ai (GLM) are still blank: Max output, TTFT p50, Output tok/s, $/M cache read, Batch discount, Cache TTL and Notice period. Those render as dashes rather than as zeroes or assumptions, because an empty cell and a bad cell are not the same thing and only one of them is honest. If you know any of these figures and can point at a document, tell us.

Every figure on this page traces back to a document you can open. Where a vendor claims a number we could not reproduce, the cell says so.