---
title: Z.ai (GLM)
slug: zhipu-zai
url: "https://toolweight.com/options/zhipu-zai"
homepage: "https://z.ai"
categories: llm-apis
last_verified: 2026-01-15
license: CC-BY-4.0
---

# Z.ai (GLM)

> GLM models with an unusually cheap flat-rate coding plan

Zhipu's GLM series, sold internationally as Z.ai, targets coding and agentic workloads directly and releases its flagship weights under MIT. Its most disruptive product is not the token price but the flat-rate coding subscription that undercuts per-token billing for anyone driving an agent hard all day. Widely mirrored by routers, which is how most Western teams reach it.

## Identity

|  |  |
| --- | --- |
| Name | Z.ai (GLM) |
| Company | Zhipu AI |
| One-liner | GLM models with an unusually cheap flat-rate coding plan |
| Site | https://z.ai |
| Docs | https://docs.z.ai |
| Founded | 2019 |
| Open source | Yes |
| Licence | MIT |
| Compared in | 1 |

## Where it is compared

### [LLM APIs](https://toolweight.com/compare/llm-apis)
Ranked **#14 of 18** on default weights.
| Field | Value | Confidence | Verified | Source | Note |
| --- | --- | --- | --- | --- | --- |
| Flagship model | GLM-4.6 | Inferred | 2026-01-15 | - | A newer GLM generation may have shipped since; verify before quoting. |
| Context window | 200,000 tokens | Inferred | 2026-01-15 | - | - |
| Max output | - | Unknown | - | - | - |
| Image input | ◐ | Inferred | 2026-01-15 | - | Via the separate GLM-V line. |
| Audio in/out | ○ | Inferred | 2026-01-15 | - | - |
| Open weights | ● | Vendor-claimed | 2026-01-15 | https://huggingface.co/zai-org/GLM-4.6 | GLM-4.6 weights released under MIT. |
| OpenAI-compat API | ● | Inferred | 2026-01-15 | - | - |
| TTFT p50 | - | Unknown | - | - | - |
| Output tok/s | - | Unknown | - | - | - |
| $/M input | $0.6 /M tok | Inferred | 2026-01-15 | - | - |
| $/M output | $2.2 /M tok | Inferred | 2026-01-15 | - | Token pricing is beside the point for heavy users, the flat-rate coding subscription is roughly the price of a coffee a month and covers a very large quota. |
| $/M cache read | - | Unknown | - | - | - |
| Batch discount | - | Unknown | - | - | - |
| Tool use | ● | Community-reported | 2026-01-15 | - | - |
| Schema output | ◐ | Inferred | 2026-01-15 | - | - |
| Effort control | ◐ | Inferred | 2026-01-15 | - | - |
| Computer use | ○ | Inferred | 2026-01-15 | - | - |
| MCP support | ○ | Inferred | 2026-01-15 | - | - |
| Cache TTL | - | Unknown | - | - | - |
| Continuity policy | Versioned model names (GLM-4.5, 4.6) stay addressable after a successor ships, but there is no published deprecation policy or notice period. MIT-licensed weights are the fallback: every flagship generation so far has been released publicly, so a retired endpoint does not strand the model. | Inferred | 2026-01-15 | - | - |
| Notice period | - | Unknown | - | - | - |
| Pinnable versions | ● | Inferred | 2026-01-15 | - | - |
| Zero retention | ○ | Inferred | 2026-01-15 | - | - |
| Trains on your data | ◐ | Inferred | 2026-01-15 | - | - |
| AU region | ○ | Inferred | 2026-01-15 | - | - |
| API since | 2023 | Inferred | - | - | - |
| Positioning | Coding-focused GLM models with a flat-rate subscription | Inferred | - | - | - |

**Verdict.** The flat-rate coding plan is the genuinely disruptive product here, it makes running an agent hard all day a fixed cost rather than a variable one, which no per-token vendor can match. Model quality is a step below the frontier on hard reasoning; for routine coding turns most users do not notice.

## Alternatives

- [Qwen](https://toolweight.com/options/alibaba-qwen), Alibaba's model family, huge open-weight range, closed flagship
- [Amazon Bedrock](https://toolweight.com/options/amazon-bedrock), Multi-vendor model access inside your existing AWS account
- [Anthropic](https://toolweight.com/options/anthropic), Claude models, built around long agentic runs and tool use
- [Cerebras](https://toolweight.com/options/cerebras), Wafer-scale inference, the fastest tokens per second available
- [Cohere](https://toolweight.com/options/cohere), Enterprise-focused models built for RAG and private deployment
- [DeepSeek](https://toolweight.com/options/deepseek), Frontier-adjacent models at a small fraction of Western prices
- [Fireworks AI](https://toolweight.com/options/fireworks-ai), Fast open-weight inference with strong structured-output support
- [Google Gemini](https://toolweight.com/options/google-gemini), Gemini via AI Studio for prototyping or Vertex AI for production
- [Groq](https://toolweight.com/options/groq), Custom LPU silicon serving open-weight models at extreme speed
- [Meta Llama](https://toolweight.com/options/meta-llama), Open-weight Llama models, hosted almost everywhere but Meta
- [Mistral AI](https://toolweight.com/options/mistral-ai), European lab with an open-weight lineage and EU-resident hosting
- [Moonshot AI](https://toolweight.com/options/moonshot-ai), Kimi models, open-weight agentic performance at low cost

## Licence and attribution

Data from toolweight (https://toolweight.com), licensed CC-BY-4.0.

- Licence: [CC-BY-4.0](https://creativecommons.org/licenses/by/4.0/)
- Canonical HTML: https://toolweight.com/options/zhipu-zai
- Machine-readable: https://toolweight.com/options/zhipu-zai.md · https://toolweight.com/api/v1 · https://toolweight.com/mcp
- toolweight takes no affiliate revenue and sells no placements. Corrections: https://toolweight.com/suggest
