Z.ai GLM-5.3 API Pricing
Last updated:
GLM-5.3 costs $1.40 input / $4.40 output per 1M tokens. Z.ai also lists cached input at $0.26 per 1M tokens and a 1M-token context window. For coding-tool users, Z.ai sells a separate GLM Coding Plan subscription that includes GLM-5.3 for supported tools.
Z.ai API token prices
| Try it | |||||||
|---|---|---|---|---|---|---|---|
GLM-5.2 | Z.ai | Mid | 1M | $1.40 | $0.26 | $4.40 | Coding plan → |
GLM-5.3 | Z.ai | Mid | 1M | $1.40 | $0.26 | $4.40 | Coding plan → |
- GLM-5.2Z.aiMid
- Input
- $1.40
- Cached
- $0.26
- Output
- $4.40
- GLM-5.3Z.aiMid
- Input
- $1.40
- Cached
- $0.26
- Output
- $4.40
Sponsored links may earn us a commission at no extra cost to you. Affiliate status never changes model ordering.
All product names, logos, and brands are property of their respective owners and are used for identification purposes only.
If you are comparing GLM against other open models, Novita is a managed multi-model API to quote alongside first-party and hosted options. For very steady volume, check the GPU break-even guide.
Affiliate disclosure: this sponsored link may earn us a commission. It does not affect Z.ai table order or pricing data.
Ox Alpha pricing status
Confirmed as GLM-family, but not a priced Z.ai SKU
Z.ai told Bloomberg that Ox Alpha is a new GLM-series model and that it plans to release the weights. Its official pricing table, model documentation, and public Hugging Face organization do not yet publish an Ox Alpha model ID, API rate, artifact, or license. OpenRouter's zero-cost stealth/ox-alpha preview is not copied into the canonical Z.ai table. Read the pricing and weights-status analysis →
Direct video benchmark
CogVideoX-3 cost per accepted second
Z.ai lists CogVideoX-3 at $0.20 per video. In our bounded text-to-video and image-to-video test, billed failures and rejected footage raised the observed rate to $0.0516 per accepted second. Read the prompts, modality results, latency, and caveats →
API vs GLM Coding Plan
Use the API when you are building software, agents, automations, customer-facing features, or anything outside Z.ai's supported coding-tool flow. Use the GLM Coding Plan when one developer wants high-volume GLM access inside Claude Code, OpenClaw, Cline, Kilo Code, OpenCode, Crush, Goose, or another supported coding tool.
The important catch: Z.ai says the GLM Coding Plan is strictly limited to supported tools and products. Calls outside the plan use normal API billing, and GLM-5.2 consumes more quota than GLM-4.7 during peak windows. Z.ai's DevPack docs list GLM-5.2 and GLM-5-Turbo at 3x quota during 14:00-18:00 UTC+8, 2x off-peak, and a limited-time 1x off-peak benefit through the end of September.
Frequently asked questions
How much does GLM-5.3 cost on the Z.ai API?
Z.ai lists GLM-5.3 at $1.4 per 1M input tokens, $0.26 per 1M cached input tokens, and $4.4 per 1M output tokens. Cached input storage is listed as limited-time free.
Does the GLM Coding Plan include GLM-5.3?
Yes. Z.ai says GLM-5.3 is available to GLM Coding Plan users. The plan is limited to officially supported coding tools and is not a general API subscription.
Is GLM-5.3 cheaper through the subscription or API?
For supported coding tools, the subscription can be much cheaper because Z.ai describes the monthly quota as equivalent to roughly 15-30x the monthly subscription fee at API prices. For custom software, agents, SDK usage, or production apps, use the API pricing instead.
Does Ox Alpha have an official Z.ai API price?
No. Z.ai confirmed Ox Alpha is a new GLM-family model and plans to release its weights, but its official pricing table does not list an Ox Alpha SKU. OpenRouter's zero-cost stealth preview is a temporary third-party route, not a durable Z.ai list price.