GLM-5.3 pricing cover: $6.30 per 1,000 chatbot replies on Z.ai, $0.98 on Relace, $12.45 on GLM-5.3 Prime

GLM-5.3 Pricing: One Model, 32 Hosts, 6x Price Gap

GLM-5.3 has one official price: Z.ai charges $1.40 per million input tokens and $4.40 per million output tokens, with cached input at $0.26. But GLM-5.3 is an open-weight model, and on 6 October 2026 OpenRouter listed 41 endpoints for it from 32 different hosts. Their input prices ran from $0.03 to $2.80 per million. If you are pricing GLM-5.3 for a WordPress chatbot, the question is not only what the model costs. It is who you are paying to run it.

We wanted to see the bill rather than the rate card. On 6 October 2026 we sent the same three customer-support questions, with the same system prompt and the same retrieved help pages, to GLM-5.3 on four hosts, and to GLM-5.3 Prime, Flash and FlashX. Every request went from our production WordPress server through OpenRouter, pinned to one host with fallbacks off. This is what GLM-5.3 pricing works out to per 1,000 chatbot replies, as billed.

  • The same model cost $6.30 per 1,000 replies on Z.ai’s own endpoint and $0.98 on Relace. Novita cost $1.95 and DeepInfra $2.54. All four gave correct answers to all three questions.
  • GLM-5.3 Prime, served only by Alibaba, cost $12.45 per 1,000 replies, twice Z.ai’s standard model, and its answers were no better on our questions.
  • GLM-5.3 Flash cost $0.77 and FlashX $1.78. Flash was the cheapest configuration we measured, but it took 8.5 seconds a reply.
  • You cannot switch reasoning off on Z.ai’s endpoint. The request returned HTTP 400, “Reasoning is mandatory for this endpoint and cannot be disabled.” Setting effort to low did work: it cut the bill 14% and the wait 47%.
  • The cheapest host was also the fastest, at 1.1 seconds a reply. DeepInfra averaged 31 seconds. For a live chat window that difference matters more than the price.

GLM-5.3 pricing: the rate card, and the hosts

Z.ai’s first-party rates are below, read from its OpenRouter listing on 6 October 2026 and matching the rates Z.ai announced in August. GLM-5.3 Prime is a separate model ID that only Alibaba serves. Z.ai lists cached-input storage as free for a limited time, with no end date.

ModelInput / 1MCached input / 1MOutput / 1MContextWho serves it
GLM-5.3$1.40$0.26$4.401MZ.ai and 31 other hosts
GLM-5.3 Prime$2.80$0.56$8.801MAlibaba only
GLM-5.3 FlashX$0.37$0.09$1.251MZ.ai only
GLM-5.3 Flash$0.15$0.03$0.501MZ.ai and about 30 other hosts
GLM-5.1 (previous)$1.40$0.26$4.40205KZ.ai

GLM-5.3 kept GLM-5.1’s price and raised the context window from about 205,000 tokens to a million. For a chatbot that sends 4,000 tokens a turn, the bigger window does not change anything. The price per token is what counts.

The interesting part is the other hosts. Here are the GLM-5.3 endpoints we tested, plus a few others, as OpenRouter listed them on 6 October:

HostInput / 1MOutput / 1MCache read / 1MWeights format
Relace$0.03$12.00$0.03Not stated
Novita$0.42$1.32$0.078FP8
DeepInfra$0.5625$2.50$0.125FP4
Together, Fireworks, Cloudflare, Parasail and others$1.40$4.40$0.26Varies
Z.ai (first party)$1.40$4.40$0.26FP8
Fireworks / BaseTen “fast”$2.10$6.60$0.21 to $0.39Varies
Alibaba “fast”$2.80$8.80$0.56Not stated

Two things stand out. Relace prices input at almost nothing and output at nearly three times Z.ai’s rate. Several hosts serve the model in FP4 or FP8, which are compressed number formats that make it cheaper to run and can change output slightly. Most hosts charge exactly Z.ai’s price. Those are the ones you get by default.

How we tested

We use the same setup for every model we price, so the numbers in this post can be compared with our Mistral API pricing test, our Claude Haiku 4.5 test and the others linked below.

  • System prompt: the one our own site’s chatbot uses, about 800 tokens of instructions on tone, length and when to offer a discount code.
  • Retrieved sources: four help-page excerpts per question, about 16,000 characters, taken from our documentation and FAQ the way a retrieval chatbot would.
  • Three questions: whether the bot can look up a WooCommerce order status for a logged-in customer, whether the knowledge base needs Pinecone, and whether a visitor can be handed to a human on Slack or Telegram.
  • Two passes: each request sent once, then again unchanged, so we could see what a cache hit does.
  • Costs as billed: taken from OpenRouter’s usage report on each call. We checked them against the rate cards and they matched to the cent.

That is 54 chatbot replies across nine configurations, plus 12 short probe calls. Every request was about 4,130 prompt tokens. GLM-5.3 counts tokens the same way on every host, because it is the same model.

Cost per 1,000 chatbot replies, by host

This is the first ask: a new question with new retrieved text, which is what most chatbot turns look like. The amber part of each bar is what the model’s hidden reasoning tokens cost.

Bar chart of GLM-5.3 cost per 1,000 chatbot replies: Relace $0.98, Novita $1.95, DeepInfra $2.54, Z.ai $6.30, Z.ai effort low $5.43, Prime $12.45, Flash $0.77, FlashX $1.78
ConfigurationPer 1,000 replies (first ask)Per 1,000 (repeat, cached)Reasoning tokens / replyAnswer tokens / replySeconds / reply
GLM-5.3 on Relace$0.98$1.0010611.1
GLM-5.3 on Novita$1.95$0.52178548.8
GLM-5.3 on DeepInfra$2.54$1.271165531.2
GLM-5.3 on Z.ai$6.30$1.95200523.5
GLM-5.3 on Z.ai, effort low$5.43$1.350541.9
GLM-5.3 Prime (Alibaba)$12.45$4.48180523.7
GLM-5.3 FlashX (Z.ai)$1.78$0.97196643.0
GLM-5.3 Flash (Z.ai)$0.77$0.71286748.5

On Z.ai, Novita, DeepInfra and Prime, input was 82% to 84% of the first-ask bill. That is normal for a support chatbot: it sends about 4,100 tokens of instructions and help text to get back a 40-word answer. So for those hosts, the input rate is the price that matters, and the 6x gap between Relace and Z.ai comes almost entirely from the $0.03 versus $1.40 input price.

Relace turns that around. Its input is so cheap that output was 87% of its bill. That makes Relace’s price depend on how much the model writes, and on Relace GLM-5.3 wrote very little. It used about 10 reasoning tokens per reply, against about 200 on Z.ai, for the same questions and the same settings. We do not know why. Relace’s answers were as accurate as everyone else’s on our three questions. But if Relace’s endpoint reasoned as much as Z.ai’s does, its $12 output rate would put the same reply at about $3.15 per 1,000, not $0.98. If you route a chatbot to Relace, keep an eye on output tokens.

Reasoning is always on, and effort low is the setting to use

GLM-5.3 is a reasoning model. Before it answers, it writes hidden “thinking” tokens that you pay for at the output rate. On Z.ai’s endpoint you cannot turn this off. A request with reasoning disabled came back with HTTP 400 and the message “Reasoning is mandatory for this endpoint and cannot be disabled.”

You can lower the reasoning effort, and for a support chatbot you should. With effort: low, GLM-5.3 on Z.ai used between 0 and 3 reasoning tokens per reply instead of about 200. The bill fell from $6.30 to $5.43 per 1,000 replies, 14% less. The average wait fell from 3.5 seconds to 1.9. The answers were the same length and just as accurate. One of them was a little more detailed than the default-effort version, telling the customer what an order lookup shows (date, status, items and totals).

The saving is small because reasoning is a small part of the bill when the input is 4,000 tokens. The speed is the bigger win. For a chat widget, getting under two seconds matters more than 87 cents per 1,000 replies.

What the prompt cache did

A chatbot’s system prompt is the same on every turn, and on a follow-up question the whole earlier conversation is repeated. Hosts that cache that repeated text can charge less for it. To see the effect, we sent every request a second time, unchanged.

Bar chart of GLM-5.3 cost per 1,000 replies on first ask versus an unchanged repeat: Z.ai $6.30 to $1.95, Novita $1.95 to $0.52, Relace $0.98 to $1.00, Prime $12.45 to $4.48

On the repeat, 3,900 to 4,100 of the 4,130 prompt tokens came from cache on Z.ai, Novita, Relace and Prime. The cache cut Z.ai’s bill by 69% and Novita’s by 73%. Relace’s bill did not move, because its input was already priced the same as its cache reads, $0.03 per million.

Flash was the exception. On both passes Z.ai’s Flash endpoint reported only 256 cached tokens, so the repeat cost almost the same as the first ask ($0.71 against $0.77). FlashX cached about three-quarters of the prompt on the repeat.

Do not plan a budget around the repeat numbers. In a retrieval chatbot, each new question brings new help text, so most of the prompt changes from turn to turn. Only the system prompt and earlier turns are repeated. The first-ask column is the honest figure for new questions. The repeat column is the floor for follow-ups. We tested this in detail for OpenAI in our post on OpenAI prompt caching, and the same pattern holds here.

Speed: the cheapest host was the fastest

Bar chart of seconds per GLM-5.3 chatbot reply by host: Relace 1.1, Z.ai effort low 1.9, FlashX 3.0, Z.ai 3.5, Prime 3.7, Flash 8.5, Novita 8.8, DeepInfra 31.2

The response times varied even more than the prices. Relace answered in about one second. DeepInfra took 31 seconds on average for a new question and 33.6 seconds at worst, which is far too long for a chat window. Most visitors will have closed the tab. On the repeat pass DeepInfra was faster (6.7 seconds), so some of that delay may be queueing at the time we tested, but we can only report what we saw.

Novita was 8.8 seconds, and Flash was 8.5. Flash is a smaller model, but it reasoned more than any other configuration (286 tokens per reply), which explains the wait.

Answer quality

We read all 54 replies. Every configuration answered all three questions correctly: the WooCommerce order lookup needs the WooCommerce add-on and MxChat Pro, Pinecone is optional because embeddings can be stored in WordPress, and live handoff to Slack and Telegram is in the free plugin. The differences were small.

  • Flash and FlashX gave the fullest answers (45 words on average). They mentioned details such as Telegram giving each visitor their own forum topic. Flash linked to our documentation page, which exists.
  • Novita linked a customer to /pricing, a page we do not have. It redirects to the Pro product page, so the link works, but the model made up the URL.
  • Prime’s answers were no better than standard GLM-5.3’s on these questions, at twice the price. Prime may do better on harder tasks. A support chatbot reading help pages does not need it.
  • Only one of 54 replies used the system prompt’s discount offer. In our Claude Sonnet 5.5 test, Sonnet used it in 14 of 18 replies. If your bot’s instructions include a sales prompt, GLM-5.3 mostly ignores it.

How GLM-5.3 compares with other models on the same request

We have run this same chatbot request on most of the major models over the last two weeks. These are first-ask costs per 1,000 replies, each billed by the vendor or by OpenRouter.

ModelPer 1,000 repliesTested
Ministral 3 14B$0.455 October
GPT-6 Luna$0.5728 September
Mistral Small 4$0.615 October
GLM-5.3 Flash (Z.ai)$0.776 October
GLM-5.3 on Relace$0.986 October
Gemini 3.5 Flash-Lite$1.353 October
DeepSeek V4.1 Flash$1.5327 September
GLM-5.3 on Novita$1.956 October
Mistral Large 3$2.085 October
Claude Haiku 4.5$4.934 October
GLM-5.3 on Z.ai$6.306 October
Mistral Medium 3.5$6.785 October
GPT-6 Sol (effort none)$10.6730 September
GLM-5.3 Prime$12.456 October
Claude Sonnet 5.5 (low effort)$13.024 October

Bought from Z.ai, GLM-5.3 costs more per chatbot reply than Claude Haiku 4.5 and almost as much as Mistral Medium 3.5. Bought from a cheap host, it sits with DeepSeek V4.1 Flash and Gemini Flash-Lite. Few other models can move that far down the table, because closed models like Claude, GPT and Gemini are sold only by their makers (plus a few cloud resellers at the same price). An open-weight model is sold by anyone who wants to run it. That is the main reason to consider GLM-5.3, and also why its price is hard to pin down.

For comparison, GPT-6 Luna is still cheaper than any standard GLM-5.3 host on this request. If you want the lowest bill from one vendor and no host-choosing, Luna or Mistral Small 4 is simpler. We also priced Qwen3.8 Max, the other big model family from China, in September.

How OpenRouter picks a host for you

If you send z-ai/glm-5.3 to OpenRouter without any routing settings, you do not get the cheapest host. In our test, OpenRouter’s default routing sent five of six requests to Z.ai and one to Mistral, both at the full $1.40 / $4.40 rate. OpenRouter’s documentation says the default favours price but also weighs uptime and recent performance, so the cheapest hosts are not always chosen.

Two settings changed that in our probes:

  • Add :floor to the model name (z-ai/glm-5.3:floor). The request went to Relace.
  • Send "provider": {"sort": "price"} in the request. It also went to Relace.
  • :nitro sorts by speed instead. Our probe went to BaseTen.

The model suffix is the easiest to use from a WordPress plugin, because it goes in the model name field and needs no code. You can also block hosts you do not want in your OpenRouter account settings, which is worth doing for any host that was very slow in your own tests.

One warning. The cheapest hosts often run compressed versions of the model (FP4 or FP8), and their output can differ slightly from Z.ai’s. On our three questions we saw no difference in accuracy. If your chatbot does anything where a small error matters, such as quoting prices or policies, test the host you choose on your own content first.

Which GLM-5.3 option to use for a WordPress chatbot

If you wantUseExpect per 1,000 replies
The lowest bill, and you can watch output tokensGLM-5.3 via :floor (Relace on 6 October)About $1, fast
Low cost from a single host with a normal price structureGLM-5.3 on NovitaAbout $2, slow (9 s)
Z.ai’s own service, answering quicklyGLM-5.3 on Z.ai with effort lowAbout $5.40, under 2 s
The cheapest GLM model and you can waitGLM-5.3 FlashAbout $0.77, 8.5 s
A middle option from Z.aiGLM-5.3 FlashXAbout $1.80, 3 s
A support chatbotNot Prime$12.45, no better here

Our own pick for a support bot would be standard GLM-5.3 on a cheap, fast host, with a monthly spending cap set in OpenRouter. If speed is not a problem for your visitors, Flash is cheaper still. If you need to stay with Z.ai, set effort to low.

Using GLM-5.3 with MxChat

MxChat connects to GLM-5.3 through an OpenRouter key, which reaches every host and variant in this post, or through the OpenAI-compatible custom endpoint if you have a Z.ai key of your own. Add the key under MxChat → Settings → API Keys and pick the model. The MxChat documentation covers model setup, and MxChat Pro adds the WooCommerce order features our first test question asked about. Whichever host you choose, run a week of your own visitors’ questions before you trust any per-reply figure, including ours. Hosts change their prices and speeds often.

Frequently asked questions

How much does GLM-5.3 cost?

On Z.ai’s own API, as of 6 October 2026, GLM-5.3 costs $1.40 per million input tokens, $0.26 per million cached input tokens and $4.40 per million output tokens. GLM-5.3 Prime costs $2.80 / $8.80, FlashX $0.37 / $1.25 and Flash $0.15 / $0.50. Other hosts charge from $0.03 to $2.80 per million input tokens for standard GLM-5.3.

What does GLM-5.3 cost per chatbot reply?

On our test request (about 4,130 prompt tokens, a short answer), GLM-5.3 cost $6.30 per 1,000 replies on Z.ai, $5.43 with effort set to low, $2.54 on DeepInfra, $1.95 on Novita and $0.98 on Relace. GLM-5.3 Flash cost $0.77 and Prime $12.45.

Can I turn off reasoning on GLM-5.3?

Not on Z.ai’s endpoint. A request with reasoning disabled returned HTTP 400, “Reasoning is mandatory for this endpoint and cannot be disabled.” Setting reasoning effort to low reduced reasoning to between 0 and 3 tokens per reply, cut the cost 14% and nearly halved the response time.

Is GLM-5.3 Prime worth it for a chatbot?

Not in our test. Prime costs twice as much as standard GLM-5.3 and is only served by Alibaba. Its answers to our support questions were no more accurate or complete than standard GLM-5.3’s.

Why is GLM-5.3 cheaper on some hosts than on Z.ai?

GLM-5.3’s weights are open, so any provider can run it and set its own price. Some hosts run compressed versions (FP4 or FP8) that cost less to serve. Some, like Relace, price input very low and output high. Check both rates and the format before you choose a host.

How do I get the cheapest GLM-5.3 host on OpenRouter?

Use the model name z-ai/glm-5.3:floor, or send "provider": {"sort": "price"} with the request. Both routed to Relace in our probes on 6 October 2026. Without either setting, OpenRouter sent our requests to Z.ai and Mistral at full price.

Similar Posts