Cover graphic reading The 20 percent cut that raises your bill, with three figures: minus 20.0 percent versus the old model's list price, plus 35.6 percent versus OpenRouter's current rate, and 8.54 dollars per 1,000 chatbot replies.

Qwen API Pricing: The 20% Cut That Raises Your Bill

Alibaba’s current flagship, Qwen3.8 Max, lists at $2.00 per million input tokens and $6.00 per million output. Its predecessor, Qwen3.7 Max, lists at $2.50 and $7.50. Put those two rate cards side by side and the arithmetic is obvious: the newer model is 20% cheaper. That is the comparison the rate card invites, and it is the one doing the rounds.

It is also, for almost anyone actually running a WordPress chatbot today, backwards. Qwen3.7 Max is not selling at list. It is on a 50% promotion, and on OpenRouter — the only route an MxChat install can reach Qwen through at all — it was priced at $1.475 in and $4.425 out when we fetched the page for this article. Measured against the price you would actually stop paying, moving to Qwen3.8 Max raises the bill by between 36% and 60%. Same two models, same day, opposite conclusion. The difference is entirely which price of the old model you measure from.

The two rate cards, and the third price nobody quotes

Here is every published price for both models, as of 4 September 2026. The list column is what Alibaba Cloud Model Studio publishes; the promotional column is the discount currently running on the older model; the OpenRouter column is what the reseller was charging when we checked.

Model / routeInput $/MOutput $/MCached read $/M
Qwen3.8 Max — Model Studio list$2.00$6.00$0.20
Qwen3.8 Max — OpenRouter$2.00$6.00$0.25
Qwen3.7 Max — Model Studio list$2.50$7.50
Qwen3.7 Max — Model Studio, 50% promo$1.25$3.75
Qwen3.7 Max — OpenRouter$1.475$4.425$0.295

Two things in that table are worth pausing on. First, Qwen3.8 Max costs the same on both routes — OpenRouter is not marking up the headline rate. Second, Qwen3.7 Max on OpenRouter sits at $1.475, which is 41% below Alibaba’s list and 18% above the direct promotional price. That is the shape you get when a model has been out long enough for the vendor to discount it and for a reseller to price against that discount. A brand-new flagship has no such history, so it sells at list everywhere.

What it costs to answer a thousand questions

Rate cards are per million tokens. Nobody buys a million tokens; they answer questions. So the only figure worth comparing is what one route costs to serve a real reply on a real site.

We measured this install’s own stored transcripts on 2 September: the system prompt is 3,478 characters, the retrieval context averages 11,426 characters across six samples, and the visitor’s message averages 135. The bot’s reply averages 681 characters across ten. At four characters to a token that is 3,760 input tokens and 170 output tokens per reply. The samples are small and we are stating them rather than smoothing them — but they are this site’s numbers, not a vendor’s example workload, and the same profile has been used for every pricing piece here since it was measured.

Horizontal bar chart titled Five ways to buy a Qwen flagship, comparing US dollars per 1,000 chatbot replies: Model Studio 50 percent promo $5.34, via OpenRouter $6.30, Qwen3.8 Max with prompt cached $7.02, Qwen3.8 Max list rate $8.54, and Model Studio list price $10.68.
Five prices for two models, on one measured workload. The grey bars need a direct Alibaba account.
RouteInput costOutput costPer 1,000 replies
Qwen3.7 Max — Model Studio promo$4.70$0.64$5.34
Qwen3.7 Max — OpenRouter$5.55$0.75$6.30
Qwen3.8 Max — with system prompt cached$6.00$1.02$7.02
Qwen3.8 Max — list, either route$7.52$1.02$8.54
Qwen3.7 Max — Model Studio list$9.40$1.28$10.68

The new model is the fourth-cheapest of five ways to buy a Qwen flagship. It beats exactly one thing: the old model’s undiscounted list price, which is the one price in the table nobody is currently paying.

Three answers to one question

“Should I move from Qwen3.7 Max to Qwen3.8 Max?” has three defensible numeric answers depending on nothing but the baseline.

Diverging bar chart titled Three answers to one model swap, showing the cost change from Qwen3.7 Max to Qwen3.8 Max: plus 35.6 percent versus OpenRouter, plus 60.0 percent versus the Model Studio promo, and minus 20.0 percent versus the Model Studio list price, drawn against a zero axis.
The sign flips depending on which price of the older model you measure from.

Against Alibaba’s list price for the old model, the swap saves 20.0%. Against the promotional price the old model is actually selling at, it costs 60.0% more. Against OpenRouter’s live rate — the only one of the three a WordPress plugin can reach — it costs 35.6% more. All three are arithmetically correct. Only the last one describes a bill anyone reading this is going to receive.

One honest caveat on the middle row. We fetched the OpenRouter figures ourselves on the day of writing, so the $1.475 and the +35.6% are current. The 50% Model Studio promotion is sourced from a pricing roundup whose page is dated June 2026, which makes its present status secondary evidence rather than something we verified at Alibaba directly. We are keeping it in because OpenRouter pricing the same model 41% below list is independent confirmation that Qwen3.7 Max is being sold well under its rate card right now — but the exact promotional figure deserves a check against Model Studio before you budget on it. Vendor prices quoted second-hand going stale is a recurring problem; we have hit it repeatedly, most recently on a tracker still showing Gemini Flash at a rate that expired months earlier.

Where the money actually goes

A retrieval chatbot is a lopsided workload. Of the 3,930 tokens in an average exchange here, 95.7% are input — system prompt, retrieved context, and a short question. The model writes 170 tokens back.

That has a consequence people consistently get wrong when comparing models. Output rates are where rate cards differ most dramatically and where spec comparisons focus, but on this workload the output rate contributes 11.9% of a Qwen3.8 Max bill. The input rate contributes 88.1%. When you are choosing a model to power site search or a documentation assistant, you are making an input-rate decision and very little else, which is the same conclusion we reached pricing DeepSeek’s off-peak window and the Grok 4.6 rate card.

ComponentTokens per replyShare of tokensShare of the Qwen3.8 Max bill
System prompt87022.1%20.4%
Retrieved context2,85672.7%66.9%
Visitor’s question340.9%0.8%
Model’s reply1704.3%11.9%

The cache lane, and an honest deflation

The system prompt is 870 tokens and it is byte-identical on every single reply — 23.1% of every input. That is exactly what prompt caching exists for, and it is the one lever on this workload that does not require changing models at all.

Alibaba prices a cached read on the flagship at $0.20 per million, a 90% discount. Caching the system prompt drops Qwen3.8 Max from $8.54 to $6.97 per thousand replies — an 18.3% cut. That is enough to clear the old model’s list price comfortably, but not enough to catch it on OpenRouter at $6.30: caching narrows the gap the migration opens, it does not close it. OpenRouter’s cached read is $0.25, a 25% premium on the same lane while matching list on the uncached one, so the reseller’s margin here lives entirely inside the discount tier.

Cache routeCached read $/MSaving per 1,000 repliesTotal per 1,000
No caching$8.54
OpenRouter$0.25$1.52$7.02
Model Studio$0.20$1.57$6.97

A 25% premium sounds like something to be annoyed about. It is worth four and a third cents per thousand replies on this workload. We are saying so plainly rather than leaving “25% markup” hanging as the takeaway, because it is a structurally interesting fact and a financially trivial one, and those are easy to confuse. If you serve fifty thousand replies a month, the difference between the two cache lanes is about two dollars a year.

Batch is half price and your chatbot cannot use it

Model Studio applies a 50% discount to batch calls, which prices Qwen3.8 Max at $1.00 in and $3.00 out — $4.27 per thousand replies, comfortably the cheapest number on this page. It is also unreachable. Batch processing is asynchronous: you submit a job and collect results later. A visitor waiting on a chat widget is the definition of a synchronous request.

This is now the third vendor in a row where the headline discount turns out to be attached to a lane a live chatbot cannot enter. We found it on OpenAI’s Batch API, again on Bedrock’s Flex tier for Grok, and now on Model Studio. The pattern is consistent enough to be worth stating as a rule: when a vendor advertises 50% off, check whether the discount applies to requests someone is waiting for. Batch discounts are real money for content generation, index embedding, and bulk classification — just not for the reply itself.

How Qwen sits against everything else we have priced

Horizontal bar chart titled Seven rate cards, one workload, showing US dollars per 1,000 chatbot replies: GPT-5.6 Luna $0.96, Claude Haiku 4.5 $4.61, Qwen3.7 Max via OpenRouter $6.30, Qwen3.8 Max $8.54, Grok 4.6 $8.54, Claude Sonnet 5 $9.22, and Claude Opus 5 $23.05.
Same measured workload throughout, so the only variable between rows is the rate card.

Two observations. Qwen3.8 Max and Grok 4.6 land on precisely the same figure, to the cent, because their published rates are identical — $2.00 in, $6.00 out. That is not a rendering error; it is two vendors arriving at the same price for a frontier model, which tells you something about where the market has settled.

The other is the spread. Qwen3.8 Max costs 8.9 times what this site’s current model does. GPT-5.6 Luna prices the same reply at $0.956 — a figure we published at $1.58 and then corrected downward by 39.5% after measuring actual transcripts instead of trusting the plugin’s configured ceiling. Whether that 8.9× gap is worth paying is a question about answer quality on your content, not about pricing, and no rate card will settle it. But it should be framed as a fit decision rather than as Qwen being expensive: at $8.54 it sits between Grok 4.6 and Claude Sonnet 5, which is exactly where a frontier model belongs.

The same trap scales up from one model pair to the whole market. Every index this week says LLM prices fell 84% since 2023, and on the identical measured reply OpenAI’s new GPT-6 Astra costs 48 times what this site pays today — because an index is a median across a catalogue that keeps widening downward, not the price of the newest model.

What this means if you run MxChat

Here is the part that changes the recommendation. MxChat ships selectable models for OpenAI, Anthropic, Google, xAI and DeepSeek, and switching between those is a dropdown in the plugin settings. Qwen is not among them. We checked this install directly: there is no Qwen entry in the model catalogue.

What the plugin does support is OpenRouter, as a provider whose model list is fetched dynamically once you supply a key. So Qwen is reachable — but only through the reseller, which has three consequences worth spelling out:

  • The Model Studio 50% promotion on Qwen3.7 Max is not available to you. That $5.34 row is a direct-account price.
  • Your cached reads cost $0.25 rather than $0.20. Worth about four cents per thousand replies, as above.
  • The cheapest Qwen lane you can actually configure is Qwen3.7 Max via OpenRouter at $6.30 — the older model, not the new one.

That last point is the practical version of this whole article. If you want a Qwen flagship behind your WordPress chatbot today, the newer, more capable, “20% cheaper” model will cost you 36% more than the one it replaced. That may well be worth it for the capability gain. It is not worth it because of a price cut, because from where you are standing there wasn’t one.

If you are still deciding whether an AI chatbot earns its keep at all, our running cost breakdown works through the whole bill rather than just the model, and the plugin comparison covers the options that don’t involve managing an API key. Setup for any of the supported providers is in the documentation.

Frequently asked questions

Is Qwen3.8 Max cheaper than Qwen3.7 Max?

Against Alibaba’s list prices, yes — $2.00/$6.00 against $2.50/$7.50, a 20% reduction. Against the prices the older model is actually selling at, no. Qwen3.7 Max is discounted 50% on Model Studio and sells at $1.475/$4.425 on OpenRouter, so the newer model costs 60% or 36% more depending on the route.

What does Qwen3.8 Max cost per chatbot reply?

On our measured profile of 3,760 input and 170 output tokens, $0.00854 per reply, or $8.54 per thousand. Caching the system prompt brings that to $6.97 on a direct account and $7.02 through OpenRouter.

Can I use Qwen with a WordPress chatbot plugin?

With MxChat, yes, through OpenRouter — add an OpenRouter API key and the Qwen models appear in the dynamically-loaded model list. There is no direct Alibaba Model Studio integration, so promotional pricing on a direct account is not reachable from the plugin.

Does the batch discount help a chatbot?

No. Batch processing is asynchronous and a chat reply is synchronous by definition. The 50% batch rate is genuinely useful for embedding your knowledge base or generating content in bulk, but it can never price a request a visitor is waiting on.

Why does the output price barely matter?

Because a retrieval chatbot sends far more than it receives. System prompt and retrieved context make up 95.7% of the tokens in an average exchange here, so the input rate drives 88.1% of the bill. A model with a dramatically cheaper output rate and a similar input rate will barely move your total.

The short version

Qwen3.8 Max is a competent frontier model at a competitive price, sitting between Grok 4.6 and Claude Sonnet 5 on a real WordPress chatbot workload. What it is not is a price cut. The 20% figure compares it to a list price its predecessor is not selling at, and on the only route a WordPress plugin can reach — OpenRouter — adopting it raises your model bill by 36%.

The general lesson is cheaper than the specific one: a percentage change is meaningless until someone tells you what it was measured from. When a vendor announces a cut, the question to ask is not “cheaper than what it used to list at” but “cheaper than what I am paying this month”. Those are frequently different numbers, and occasionally, as here, they point in opposite directions.

Similar Posts