Claude Sonnet 5.5 Pricing: Same Price, Not 30% Cheaper
Claude Sonnet 5.5 pricing is identical to Sonnet 5: $2 per million input tokens and $10 per million output tokens. Anthropic released it on 28 September 2026 and said it “costs up to 30% less per task” because it needs fewer tokens for the same work. If you run a WordPress chatbot on Sonnet, that is the number you care about, so we tested it the same day on the workload a site chatbot actually sends.
The short version: on a retrieval chatbot, Sonnet 5.5 is not 30% cheaper. With the API’s default settings it cost 6% more per reply than Sonnet 5. Set one parameter and it costs the same as Sonnet 5 and answers 28% faster. It was also the more careful of the two models on the one question where our documentation had a gap.
- Per 1,000 replies: Sonnet 5 $12.94, Sonnet 5.5 at its default $13.70 (+5.9%), Sonnet 5.5 at
mediumeffort $12.99 (+0.4%). Measured on 48 real Messages API calls from a production WordPress server, with the site’s own system prompt, retrieved documentation and three questions visitors actually ask. - Sonnet 5.5’s default effort on the API is
high. At that setting it thought before answering on the hardest question, adding about 210 thinking tokens and a second of wait. Atmediumit never thought on these questions. - The tokenizer did not change. The same request counted 6,765 tokens on Sonnet 5 and 6,767 on Sonnet 5.5. The rate card is the bill.
- Sonnet 5.5 rejects
temperaturewith HTTP 400, like Sonnet 5, and rejectsthinking: disabled. A plugin that sends a temperature to any model it does not recognise will fail on every message.
The Claude Sonnet 5.5 rate card
These are Anthropic’s published rates per million tokens as they read on 29 September 2026. Sonnet 5 kept its $2 / $10 price; the increase to $3 / $15 that was scheduled for 1 September was cancelled, which we covered in our Sonnet 5 pricing page. Sonnet 5.5 launched at the same numbers, and Sonnet 5 is now listed as a legacy model that is still available.
| Per 1M tokens | Claude Sonnet 5 | Claude Sonnet 5.5 | Claude Opus 5.5 |
|---|---|---|---|
| Input | $2.00 | $2.00 | $4.00 |
| Cache write (5 minutes) | $2.50 | $2.50 | $5.00 |
| Cache write (1 hour) | $4.00 | $4.00 | $8.00 |
| Cache read | $0.20 | $0.20 | $0.20 |
| Output (includes thinking) | $10.00 | $10.00 | $20.00 |
| Batch API input / output | $1 / $5 | $1 / $5 | $2 / $10 |
| Default effort on the API | — | high | medium |
| Context window / max output | — | 1M / 128K | 1M / 128K |
| Retirement | legacy, still available | not before 28 Sep 2027 | not before 22 Sep 2027 |
Because the prices did not move, any saving from Sonnet 5.5 has to come from using fewer tokens. That is exactly Anthropic’s claim, and it is a claim about output: fewer steps, shorter reasoning, less rework. Whether it reaches your bill depends on how much of your bill is output in the first place.
How we tested
Every call went from the mxchat.ai WordPress server to Anthropic’s Messages API with the site’s own Claude API key, which never left the server. Each request had the same layout MxChat uses for a knowledge-base chatbot: the live system prompt (3,478 characters, 1,212 tokens) marked for prompt caching, and a user message carrying four retrieved documentation excerpts plus the visitor’s question. max_tokens was 1,000, which is what our plugin sends for chat. The three questions were real ones from our own chat logs:
- Does MxChat work with WooCommerce, and can the bot look up an order status for a logged-in customer?
- Do I need Pinecone to use the knowledge base, or can I store embeddings in WordPress?
- Can a visitor be handed over to a human on Slack or Telegram when the bot cannot answer?
We ran four configurations, each on every question twice: Sonnet 5 as a plugin calls it today (no effort parameter), Sonnet 5.5 with no effort parameter (so the API default, high), and Sonnet 5.5 at medium and low. That is 24 calls per batch. We ran two batches, because the first one taught us something.
The first batch used a simple keyword search to pick the retrieved pages, and for two of the three questions it pulled in our own pricing articles instead of the documentation. The bots had less to work with: the Pinecone question got “I don’t have enough information” from both models every time, and the handoff question was answered from one Telegram page. We kept those numbers as the “thin sources” case, which every real chatbot hits some of the time. The second batch used the matching documentation sections (the WooCommerce docs, the knowledge-base and Pinecone sections, the Telegram and Slack handoff docs), and every reply was a real answer. The headline numbers below come from the second batch. Six further calls tested request parameters, and 18 token-count calls checked the tokenizer.
We read every reply before pricing it. A cheaper model that answers worse is not cheaper, and a test set that makes every model refuse measures nothing about quality.
What a reply costs on each model
Costs are per 1,000 replies at list price, computed from the usage fields the API returned. “Cached system prompt” is the steady state of a busy site: the 1,212-token system prompt is read from cache at $0.20 per million and the retrieved sources are billed as fresh input, because they change with every question. “Nothing cached” is the same request with no caching at all.
| Configuration | Per 1,000 replies, cached system prompt | Nothing cached | vs Sonnet 5 | Output tokens (thinking) | Mean response time |
|---|---|---|---|---|---|
| Sonnet 5 (no effort sent) | $12.94 | $15.12 | — | 145 (0) | 2.66 s |
Sonnet 5.5, default (high) | $13.70 | $15.88 | +5.9% | 221 (71) | 2.91 s |
Sonnet 5.5, medium | $12.99 | $15.17 | +0.4% | 150 (0) | 1.91 s |
Sonnet 5.5, low | $13.24 | $15.42 | +2.3% | 174 (0) | 2.19 s |
Two things stand out. First, no setting made Sonnet 5.5 cheaper than Sonnet 5 on this workload. Second, the spread between the best and worst setting is only 6%, because the part of the bill Sonnet 5.5 improves is small. The next section shows why.
For the “thin sources” batch the gap was wider. Sonnet 5 cost $13.18 per 1,000 replies and Sonnet 5.5 at its default cost $14.57 (+10.5%), because it thought before answering on four of its six replies. At medium it was $13.87. When the retrieved pages only half answer the question, the high-effort default spends tokens working out what it can and cannot say.
Why “30% cheaper” cannot reach a chatbot
A retrieval chatbot sends a lot and receives a little. Each of our requests carried about 5,600 tokens of fresh input (the retrieved sources and the question) plus the cached system prompt, and got back 150 to 220 tokens. At $2 input and $10 output, that makes input about 89% of the bill on Sonnet 5 and 84% on Sonnet 5.5 at its default.
Anthropic’s efficiency claim is about output: the model does the same work in fewer tokens, so an agent that writes code or runs tools for twenty turns finishes with a smaller bill. A chatbot answer is one turn, and its output is already short. On our Sonnet 5 numbers, output was $1.45 of the $12.94. If Sonnet 5.5 produced no output tokens at all, the reply would still cost $11.49. The best any output-side improvement could do on this workload is about 11%, and a model that thinks more by default goes the other way.
The lever that does move a chatbot’s bill is the input. Four documentation excerpts at about 1,350 tokens each were roughly 5,400 of the 5,600 fresh tokens. Retrieving three excerpts instead of four, or trimming each one to the relevant paragraph, saves more than any model switch in this family. Our Claude tokenizer measurement goes through how much text a token buys on the current tokenizer, which Sonnet 5.5 shares.
The setting that decides it: effort
Sonnet 5.5 uses adaptive thinking: it decides for itself whether to reason before answering, and the effort parameter (sent as output_config.effort) steers how readily it does. Anthropic’s model table lists the API default as high. Opus 5.5, by comparison, defaults to medium. If your code sends nothing, Sonnet 5.5 runs at high.
| Question | Sonnet 5 | Sonnet 5.5 default (high) | Sonnet 5.5 medium | Sonnet 5.5 low |
|---|---|---|---|---|
| WooCommerce order status | 175 tokens, 3.3 s | 140 tokens, 1.8 s | 164 tokens, 1.8 s | 178 tokens, 2.0 s |
| Pinecone or WordPress storage | 81 tokens, 1.9 s | 110 tokens, 2.9 s | 89 tokens, 1.4 s | 106 tokens, 2.0 s |
| Human handoff on Slack or Telegram | 179 tokens, 2.8 s | 413 tokens (214 thinking), 4.1 s | 197 tokens, 2.6 s | 239 tokens, 2.6 s |
The pattern is clear even at two runs per cell. On the two questions the documentation answered directly, Sonnet 5.5 at any effort was as fast or faster than Sonnet 5. On the handoff question, where the docs describe a visitor asking for a human but say nothing about the bot escalating by itself, the default high effort thought for about 210 tokens both times and roughly doubled the reply length. At medium and low it answered directly.
Two further details from the calls. low produced slightly longer visible answers than medium (77 words against 66 on average), so it was not the cheapest setting here. And effort: "minimal" is rejected with HTTP 400; the accepted values are low, medium, high, xhigh and max. If a plugin maps a generic “fastest” option to minimal, as several do for OpenAI models, it will fail on Claude. We found the same trap on the OpenAI side last week in our GPT-6 Luna test.
Speed: the claim that holds
Anthropic says Sonnet 5.5 is more than 30% faster. At medium, the mean response time across all six calls was 1.91 seconds against 2.66 for Sonnet 5, which is 28% faster. That is wall time from our server to the API and back, including the network, so the model-side gain is somewhat larger. At the default high, the thinking on one question pulled the mean up to 2.91 seconds, slower than Sonnet 5.
For a chatbot, response time is the part visitors notice. A reply that starts arriving a second sooner matters more to them than a cent per thousand replies matters to you, and medium gets both.
Quality: did Sonnet 5.5 answer better?
On the WooCommerce and Pinecone questions both models gave correct, short answers drawn from the documentation: the WooCommerce add-on shows order history to logged-in customers only and needs an active Pro licence, and Pinecone is optional, with embeddings stored in the WordPress database by default.
The handoff question separated them. Our documentation says a visitor is handed to Slack or Telegram when they ask for a human. It does not say the bot escalates on its own when it cannot answer. In one of its two runs, Sonnet 5 told the visitor the conversation is routed to a human “when a visitor requests human help (or the bot can’t answer)”. That parenthesis is not in the sources. Sonnet 5.5 declined to claim it in all six of its runs, at every effort level, with lines such as “I don’t have details on whether the bot hands off automatically when it can’t answer.”
That is one question and eight replies, not a benchmark. But it is the behaviour a site owner wants from a support bot: say what the documentation says and stop there. It is also a reason to switch that has nothing to do with price.
Prompt caching carries over, per model
The first Sonnet 5 call wrote the 1,212-token system prompt to the cache at $2.50 per million; every later Sonnet 5 call read it at $0.20. The first Sonnet 5.5 call wrote it again. The cache belongs to the model, not to the account, so switching models costs one cache write per prompt. Changing effort does not: the medium and low calls read the cache the default calls had written. A 1,212-token system prompt was long enough to cache on both models. If your system prompt is shorter than about 1,024 tokens it will not cache at all, which our prompt caching measurement covers for OpenAI’s side.
Should a WordPress chatbot move to Sonnet 5.5?
Yes, with effort set to medium. It costs the same as Sonnet 5, answers faster, stayed closer to the sources on our test, and has a retirement commitment a year out, while Sonnet 5 is now a legacy model. The steps:
- Change the model ID to
claude-sonnet-5-5. - Send
output_config: {"effort": "medium"}. Leaving it out gives youhigh, which on our test cost 6% more and was slower whenever the sources left a gap. - Do not send
temperatureorthinking: {"type": "disabled"}. Both return HTTP 400 on Sonnet 5.5. - Look at your retrieval before your model. Input is 84 to 89% of a chatbot reply on Sonnet. Fewer, tighter excerpts beat any model change in this family.
- Keep a spending cap on the key. A public chat box is an open door to your API bill; our denial-of-wallet notes cover the limits worth setting.
If Sonnet is more model than you need, Haiku 5.5 is due “in the coming weeks” and Haiku 4.5 has a retirement date of no sooner than 15 October 2026. If you need more, our Opus 5.5 test priced the step up on the same kind of request.
What this means for MxChat sites
MxChat’s model list includes Claude Sonnet 5, and that keeps working at the same price. Sonnet 5.5 is not in the list yet. Adding it takes two entries in the plugin’s model catalog: the model itself, and Sonnet 5.5 on the list of Claude models that must not receive a temperature. Without the second, every chat message would return an error. We have logged both for the next release. Until then, Sonnet 5 costs the same per reply, so nothing is lost by waiting. The MxChat documentation covers model selection and the knowledge base, and MxChat Pro adds the WooCommerce and live-agent features our test questions asked about.
FAQ
How much does Claude Sonnet 5.5 cost?
$2 per million input tokens, $10 per million output tokens, $0.20 per million cache-read tokens and $2.50 per million 5-minute cache-write tokens, the same as Sonnet 5. The Batch API halves input and output to $1 and $5. On our retrieval chatbot test that came to $12.99 per 1,000 replies at medium effort.
Is Claude Sonnet 5.5 cheaper than Sonnet 5?
Not for a chatbot. The prices are identical and the tokenizer is the same. At the API default effort Sonnet 5.5 cost 5.9% more per reply on our test, and at medium it cost 0.4% more. Anthropic’s “up to 30% less per task” applies to longer, output-heavy work such as agents and coding.
What is Claude Sonnet 5.5’s default effort?
high on the Claude API. Set output_config.effort to medium for chat: on our test it removed the thinking tokens, matched Sonnet 5’s cost and cut response time by 28%.
Does Claude Sonnet 5.5 support temperature?
No. Sending temperature returns HTTP 400 with the message that it is deprecated for this model. Sonnet 5 behaves the same way.
Is Claude Sonnet 5 being retired?
Not yet. Anthropic lists Sonnet 5 as a legacy model that is still available, with no retirement date published as of 29 September 2026. Sonnet 5.5 is committed to stay available until at least 28 September 2027.
Method note: 48 Messages API calls on 29 September 2026 from the mxchat.ai production server (two batches of 24), models claude-sonnet-5 and claude-sonnet-5-5, the site’s live system prompt with prompt caching, four retrieved documentation excerpts per question, three questions, each configuration run twice per batch, max_tokens 1,000. Headline costs are from the second batch, in which every reply answered the question; the first batch, where retrieval returned weaker sources, is reported as the thin-sources case. Costs are computed from the returned usage fields at Anthropic’s list rates with the system prompt read from cache. Latency is wall time from the server, including network. Six further calls tested request parameters and 18 count_tokens calls compared the tokenizers.