Cover card: Anthropic halved the Claude Sonnet 5.5 cache-read price from $0.20 to $0.10 per million tokens; a chatbot saves 0.9% on a new question, 15% on a cached follow-up and 35% with its whole knowledge base cached

Claude Prompt Caching: What Sonnet 5.5’s Cut Saves

On 7 October 2026 Anthropic cut the price of a cache read on Claude Sonnet 5.5 from $0.20 to $0.10 per million tokens. Input, output and cache writes stayed where they were. Coverage of the change carried Anthropic’s framing that Sonnet 5.5 becomes about 20% cheaper on most agentic work. A WordPress chatbot is not most agentic work, so we measured it: 27 Sonnet 5.5 calls from our production server on 9 October, two days after the change, in the three ways a site chatbot can use Claude prompt caching.

The saving ran from almost nothing to about a third of the bill. It depends entirely on how much of each request is read from cache, and on a typical retrieval chatbot that share is small.

  • A new question in MxChat’s default layout got 0.9% cheaper: $13.29 per 1,000 replies, down from $13.41. Only the 1,212-token system prompt is cached; the 5,600 tokens of retrieved pages and the question are fresh every time.
  • A follow-up in a conversation with automatic caching got 14.6% cheaper: $4.08 per 1,000, down from $4.78. Here most of the prompt is the earlier conversation, read back from cache.
  • A bot with its whole knowledge base in a cached prompt got 35.4% cheaper: $3.27 per 1,000, down from $5.06, on a 17,946-token prompt. Without caching the same request costs $37.43.
  • The arithmetic is simple: the cut halves one line of the bill, so you save half of whatever share cache reads were. A 20% saving needs cache reads to be about 40% of what you pay.
  • Automatic caching is not free on the first message. It writes the whole first question to cache at 1.25 times the input price, adding $2.81 per 1,000 conversations. It pays that back if roughly one conversation in four gets a follow-up within five minutes.

Claude prompt caching pricing after 7 October

Anthropic bills a cached prompt in two parts. Writing a prefix to the cache costs 1.25 times the input price for a 5-minute cache, or twice the input price for a 1-hour cache. Reading it back, which also refreshes the timer, costs a fraction of the input price. For most Claude models that fraction is 10%. Since the 7 October change, Anthropic’s pricing page lists it as 5% for Sonnet 5.5 and Opus 5.5, and 2.5% for Fable 5.1. These are the rates we read on 9 October 2026, per million tokens.

Per 1M tokensInput5-minute cache write1-hour cache writeCache readRead as share of inputOutput
Claude Sonnet 5.5, before 7 Oct$2$2.50$4$0.2010%$10
Claude Sonnet 5.5, from 7 Oct$2$2.50$4$0.105%$10
Claude Sonnet 5$2$2.50$4$0.2010%$10
Claude Opus 5.5$4$5$8$0.205%$20
Claude Fable 5.1$10$12.50$20$0.252.5%$50
Claude Haiku 5.5 (prompt up to 100K)$0.10$0.125$0.20$0.0110%$0.50
Claude Haiku 4.5$1$1.25$2$0.1010%$5

Two comparisons put the new number in context. A cached token on Sonnet 5.5 now costs the same $0.10 as a cached token on Haiku 4.5, a model that charges half as much for fresh input. And it matches OpenAI’s cached-input rate for GPT-6.1 Sol, which also lists at $2 input and $10 output; we covered that model in our GPT-6.1 Sol pricing test. On the cache line, the two mid-tier flagships are now level.

Nothing else about Sonnet 5.5 changed. Our Claude Sonnet 5.5 pricing test from launch week still describes the model; only its cache-read line is out of date, and we have added a note to it.

How we tested

All calls went from our WordPress server to Anthropic’s Messages API with the site’s own API key, on 9 October 2026, using model claude-sonnet-5-5 with effort set to low unless stated. Each request carried our live chatbot system prompt (3,478 characters, 1,212 tokens on Sonnet 5.5) and one of three real pre-sales questions a visitor might ask:

  • Does MxChat work with WooCommerce, and can the bot look up an order status for a logged-in customer?
  • Do I need Pinecone to use the knowledge base, or can I store embeddings in WordPress?
  • Can a visitor be handed over to a human on Slack or Telegram when the bot cannot answer?

For each question we pulled four passages from our own documentation and posts, about 16,000 characters, the way a retrieval chatbot would. Costs below come from the token counts each response reported in its usage block (fresh input, cache writes, cache reads and output), priced at the old and new cache-read rates. Every other rate is unchanged, so the old and new figures differ only on the cache line. We read all 27 replies. None contradicted the sources it was given, and where the sources did not cover something (one follow-up asked about performance under heavy traffic), the bot said so instead of guessing.

We tested three layouts, because where the cache breakpoint sits decides how much of each request can be cached at all:

LayoutWhat is cachedWhat is sent freshCalls
MxChat layout (system prompt breakpoint)System prompt, 1,212 tokensRetrieved passages + question, about 5,640 tokens9 (6 at effort low, 3 at default effort)
Conversation with automatic cachingEverything up to the latest messageThe new message and a few formatting tokens12 (3 conversations, 4 turns each)
Whole knowledge base in the system promptSystem prompt + all 12 passages, 17,946 tokensThe question, about 40 tokens6
Bar chart of Claude Sonnet 5.5 cost per 1,000 chatbot replies after the cache-read cut: MxChat layout $13.29, was $13.41; conversation follow-up $4.08, was $4.78; whole knowledge base cached $3.27, was $5.06

Layout 1: a retrieval chatbot saves 0.9%

This is how most WordPress chatbot plugins call Claude, MxChat included. The system prompt never changes, so it carries the cache breakpoint. The knowledge base passages change with every question, because retrieval picks different pages each time, so they cannot be cached and are billed as fresh input at $2 per million.

On that layout the cached part is small. Of roughly 6,850 input tokens per request, 1,212 were cache reads. At the old rate those reads cost $0.24 per 1,000 replies; at the new rate, $0.12. The whole bill is $13.29, so the cut is worth 12 cents per 1,000 replies.

Per 1,000 replies, MxChat layoutBefore 7 OctFrom 7 OctChange
Fresh input (passages + question, about 5,640 tokens)$11.28$11.28—
Cache read (system prompt, 1,212 tokens)$0.24$0.12−$0.12
Output (about 188 tokens)$1.88$1.88—
Total, effort low$13.41$13.29−0.9%
Total, default effort$13.96$13.84−0.9%

The first call of the run also wrote the system prompt to cache, which costs $2.50 per million instead of $2. On a site with steady traffic that write happens once every five minutes at most, and it is too small to matter: 1,212 tokens at $2.50 per million is a third of a cent.

Fresh input is 85% of this bill. That is the line to work on if you want a cheaper Sonnet bot: fewer or shorter retrieved passages, or a cheaper model. In our Claude Haiku 5.5 test the day before, the same three questions with the same passages cost $0.75 per 1,000 replies on Haiku 5.5 at its default settings. No change to cache pricing comes close to that.

Layout 2: conversations save 14.6% per follow-up

Automatic caching is the simplest way to cache more. You add one cache_control field at the top level of the request and Anthropic moves the breakpoint to the end of the conversation on each call. The next message then reads everything before it from cache: the system prompt, the retrieved passages from the first question, and every earlier reply.

We ran three four-turn conversations. The first message was one of the three questions above; the follow-ups were “Is that included in the free plugin, or do I need MxChat Pro?”, “Which settings page do I open first?” and “Will it slow my site down if I have a lot of traffic?”

Per 1,000 conversations, Sonnet 5.5Cache reads (tokens)Cache writes (tokens)Before 7 OctFrom 7 OctChange
Turn 11,2125,622$15.80$15.68−0.8%
Turn 26,834181$4.42$3.74−15.5%
Turn 37,015138$5.04$4.34−13.9%
Turn 47,153136$4.90$4.18−14.6%
Mean of turns 2 to 47,001152$4.78$4.08−14.6%

On follow-ups, cache reads were 29% of the bill before the cut, so halving them took off about 15%. Output is now the biggest line on a follow-up: 73% of what you pay. Follow-ups also produced longer answers than first questions here (about 300 output tokens against 150), which is why turn 3 cost more than turn 2.

The catch: automatic caching costs more on the first message

The cut makes follow-ups cheaper, but it does nothing for the first message, and automatic caching makes the first message dearer. Turn 1 wrote the whole question, passages included, to the cache: 5,622 tokens at $2.50 per million instead of $2. That costs $15.68 per 1,000 first messages, against $12.87 if only the system prompt were cached.

Bar chart of a four-turn Claude Sonnet 5.5 chatbot conversation, cost per 1,000 conversations: automatic caching $15.68, $3.74, $4.34 and $4.18 by turn, against system-prompt-only caching $12.87, $14.33, $15.29 and $15.40

So whether automatic caching pays depends on how many of your visitors ask a second question within five minutes. The extra write costs $2.81 per 1,000 conversations. Each follow-up that reads from cache saves about $10.60 per 1,000 against caching the system prompt alone ($3.74 instead of $14.33 on turn 2). That breaks even at roughly one follow-up for every four conversations. Before the cut, the saving per follow-up was about $10.00, so the break-even point has moved slightly in favour of caching, not dramatically.

If most of your chats are one question and done, as many pre-sales bots are, automatic caching costs you about 22% more per conversation. If visitors routinely ask two or three questions, it roughly halves the conversation bill. In our four-turn conversations the mean cost per turn was $6.98 with automatic caching against $14.47 with the system prompt cached alone.

The five-minute window matters. A cache read also refreshes the timer, so an active conversation stays warm, but a visitor who comes back after a coffee pays the full write again. The 1-hour cache keeps the prefix longer at twice the input price for the write.

Layout 3: a cached knowledge base saves 35%

This layout is the closest to the agentic shape Anthropic was describing: a large, fixed prompt read from cache over and over, with very little fresh input per call. Instead of retrieving passages per question, we put all twelve passages (48,415 characters) into the system prompt behind the cache breakpoint and sent only the question.

Per 1,000 replies, 17,946-token cached promptBefore 7 OctFrom 7 OctChange
Call that writes the cache$46.71$46.71—
Call that reads the cache$5.06$3.27−35.4%
Same request with no caching (computed)$37.43$37.43—

On a read, cache tokens were 71% of the old bill, so the cut took off 35%. That is the only one of our three layouts where the “about 20%” figure holds, and it beats it. Answers were as good as on the retrieval layout, and slightly faster: 1.85 seconds on average for a cached read, against 2.09 seconds for the MxChat layout.

Bar chart of cache reads as a share of the Claude Sonnet 5.5 bill before the cut: MxChat layout 1.8 percent, saving 0.9 percent; conversation follow-up 29.3 percent, saving 14.6 percent; whole knowledge base cached 70.9 percent, saving 35.4 percent; 40 percent needed for a 20 percent saving

There are two conditions, and most sites fail at least one of them.

  • The whole knowledge base has to fit. Ours was twelve passages and 17,946 tokens. A site with hundreds of documentation pages, products or posts cannot put them all in every request, which is why retrieval exists. This layout suits a small, stable knowledge base: one product, a short FAQ, a booking policy.
  • Traffic has to keep the cache warm. A call that writes the cache costs $46.71 per 1,000, 25% more than not caching at all. With a 5-minute cache the write pays for itself after a single read, but if questions arrive less often than every five minutes, almost every call is a write. The 1-hour cache writes at $4 per million, about $73 per 1,000 calls, and needs a little over one read per write to break even.

Where the money goes, by layout

The same five lines make up every Claude bill. The share each one takes is what decides whether a cache price cut reaches you.

Share of the bill, new ratesFresh inputCache writesCache readsOutputSaving from the cut
MxChat layout, new question85%0%1%14%0.9%
Conversation follow-up0%9%17%73%14.6%
Whole knowledge base, cached read2%0%55%43%35.4%

If you want to estimate your own saving, look at the usage block of a few real requests. Multiply cache_read_input_tokens by $0.20 per million, divide by the total cost of the request, and halve it. That is what the 7 October change saves you on Sonnet 5.5.

What to change on a WordPress chatbot

If your bot retrieves passages per question, the cut is a rounding error. Leave your settings alone. If Sonnet 5.5’s bill is the problem, the levers are fresh input and model choice: fewer passages per question, shorter passages, or a smaller model for routine questions. Our earlier OpenAI prompt caching test found the same thing from the other side: caching only pays when the cached part of the request is large and repeated.

If visitors hold real conversations, automatic caching is worth testing, and the cut makes it a little more attractive. Check how many of your chats have a second message within five minutes. Above about one in four, turn it on.

If your knowledge base is small and your traffic is steady, putting the whole thing in a cached system prompt is now the cheapest way to run Sonnet 5.5: $3.27 per 1,000 replies on our content, a quarter of the retrieval layout’s $13.29. Check the token count first, and remember that each cold call costs more than no cache at all.

If you run Opus 5.5, nothing changed. Its cache reads were already 5% of input at launch, $0.20 per million; our Opus 5.5 pricing test covers the rest of its rate card.

MxChat connects to Claude, OpenAI, Gemini and other providers with your own API key, and you pay the provider directly with no markup on tokens. The MxChat documentation covers model selection and knowledge base setup, and MxChat Pro adds the WooCommerce, live-agent and knowledge-base add-ons the test questions were about.

FAQ

How much does Claude prompt caching cost?

A 5-minute cache write costs 1.25 times the model’s input price and a 1-hour write costs twice the input price. A cache read costs 10% of the input price on most Claude models, 5% on Sonnet 5.5 and Opus 5.5, and 2.5% on Fable 5.1. On Sonnet 5.5 that is $2.50 and $4 per million tokens for writes and $0.10 for reads, as listed by Anthropic on 9 October 2026.

What changed in Claude Sonnet 5.5 pricing on 7 October 2026?

Only the cache-read price, which fell from $0.20 to $0.10 per million tokens. Input stays at $2, output at $10, the 5-minute cache write at $2.50 and the 1-hour write at $4 per million tokens.

Does the Sonnet 5.5 cache cut make a chatbot 20% cheaper?

Only if cache reads were about 40% of its bill. On our tests a retrieval chatbot saved 0.9% per new question, a conversation with automatic caching saved 14.6% per follow-up, and a bot with its whole 17,946-token knowledge base in a cached prompt saved 35.4% per reply.

Should I turn on automatic prompt caching for my chatbot?

If visitors often ask follow-up questions, yes. Automatic caching writes the first message to cache at 1.25 times the input price, which cost $2.81 more per 1,000 conversations in our test, and then saves about $10.60 per 1,000 on each follow-up. It breaks even at roughly one follow-up for every four conversations, as long as the follow-up arrives within the five-minute cache window.

Is Sonnet 5.5 caching now cheaper than OpenAI’s?

It is level with GPT-6.1 Sol: both list at $2 input and $10 output, and both now charge $0.10 per million for cached input. Both now bill for writing to cache as well (our OpenAI prompt caching test covers OpenAI’s side), and which works out cheaper also depends on how many tokens each model counts for the same text.

How do I check how much of my bill is cache reads?

Every Messages API response includes a usage block with input_tokens, cache_creation_input_tokens, cache_read_input_tokens and output_tokens. Price each at its rate. The share that comes from cache reads, halved, is what the 7 October cut saves you on Sonnet 5.5.

Similar Posts