Cover card: Claude Fable 5.1 pricing, the 75% cache cut tested on a chatbot. 1.4% saved on a new question, 33% on a follow-up turn, 15 follow-ups to reach 25%

Claude Fable 5.1 Pricing: The 75% Cache Cut, Tested

Claude Fable 5.1 pricing is $10 per million input tokens and $50 per million output tokens, the same as Claude Fable 5. One number changed when Anthropic released it on 1 September 2026: reading from the prompt cache now costs $0.25 per million tokens instead of $1.00, a 75% cut. Anthropic says that makes typical workloads around 25% cheaper than on Fable 5, and highly agentic ones up to around 45% cheaper.

Those figures come from Anthropic’s own measurement of customer usage in August 2026. We wanted the number for a WordPress chatbot, so we sent Claude Fable 5.1 and Fable 5 the same chatbot requests from our production server, 45 calls in all, and priced what the API billed.

  • A new question cost 1.4% less. $64.08 per 1,000 replies on Fable 5.1 against $64.99 for the same tokens at Fable 5’s rates. The only thing read from the cache was a 1,212-token system prompt.
  • A follow-up turn cost 33% less. $10.54 against $15.66 per 1,000 replies, when the whole first turn was read from the cache.
  • Reaching Anthropic’s 25% took about 15 follow-ups per retrieved question. Support chats are rarely that long.
  • Fable 5.1 was slower. Mean response time was 5.91 seconds against 3.65 seconds for Fable 5 on the same requests.
  • It cost 4.8 times what Claude Sonnet 5.5 did. Sonnet 5.5 answered the same questions correctly for $13.24 per 1,000 replies in about 2 seconds.

Claude Fable 5.1 pricing, next to the models around it

These are the list rates per million tokens as they read on Anthropic’s pricing page on 2 October 2026.

Per 1M tokensClaude Fable 5.1Claude Fable 5Claude Opus 5.5Claude Sonnet 5.5
Input$10.00$10.00$4.00$2.00
Cache read$0.25$1.00$0.20$0.20
Cache write, 5 minutes$12.50$12.50$5.00$2.50
Cache write, 1 hour$20.00$20.00$8.00$4.00
Output (includes thinking)$50.00$50.00$20.00$10.00
Cache read as a share of input2.5%10%5%10%

Input, output and both cache-write rates match Fable 5 to the cent. The Batch API halves input and output to $5 and $25. The API’s model record lists a 1,000,000-token context window and a 128,000-token output limit for both Fable models, and Anthropic’s caching documentation gives both a 512-token minimum for a cacheable prompt. Claude Mythos 5.1 carries the same rates but is offered by invitation only, so Fable 5.1 is the model most accounts can actually call.

So the whole price change is one row. A cache read fell from 10% of the input rate to 2.5%, the lowest ratio on Anthropic’s list. Opus 5.5 made the same kind of move a few weeks later, to 5%; we measured that one in our Claude Opus 5.5 pricing test.

Where Anthropic’s 25% comes from

When only the cache-read rate changes, the saving is simple arithmetic. Every dollar you spent on cache reads under Fable 5 becomes 25 cents. Nothing else moves. So the saving on your whole bill is 75% of whatever share of that bill was cache reads.

Cache reads as a share of your Fable 5 billSaving on Fable 5.1
2%1.5%
10%7.5%
33%25% (Anthropic’s typical figure)
44%33%
60%45% (Anthropic’s agentic figure)

Read that way, Anthropic’s claim says something about its customers: on a typical Fable workload a third of the bill was cache reads, and on agent-heavy ones it was closer to 60%. That fits coding agents and long tool loops, which resend a large unchanged history on every step. The question for a site owner is what share of a chatbot’s bill is cache reads. We measured it.

How we tested

Every call went from the mxchat.ai WordPress server to Anthropic’s Messages API with the site’s own API key, which never left the server. Each request was built like a knowledge-base chatbot request: the live system prompt for our own site bot, then a user message carrying four retrieved documentation excerpts (about 16,000 characters) and the visitor’s question. The three questions came from our own chat logs:

  1. Does MxChat work with WooCommerce, and can the bot look up an order status for a logged-in customer?
  2. Do I need Pinecone to use the knowledge base, or can I store embeddings in WordPress?
  3. Can a visitor be handed over to a human on Slack or Telegram when the bot cannot answer?

These are the same requests, character for character, that we used in our GPT-6.1 Sol test the day before. We ran two layouts.

Layout 1, cached system prompt. A cache breakpoint sits on the system prompt and nothing else. The retrieved excerpts and the question go in the user message and are billed as ordinary input. This is how MxChat builds a Claude request. We ran Fable 5.1 with no effort setting, at medium and at low, and Fable 5 with no effort setting and at low, 27 calls. Then we ran Claude Sonnet 5.5 at medium and Claude Opus 5.5 at its default on the same three questions as reference points, 6 calls.

Layout 2, automatic caching. A single cache_control field at the top of the request, which tells the API to cache everything up to the last message. For each question we sent the first turn, appended the model’s answer, and asked a second question in the same conversation: “Is that included in the free plugin, or do I need MxChat Pro for it?” That was 12 calls, six per Fable model.

The token counts barely differed between the two models. Fable 5.1 counted the three requests as 6,767, 6,850 and 6,897 input tokens, and Fable 5 counted two fewer each time. Opus 5.5 and Sonnet 5.5 returned exactly the Fable 5.1 counts. So there is no tokenizer change hiding in this release, unlike the one we covered in our Claude tokenizer breakdown.

What Fable 5.1 accepts and rejects

Nine short probe calls checked which request parameters still work, because a rejected parameter means a broken chatbot rather than a dearer one.

Parameter sent to claude-fable-5-1Result
No extras200
temperature: 0.8400, “deprecated for this model”
top_p: 0.9400, “deprecated for this model”
thinking enabled with a token budget400, adaptive thinking only
thinking disabled400, adaptive thinking only
Effort low, max200
Effort minimal, none400; allowed values are low, medium, high, xhigh, max

None of this is new with 5.1. The API’s model record lists the same capabilities for Fable 5, and Anthropic deprecated temperature on Claude 4.7 and later. But it matters for any plugin that sends a fixed temperature to every model. You also cannot switch thinking off. In practice that cost nothing here: Fable 5.1 used no thinking tokens on any of its 21 first-turn replies.

What a new question costs

This is the request a retrieval chatbot sends most: a visitor asks something, the plugin fetches matching content, and the model answers. Prices below assume the system prompt is already in the cache, which is the best case for both models.

Model and effortCallsOutput tokensCost per 1,000 repliesMean response time
Claude Fable 5.1, default6150$64.085.43 s
Claude Fable 5.1, medium6163$64.706.93 s
Claude Fable 5.1, low6137$63.395.37 s
Claude Fable 5, default6144$64.653.77 s
Claude Fable 5, low3108$62.853.41 s
Claude Opus 5.5, default3330$29.345.47 s
Claude Sonnet 5.5, medium3175$13.242.01 s

Bar chart of cost per 1,000 chatbot replies to new questions: Claude Fable 5.1 default $64.08, medium $64.70, low $63.39, Claude Fable 5 default $64.65, Claude Opus 5.5 $29.34 and Claude Sonnet 5.5 $13.24

Across all 18 Fable 5.1 calls the average was $64.06 per 1,000 replies. Across the 9 Fable 5 calls it was $64.05. On a new chatbot question the two models cost the same.

The reason is in the usage the API returned. Of roughly 6,840 input tokens per request, 1,212 were the system prompt, read from the cache. The other 5,600 were retrieved excerpts and the question, which are different for every visitor and were billed at the full $10 rate. On Fable 5, those 1,212 cached tokens cost $1.21 per 1,000 replies. On Fable 5.1 they cost 30 cents. That 91-cent difference is the entire effect of the price cut on this request: 1.4% of the bill.

The small differences between rows are output. Fable 5.1 wrote slightly longer answers (64 words on average against 51), and at $50 per million every extra 20 output tokens adds a dollar per 1,000 replies. That is enough to cancel the cache saving, which is why the measured totals sit within a cent of each other.

The system prompt cache has to stay warm

Caching the system prompt is worth far more than the 5.1 price cut, on either model. But a 5-minute cache entry expires if nobody chats for five minutes, and the next request pays to write it again at 1.25 times the input rate.

Per 1,000 new-question repliesFable 5.1Fable 5 rates, same tokens
No caching at all$75.90$75.90
System prompt read from cache (warm)$64.08$64.99
System prompt written to cache (cold)$78.93$78.93
Warm against no caching16% less14% less
Cold against no caching4% more4% more
Share of requests that must hit a warm cache to break even20.4%21.7%

If more than about one request in five arrives within five minutes of the previous one, the breakpoint pays for itself. A busy site clears that easily. A site with a dozen chats a day mostly does not, and pays the 4% write premium on most of them. The 1-hour cache costs twice the input rate to write, so it needs about half of all requests to land on a warm entry before it beats no caching.

One more observation from the logs. Of 27 calls in this layout, 24 read the system prompt from the cache and 3 wrote it. Two of those writes were expected, the first call to each model. The third was a Fable 5 call that wrote the prompt again seconds after another call had read it. A cache hit is likely, not guaranteed.

Follow-up turns are where the cut is real

With automatic caching the picture changes on the second turn. The first request writes the whole prompt to the cache. When the visitor asks a follow-up, the API reads everything from the first turn and only writes the new answer and question.

Our six follow-up requests averaged 6,834 tokens read from the cache, 156 written, 3 uncached and 137 tokens of output. Priced on both rate cards:

Per 1,000 follow-up repliesFable 5.1Fable 5
Cache read, 6,834 tokens$1.71$6.83
Cache write, 156 tokens$1.95$1.95
Uncached input, 3 tokens$0.03$0.03
Output, 137 tokens$6.85$6.85
Total$10.54$15.66

That is 33% less, and it is the shape Anthropic’s claim describes. Cache reads were 44% of the Fable 5 bill for this turn, and three quarters of that went away. The same follow-up with no caching would cost $76.78, and with only the system prompt cached, $64.96.

There is a caveat in the bills as they actually arrived. Fable 5.1’s three follow-up replies averaged 183 output tokens, because one of them spent 174 tokens thinking. Fable 5’s averaged 91. As billed, both models’ follow-up turns came to $13.10 per 1,000 replies, identical to the cent. The cheaper reads were real, and longer output spent them. At $50 per million, output is the line to watch on this model.

Automatic caching also has a cost on the first turn. Writing the retrieved excerpts to the cache raised a new question from $62.90 to $76.96 per 1,000 replies on the same tokens, 22% more. That premium is repaid if about one conversation in four goes on to a follow-up that can be answered from the same excerpts. The same trade-off exists on OpenAI’s side, with different numbers; see our OpenAI prompt caching test.

How many follow-ups it takes to reach 25%

A conversation is one new question plus some number of follow-ups. Using the measured token counts for each, here is what the move from Fable 5 to Fable 5.1 saves at different conversation lengths, with automatic caching on.

Follow-ups per retrieved questionFable 5, per 1,000 repliesFable 5.1, per 1,000 repliesSaving
0$77.87$76.961.2%
1$46.76$43.756.5%
2$36.40$32.6810.2%
4$28.10$23.8215.2%
9$21.88$17.1821.5%
19$18.77$13.8626.2%
49$16.91$11.8629.8%

Bar chart of the saving from Claude Fable 5 to Fable 5.1 by conversation length: 1.2% with no follow-ups, 6.5% with one, 10.2% with two, 15.2% with four, 21.5% with nine, 26.2% with nineteen and 29.8% with forty-nine

The saving crosses 25% at about 15 follow-ups per retrieved question and levels off near 33%, because every turn still pays for its output and for writing the new messages. Getting to 45% needs a workload where cache reads are 60% of the bill: very large prefixes, short outputs, many steps. That describes an agent working through a codebase. A visitor asking about shipping does not produce it.

The table also shows something more useful than the percentage. Per reply, a long cached conversation on Fable 5.1 is cheap relative to its first turn: $13.86 per 1,000 replies at 19 follow-ups against $76.96 for a lone question. Cache reads at 2.5% of input make history nearly free to resend.

Speed and answer quality

Bar chart of mean response time for the same chatbot request: Claude Fable 5.1 default 5.43 seconds, medium 6.93, low 5.37, Claude Fable 5 default 3.77, Claude Opus 5.5 default 5.47 and Claude Sonnet 5.5 medium 2.01 seconds

Fable 5.1 was the slower of the two Fable models on every configuration. Its 18 new-question calls averaged 5.91 seconds, with a range of 4.04 to 8.77. Fable 5’s nine averaged 3.65 seconds, range 2.84 to 4.19. Lowering the effort setting did not help: low took 5.37 seconds and the default 5.43. These are full round trips without streaming, from one server on one morning, so treat the gap as indicative rather than fixed.

The two models also reason differently on a short support answer. Fable 5 used a few thinking tokens (6 to 39) on 10 of its 15 calls. Fable 5.1 used none on 23 of 24, then 174 on the remaining one. The extra response time on Fable 5.1 is not explained by billed thinking.

We read all 45 replies. Every one answered the question from the supplied excerpts, and none refused. All of them got the facts right: order lookup needs the WooCommerce add-on and a logged-in customer, Pinecone is optional, and Slack and Telegram handoff ship in the free plugin. Sonnet 5.5 and Opus 5.5 did the same. On questions the knowledge base covers, the answers do not separate these four models. Price and speed do.

What MxChat does on the Claude path

MxChat 3.2.22 builds Claude requests the way Layout 1 does. The system prompt gets a cache breakpoint, the retrieved content goes in the last user message, and each turn retrieves again instead of carrying earlier excerpts forward. So on MxChat the Fable 5.1 price cut applies to the system prompt only, which is the 1.4% case. There is one exception: if your bot instructions embed the {context} placeholder, the plugin skips the breakpoint, because a prompt that changes on every question would pay the write premium and never be read.

The model picker in 3.2.22 lists Claude Fable 5 and does not list Fable 5.1 yet. Setup for the Claude models that are listed is covered in the MxChat documentation.

Should a WordPress chatbot run on Claude Fable 5.1?

For answering visitor questions from a knowledge base, no. At 3,000 replies a month, the request we tested costs about $192 on Fable 5.1, $88 on Opus 5.5 and $40 on Sonnet 5.5, and all three gave correct answers. Sonnet 5.5 did it in a third of the time. The detail on that model is in our Claude Sonnet 5.5 pricing test, and our GPT-6 Sol comparison shows what the same request costs on OpenAI’s side.

If you already run Fable 5 for a reason, such as a bot that does multi-step tool work or long document analysis, moving to Fable 5.1 cannot raise your rates. Check two things first. Compare response times on your own requests, since ours were about 60% longer. And watch output length, because a model that writes 20 more tokens per reply gives back everything the cache cut saved on a retrieval question.

If you are building something agent-shaped, with a large fixed context and many steps, Fable 5.1’s cache rate is a real advantage. Structure the request so the stable part comes first, turn on automatic caching, and keep traffic frequent enough that the cache stays warm. That is where Anthropic’s 25 to 45% lives.

For most site chatbots the cheaper route to a lower bill is a smaller model and a tight system prompt. MxChat Pro is a one-time licence and you pay the model provider directly at these list rates, so choosing the model is the pricing decision.

FAQ

How much does Claude Fable 5.1 cost?

$10 per million input tokens, $50 per million output tokens, $0.25 per million cache-read tokens, $12.50 per million for a 5-minute cache write and $20 for a 1-hour cache write, as listed on 2 October 2026. The Batch API is half price on input and output. Thinking tokens are billed as output.

Is Claude Fable 5.1 cheaper than Claude Fable 5?

Only on cache reads, which are $0.25 instead of $1.00 per million. Every other rate is identical. In our chatbot test that saved 1.4% on a new question and 33% on a follow-up turn read from the cache. Anthropic’s figure of around 25% applies when about a third of your bill is cache reads.

What is the minimum prompt size for caching on Claude Fable 5.1?

Anthropic’s documentation lists 512 tokens for Fable 5.1 and Fable 5. Our 1,212-token system prompt cached on both. Shorter prompts are processed normally and the cache marker is ignored without an error.

Is Claude Fable 5 being retired?

Not yet. Anthropic’s deprecations page lists claude-fable-5 as active with a retirement date no sooner than 9 June 2027, and claude-fable-5-1 no sooner than 1 September 2027. Both were callable side by side on 2 October 2026.

Does Claude Fable 5.1 accept a temperature setting?

No. The API returned HTTP 400 for temperature and top_p, and for any attempt to enable or disable thinking manually. Reasoning depth is controlled with the effort setting, which accepts low, medium, high, xhigh and max.

Is Claude Fable 5.1 worth it for a customer support chatbot?

Not for knowledge-base answers. It cost about $64 per 1,000 replies in our test against $13 for Claude Sonnet 5.5, took nearly three times as long, and the answers were equally correct. It earns its price on long, multi-step work where its cheap cache reads cover most of the tokens.

Method note: 45 Messages API calls, 9 parameter probes and 16 token-count calls from the mxchat.ai server on 2 October 2026, max_tokens 1,000, no streaming, 5-minute cache. Prices are list rates read from Anthropic’s pricing page the same day. “Warm” prices 1,212 system-prompt tokens as a cache read and the rest of the prompt as input, matching the usage the API reported on 24 of 27 calls. Conversation figures price the mean token counts of the six first-turn and six follow-up calls on each rate card. Response times are full round trips from one server on one morning.

Similar Posts