Claude Haiku 4.5 Pricing: Half the Rate, 38% of the Bill
Claude Haiku 4.5 costs $1 per million input tokens and $5 per million output tokens, exactly half of Claude Sonnet 5.5’s $2 and $10. On a WordPress chatbot it costs a lot less than half. We sent both models the same support questions from our production server on 4 October 2026, and Haiku 4.5 came to $4.93 per 1,000 replies against $13.02 for Sonnet 5.5. That is 38% of the Sonnet bill, not 50%.
The gap comes from somewhere the price page only mentions in passing. Haiku 4.5 still uses Anthropic’s older tokenizer, so it counts the same text as fewer tokens. Claude Haiku 5.5 was announced on 28 September with no price and no release date, and the model ID returned 404 when we tried it. Until it ships, Haiku 4.5 is the cheapest Claude you can put behind a chatbot, and this is what Claude Haiku 4.5 pricing works out to in practice.
- Haiku 4.5 cost $4.93 per 1,000 chatbot replies. Sonnet 5.5 cost $13.02 at effort
lowand $13.75 atmedium, with its system prompt read from cache. - Haiku 4.5 counted the identical request as 4,550 tokens; Sonnet 5.5 counted 6,838. That 1.5x tokenizer gap, plus shorter answers, is why the bill ratio is 0.38 and not 0.5.
- The usual system-prompt cache does nothing on Haiku 4.5. Haiku will not cache a prefix shorter than 4,096 tokens. Our system prompt is 836 tokens on Haiku, so the cache breakpoint was ignored on all six calls.
- Automatic caching makes follow-up questions very cheap: $0.82 per 1,000 Haiku follow-ups against $6.10 for the first turn that writes the cache.
- Haiku 4.5 was faster (1.51 seconds against 2.16) but ignored the sales instructions in our system prompt. Sonnet used the scripted discount offer in 14 of 18 replies; Haiku used it in none of 15.
Claude Haiku 4.5 pricing next to the rest of the Claude line-up
These are the standard per-million-token rates from Anthropic’s pricing page, read on 4 October 2026. Output includes any thinking tokens.
| Per 1M tokens | Claude Haiku 4.5 | Claude Sonnet 5.5 | Claude Opus 5.5 | Claude Fable 5.1 |
|---|---|---|---|---|
| Input | $1 | $2 | $4 | $10 |
| Output | $5 | $10 | $20 | $50 |
| 5-minute cache write | $1.25 | $2.50 | $5 | $12.50 |
| 1-hour cache write | $2 | $4 | $8 | $20 |
| Cache read | $0.10 | $0.20 | $0.20 | $0.25 |
| Batch input / output | $0.50 / $2.50 | $1 / $5 | $2 / $10 | $5 / $25 |
| Shortest prefix it will cache | 4,096 tokens | 512 tokens | 512 tokens | 512 tokens |
| Tokenizer | Older | Newer | Newer | Newer |
effort parameter | Not supported (HTTP 400) | Supported | Supported | Supported |
Three rows matter more than the headline rates. The tokenizer row decides how many tokens you are charged for. The cache-minimum row decides whether your cache does anything. And the effort row matters if your plugin sends that parameter: Haiku 4.5 answered output_config.effort: low with “This model does not support the effort parameter.”
What we know about Claude Haiku 5.5
Release trackers report that Anthropic said on 28 September 2026 that Haiku 5.5 would join the Claude 5.5 family in the coming weeks; we could not find a dated Anthropic page saying so. As of 4 October there is no published price, context window or model ID. On our key, claude-haiku-5-5 returned HTTP 404, and the models list held 13 Claude models with Haiku 4.5 as the only Haiku. Price rumours are circulating; we found no Anthropic source for any of them.
The thing to check on launch day is not just the rate card. Every Claude model from 4.7 onwards uses the newer tokenizer, which Anthropic says “produces approximately 30% more tokens for the same text”. On our chatbot text it produced 50% more. If Haiku 5.5 moves to the newer tokenizer at Haiku 4.5’s price, a chatbot’s bill rises by about half before anything else changes. We will run the same test the day it is callable. Our Claude tokenizer test covers how the change shows up on a bill.
How we tested
Every call went from the mxchat.ai WordPress server to Anthropic’s Messages API with the site’s own API key, the way the MxChat plugin sends them. Each request carried:
- the site’s live chatbot system prompt (3,478 characters), with a cache breakpoint on the system block, which is how the plugin builds Anthropic requests;
- four retrieved pages from our own documentation and blog, about 16,000 characters, in the user message;
- one of three real visitor questions: WooCommerce order lookup, whether Pinecone is required, and handing a visitor to a human on Slack or Telegram.
Haiku 4.5 ran at its default setting and once with extended thinking (budget 1,024 tokens). Sonnet 5.5 ran at effort low and medium. Each configuration answered all three questions twice. We then ran three two-turn conversations per model with automatic caching. That is 33 calls, all of which returned HTTP 200. We also counted tokens with count_tokens and probed parameters with eight short calls. max_tokens was 1,000 and nothing was streamed. These are the same system prompt, sources and questions we used for Sonnet 5.5, Fable 5.1 and the budget models below.
Cost per 1,000 chatbot replies
| Configuration | Tokens in (mean) | Tokens out (mean) | Cost per 1,000 replies | Mean time |
|---|---|---|---|---|
| Haiku 4.5, default | 4,550 (none cached) | 77 | $4.93 | 1.51 s |
| Haiku 4.5, thinking budget 1,024 | 4,580 | 372 (313 thinking) | $6.44 | 4.83 s |
| Sonnet 5.5, effort low | 6,838 (1,212 read from cache) | 153 | $13.02 | 2.16 s |
| Sonnet 5.5, effort medium | 6,838 (cache mostly read) | 179 | $13.75 | 2.44 s |
Costs are as billed: we priced the input, cache-write, cache-read and output token counts each response reported at Anthropic’s list rates. Input was 92% of the Haiku bill and 88% of the Sonnet bill. On a retrieval chatbot the knowledge base you send with each question is most of what you pay for, which is why the tokenizer matters so much.
Where the 38% comes from
Split the gap into its parts. If Haiku 4.5 counted text the way Sonnet 5.5 does, the same request would be 6,838 input tokens and the reply would cost about $7.22 per 1,000, which is 55% of Sonnet’s bill. The older tokenizer takes that down to $4.93, a further 32% off. Haiku’s answers were also half as long (77 tokens against 153), and output is the expensive side of the card.
| Step | Haiku 4.5 per 1,000 replies | Share of Sonnet 5.5’s $13.02 |
|---|---|---|
| Rate card alone (half of Sonnet) | about $6.51 | 50% |
| Haiku prices, Sonnet’s token count, Haiku’s answers | about $7.22 | 55% |
| As billed, Haiku’s own token count | $4.93 | 38% |
The middle row is higher than the first because Sonnet reads its system prompt from cache at $0.20 and Haiku cannot. That cache saving is small next to the tokenizer.
Why the system-prompt cache does nothing on Haiku 4.5
Anthropic’s prompt caching docs set a minimum length for a cached prefix: 512 tokens on the Claude 5 and 5.5 models, 4,096 tokens on Haiku 4.5. A prefix shorter than that is processed at the normal input rate with no error and no warning. The response simply reports zero cache-write and zero cache-read tokens.
Our system prompt is 836 tokens on Haiku 4.5, so all six Haiku calls in the main test wrote nothing and read nothing. On Sonnet 5.5 the same prompt is 1,221 tokens, above the 512-token minimum, and it was read from cache on most calls. If your chatbot has a short system prompt and you switch to Haiku 4.5, the cache line on your bill drops to zero. That is expected, not a fault in your plugin.
When could the prefix reach 4,096 tokens? Tool definitions sit in front of the system prompt in Anthropic’s cache order, and Haiku 4.5 adds 496 tokens of tool-use instructions whenever tools are offered. A bot with several actions and a long system prompt can get past the limit. Our test sent no tools, so we have not measured that case.
Follow-up questions: where Haiku 4.5’s cache pays off
With automatic caching (a cache_control field at the top level of the request), the cache breakpoint moves to the end of the conversation. The first turn is now well past 4,096 tokens, so Haiku caches all of it. We asked each model one question, then a follow-up two seconds later: “Is that included in the free plugin, or do I need MxChat Pro for it?”
| Per 1,000 replies, as billed | Haiku 4.5 | Sonnet 5.5, low |
|---|---|---|
| First turn, no automatic cache (main test) | $4.93 | $13.02 |
| First turn, automatic cache (writes it) | $6.10 | $15.99 |
| Follow-up turn (reads it) | $0.82 | $4.74 |
| Follow-up turn at full input price | $4.89 | $16.94 |
Writing the cache adds 24% to Haiku’s first turn, about $1.17 per 1,000 conversations. Each follow-up inside five minutes saves about $4.07. So automatic caching pays for itself on Haiku 4.5 once about 30% of conversations include a follow-up. On a support bot most do. If your visitors usually ask one question and leave, it costs you money.
Sonnet’s follow-ups thought before answering (162 thinking tokens on average at effort low), which is why its follow-up is still $4.74. Haiku 4.5 does not think unless you switch it on.
Haiku 4.5 against the budget models from other vendors
Haiku 4.5 is cheap for a Claude model. It is not cheap next to the budget tiers from OpenAI, Google and DeepSeek. We sent each of these the same system prompt, sources and questions from the same server in the last ten days.
| Model (tested) | List price per 1M (in / out) | Cost per 1,000 replies | Our test |
|---|---|---|---|
| GPT-6 Luna, effort none (28 Sep) | $0.10 / $0.50 | $0.57 | GPT-6 Luna pricing |
| DeepSeek V4.1 Flash, off-peak / peak (27 Sep) | $0.15 / $0.60 off-peak | $0.77 / $1.53 | DeepSeek Flash pricing |
| Gemini 3.1 Flash-Lite (3 Oct) | $0.25 / $1.50 | $1.12 | Gemini Flash-Lite pricing |
| Gemini 3.5 Flash-Lite (3 Oct) | $0.30 / $2.50 | $1.35 | same |
| Claude Haiku 4.5 (4 Oct) | $1 / $5 | $4.93 | this page |
| Claude Sonnet 5.5, low (4 Oct) | $2 / $10 | $13.02 | Sonnet 5.5 pricing |
Haiku 4.5 costs about 3.6 to 8.6 times as much per reply as these four. In money that is still small. At 3,000 replies a month, Haiku 4.5 is about $15, GPT-6 Luna under $2, and Sonnet 5.5 about $39. If you already run Claude and want the cheapest Claude, Haiku 4.5 is it. If price per reply is the only thing you care about, it is not the cheapest model a WordPress chatbot can use.
What it costs per month
| Replies per month | Haiku 4.5 | Haiku 4.5 + cached follow-ups (half the replies are follow-ups) | Sonnet 5.5, low |
|---|---|---|---|
| 1,000 | $4.93 | $3.46 | $13.02 |
| 3,000 | $14.79 | $10.38 | $39.06 |
| 10,000 | $49.30 | $34.60 | $130.20 |
| 50,000 | $246.50 | $173.00 | $651.00 |
The middle column assumes every conversation is one question plus one follow-up within five minutes, with automatic caching on: $6.10 for the first turn and $0.82 for the second, averaged. Your own mix will differ. Longer knowledge-base excerpts raise every column in proportion, because input is most of the bill.
Speed
Haiku 4.5 answered in 1.51 seconds on average (1.32 to 1.86), against 2.16 seconds for Sonnet 5.5 at low and 2.44 at medium. Cached follow-ups were faster still: 0.97 seconds on Haiku. With extended thinking on, Haiku slowed to 4.83 seconds (3.60 to 6.96) and spent 313 thinking tokens per reply to produce answers that were shorter, not better. For a site chatbot we would leave thinking off.
Answer quality
We read all 33 replies. Both models got the facts right on the three first-turn questions every time: the WooCommerce add-on needs MxChat Pro and shows up to five recent orders to logged-in customers; Pinecone is optional and embeddings can live in the WordPress database; Slack and Telegram handoff is in the free plugin.
The differences were elsewhere:
- Haiku 4.5 ignored the sales script. Our system prompt asks the bot to offer a discount code. Sonnet 5.5 did so in 14 of 18 replies; Haiku 4.5 in 0 of 15. Haiku wrote its own calls to action a few times instead (“I can add MxChat Pro to your cart”). If your chatbot has a sales job, test this before you switch.
- Haiku 4.5 was more confident and once wrong. Asked whether the knowledge base is in the free plugin, it said yes (correct) and then listed “Slack integration” as a Pro add-on, which contradicts its own earlier answer. Sonnet 5.5 said it did not have enough information to answer that follow-up. One error in 15 Haiku replies is not a verdict, but it is the kind a visitor notices.
- Haiku 4.5 accepted a loaded question. Asked whether a visitor can be handed to a human “when the bot cannot answer”, Haiku said yes. Sonnet 5.5 at
mediumpointed out that our documentation describes handoff when the visitor asks, not automatic handoff. - Haiku’s answers were shorter: 47 words on average against 69 to 79 for Sonnet. For a chat widget that is mostly a good thing.
Should a WordPress chatbot use Claude Haiku 4.5?
| Your situation | What we would do |
|---|---|
| Support or FAQ bot, cost matters, you want Claude | Haiku 4.5, thinking off, automatic caching on if visitors ask follow-ups |
| Sales bot that must follow a script or offer codes | Sonnet 5.5 at low; Haiku 4.5 ignored our script |
| Lowest possible cost per reply, any vendor | GPT-6 Luna or DeepSeek V4.1 Flash off-peak |
| Short system prompt, explicit system-block caching | Expect no cache saving on Haiku 4.5 below 4,096 tokens |
| Waiting for Haiku 5.5 | Check its tokenizer and cache minimum on launch day, not just its price |
In MxChat, Claude Haiku 4.5 is in the model list under its full ID, claude-haiku-4-5-20251001, and the plugin already uses it as the default Claude model for image analysis. Setup is in the MxChat documentation; the multi-bot and WooCommerce features we asked about are in MxChat Pro.
FAQ
How much does Claude Haiku 4.5 cost?
$1 per million input tokens and $5 per million output tokens on Anthropic’s API. Cache writes are $1.25 (5 minutes) or $2 (1 hour) per million, cache reads $0.10, and the Batch API halves input and output to $0.50 and $2.50. In our WordPress chatbot test it came to $4.93 per 1,000 replies.
Is Claude Haiku 4.5 cheaper than Sonnet 5.5?
Yes, by more than the rate card suggests. The list price is half, but Haiku 4.5 counted our chatbot request as 4,550 tokens against 6,838 on Sonnet 5.5 and wrote shorter answers, so it cost 38% of Sonnet’s bill: $4.93 against $13.02 per 1,000 replies.
Has Claude Haiku 5.5 been released?
Not as of 4 October 2026. It was announced on 28 September without a price or date, and the model ID claude-haiku-5-5 returned HTTP 404 on our API key. Haiku 4.5 is the only Haiku in the API’s model list.
Why does prompt caching not work on my Haiku 4.5 chatbot?
Haiku 4.5 only caches prefixes of 4,096 tokens or more. A shorter system prompt is billed at the normal input rate, and the response reports zero cache tokens. Either the prefix gets longer (tools and system prompt together), or you use automatic caching so the whole conversation is cached from the second turn.
Does Claude Haiku 4.5 support the effort parameter?
No. A request with output_config.effort returned HTTP 400, “This model does not support the effort parameter.” It does support extended thinking with a token budget, which is off by default.
Does Haiku 4.5 use a different tokenizer?
Yes. Anthropic says Claude 4.7 and later models use a newer tokenizer that produces about 30% more tokens for the same text. Haiku 4.5 predates it. On our chatbot request the newer tokenizer counted 1.5 times as many tokens.
Method note: 33 Messages API calls, 8 parameter probes, 12 count_tokens calls and 1 models list call from the mxchat.ai server on 4 October 2026, max_tokens 1,000, no streaming. Prices are list rates read from Anthropic’s pricing page the same day; cache minimums from Anthropic’s prompt caching documentation. Costs price the input_tokens, cache_creation_input_tokens, cache_read_input_tokens and output_tokens each response reported. The budget-model rows were measured on the dates shown with the same system prompt, sources and questions. Response times are full round trips from one server on one morning.