Claude Haiku 5.5 pricing cover: $0.75 per 1,000 chatbot replies on Haiku 5.5, $4.98 on Haiku 4.5 and $0.53 on GPT-6 Luna, measured October 2026

Claude Haiku 5.5 Pricing: 85% Cheaper, Not Luna-Cheap

Claude Haiku 5.5 costs $0.10 per million input tokens and $0.50 per million output tokens, a tenth of what Claude Haiku 4.5 charges. That is also exactly what OpenAI charges for GPT-6 Luna, so on paper the two cheapest models from the two biggest labs now cost the same. On a WordPress chatbot they do not. We sent the same support questions to Haiku 5.5, Haiku 4.5 and GPT-6 Luna from our production server on 8 October 2026, the day after Haiku 5.5 went live, and Haiku 5.5 came to $0.75 per 1,000 replies, against $4.98 for Haiku 4.5 and $0.53 for GPT-6 Luna.

So Claude Haiku 5.5 pricing is a real cut, about 85% off the Haiku 4.5 bill on identical requests. It is not a tie with GPT-6 Luna, and the reasons are two things the rate card does not show: how many tokens Haiku 5.5 counts for the same text, and the fact that it thinks before answering unless you tell it not to.

  • Haiku 5.5 cost $0.75 per 1,000 chatbot replies at its default settings, 85% less than Haiku 4.5’s $4.98 on the same prompts. The per-token rate fell 90%; the newer tokenizer gives some of that back.
  • It cost 41% more than GPT-6 Luna ($0.53 per 1,000 with reasoning off), even though the two models have identical list prices.
  • Haiku 5.5 counted the identical request as 6,838 tokens. GPT-6 Luna counted 4,099. That is 1.67 times as many tokens at the same $0.10 rate.
  • Haiku 5.5 thinks by default. It spent 250 to 530 hidden thinking tokens on 4 of 6 replies. Turning thinking off cut the bill to $0.63 and the mean response time from 2.02 seconds to 0.88, the fastest of anything we tested.
  • Prompts over 100,000 tokens cost five times as much. A normal retrieval chatbot is nowhere near that line; a bot that pastes a whole site into every request can be.

Claude Haiku 5.5 pricing: the rate card

These are Anthropic’s published rates per million tokens, read from its pricing page on 8 October 2026. Haiku 5.5 is the first current Claude model with two price rows: one for prompts up to 100,000 tokens, and one for anything longer. GPT-6 Luna’s rates are OpenAI’s list prices as of 28 September 2026, which we covered in our GPT-6 Luna pricing test.

Per 1M tokensClaude Haiku 5.5 (prompt up to 100K)Claude Haiku 5.5 (prompt over 100K)Claude Haiku 4.5GPT-6 Luna
Input$0.10$0.50$1$0.10
Output (includes thinking)$0.50$2.50$5$0.50
5-minute cache write$0.125$0.625$1.25$0.125
1-hour cache write$0.20$1$2—
Cache read$0.01$0.05$0.10$0.01
Batch input / output$0.05 / $0.25$0.25 / $1.25$0.50 / $2.50—
Shortest prefix it will cache512 tokens512 tokens4,096 tokensAutomatic
Tokens for our chatbot request6,8384,5504,099

Read across the first column and the last one and the two models look interchangeable. Read down to the last row and they are not. The rate is per token, and the models do not agree on what a token is.

How we tested

The setup is the same one we used for our Claude Haiku 4.5 pricing test four days ago, so the Haiku 4.5 numbers here are a fresh re-run of that test, not copied from it. Every call went from the WordPress server that runs mxchat.ai, using the site’s own API keys.

  • The prompt: our chatbot’s real system prompt (3,478 characters), plus four retrieved passages from our own documentation and blog, plus the visitor’s question. That is the shape of a request from a retrieval chatbot such as MxChat, where the bot looks up relevant pages and passes them to the model with the question.
  • The questions: three real pre-sales questions. Does MxChat work with WooCommerce order lookups? Do I need Pinecone for the knowledge base? Can a visitor be handed to a human on Slack or Telegram?
  • The setups: Haiku 4.5 at its defaults; Haiku 5.5 at its default, at effort: low, at effort: high and with thinking switched off; Claude Sonnet 5.5 at effort: low as a reference; GPT-6 Luna through Chat Completions with reasoning_effort at none and at its default.
  • The cache layout: for Claude, a cache breakpoint on the system prompt only, which is how most WordPress chatbot plugins that cache at all do it. The retrieved pages change with every question, so they are never cached. OpenAI caches automatically, and on GPT-6 Luna it writes each new prompt to the cache at 1.25 times the input rate, which we included.

Each Claude setup answered each question twice, six replies per setup. GPT-6 Luna answered each question once as a first-time question, because a word-for-word repeat on OpenAI is read from its cache and would make Luna look cheaper than a real chatbot sees. Costs are computed from the usage fields each API returned, at list prices. In all, 66 answered requests plus token counts and parameter probes.

What 1,000 chatbot replies cost

Bar chart of cost per 1,000 chatbot replies: Claude Haiku 4.5 $4.98, Haiku 5.5 effort high $0.89, default $0.75, effort low $0.70, thinking off $0.63, GPT-6 Luna reasoning default $0.56 and reasoning none $0.53

Model and settingCost per 1,000 repliesInput tokensOutput tokens (of which thinking)Replies that thoughtMean time
Claude Haiku 4.5$4.984,55085 (0)0 of 61.53 s
Claude Haiku 5.5, default$0.756,838359 (248)4 of 62.02 s
Claude Haiku 5.5, effort: low$0.706,838255 (162)4 of 61.69 s
Claude Haiku 5.5, effort: high$0.896,838636 (525)6 of 63.20 s
Claude Haiku 5.5, thinking off$0.636,837106 (0)0 of 60.88 s
GPT-6 Luna, reasoning_effort: none$0.534,09943 (0)0 of 31.66 s
GPT-6 Luna, default reasoning$0.564,09985 (41)2 of 32.34 s
Claude Sonnet 5.5, effort: low (reference)$13.186,838181 (0)0 of 63.20 s

Against Haiku 4.5 the saving is plain. The same three questions cost 85% less, and that holds at every Haiku 5.5 setting we tried: even effort: high, the most expensive, came in at under a fifth of the Haiku 4.5 bill.

The 85% needs its baseline named, because it is not the number on the price page. Per token, Haiku 5.5 is 90% cheaper than Haiku 4.5. Per reply, it is 85% cheaper, because Haiku 5.5 counts more tokens for the same text and, at its default, writes more of them. Against Sonnet 5.5, Anthropic’s mid-tier model, Haiku 5.5 is about 94% cheaper per reply.

Against GPT-6 Luna the picture flips. Two models with the same input rate, the same output rate and the same cache-write rate produced bills $0.53 and $0.75 apart. With Haiku 5.5’s thinking switched off the gap narrows to 18% ($0.53 against $0.63). It does not close, because the remaining difference is in the input, and input is where a retrieval chatbot spends its money.

Why the same rate card gives a different bill: the tokenizer

Bar chart of tokens billed for the identical chatbot request: Claude Haiku 5.5 6,838, Claude Haiku 4.5 4,550, GPT-6 Luna 4,099, with cache minimums of 512 tokens for Haiku 5.5 and 4,096 for Haiku 4.5

A token is whatever the model’s tokenizer says it is, and each lab uses its own. Anthropic’s pricing page says its newer tokenizer, used by Claude 4.7 and later models, “produces approximately 30% more tokens for the same text,” and notes that the exact increase depends on the content. On our chatbot request it was more than that. Haiku 5.5 counted 6,838 tokens where Haiku 4.5 counted 4,550 (1.50 times) and GPT-6 Luna counted 4,099 (1.67 times).

Haiku 5.5 uses the same newer tokenizer as Sonnet 5.5 and Opus 5.5: Anthropic’s token-counting endpoint returned exactly the same numbers for Haiku 5.5 and Sonnet 5.5 on every request we sent. Haiku 4.5 is the last current Claude model on the older one. If you want the mechanics of why the counts differ and how to check your own prompts, our write-up of the Claude tokenizer change walks through it.

The practical point is a simple multiplication. A retrieval chatbot’s bill is mostly input: at Haiku 5.5’s default, input was 76% of the cost per reply, and with thinking off it was 92%. Multiply the input by 1.67 at the same rate and you get most of the gap to GPT-6 Luna. You cannot set it away; it is a property of the model. What you can control is the other part of the gap.

Haiku 5.5 thinks unless you tell it not to

Haiku 4.5 does not reason before answering unless you turn extended thinking on and give it a budget. Haiku 5.5 is the other way round. With no thinking settings in the request, it decided per question whether to think: it spent 529, 356, 352 and 253 thinking tokens on four replies, and none on the other two. At the default, the Pinecone question, the simplest of the three, never triggered thinking. The WooCommerce question always did.

Thinking tokens are billed as output, at $0.50 per million, and they are invisible to the visitor. At the default they made up about a sixth of the bill. They also cost time. Mean response time was 2.02 seconds at the default, with the slowest reply at 3.25 seconds, against 0.88 seconds with thinking off. In a chat window, under a second feels instant and three seconds is long enough to notice the typing indicator.

Here is what each setting did on our three questions:

Request settingWhat it doesCost per 1,000Mean timeCorrect answers (first run)
Nothing set (default)Model decides per question whether to think$0.752.02 s3 of 3
"output_config": {"effort": "low"}Still thinks, a bit less$0.701.69 s3 of 3
"output_config": {"effort": "high"}Thinks on every question$0.893.20 s3 of 3
"thinking": {"type": "disabled"}No thinking at all$0.630.88 s3 of 3

Unlike Haiku 4.5, which returns an HTTP 400 error if you send the effort parameter, Haiku 5.5 accepts it. But effort: low did not stop thinking, it only shortened it: four of six replies still thought. The setting that actually turns it off is thinking: {"type": "disabled"}.

For a support chatbot answering questions from retrieved documentation, thinking bought nothing we could measure. Every setting answered all three questions correctly from the sources on the run we checked by hand, and the thinking-off answers were as specific as the others: the WooCommerce add-on needs MxChat Pro and WooCommerce 6.0 or later, order lookups are limited to logged-in customers, embeddings can live in the WordPress database without Pinecone, and hand-off goes to a Slack channel or a Telegram topic. Three questions are not a benchmark, and a bot that has to reason across many documents, or follow a long multi-step policy, may get more from thinking. For short questions answered from a few retrieved passages, start with thinking off and turn it on only if you see answers that need it.

One more difference showed up in the replies. Our system prompt tells the bot to offer a discount code to visitors who are close to buying. Sonnet 5.5 used it in 4 of 6 replies, Haiku 5.5 in 3 of 24 across all its settings, and Haiku 4.5 and GPT-6 Luna in none. If your bot’s system prompt carries sales instructions, small models follow them less reliably than mid-tier ones, and that is worth testing before you switch.

Speed

Bar chart of mean seconds per reply: Claude Haiku 5.5 effort high 3.20 s, Haiku 5.5 default 2.02 s, GPT-6 Luna reasoning none 1.66 s, Claude Haiku 4.5 1.53 s, Claude Haiku 5.5 thinking off 0.88 s

With thinking off, Haiku 5.5 was the fastest model in the test: 0.88 seconds on average, and never slower than 1.26 seconds. That is faster than Haiku 4.5 (1.53 seconds) and GPT-6 Luna with reasoning off (1.66 seconds). At its default it was slower than both, because of the replies where it stopped to think.

These times are measured from a server in the United States to each provider’s API, and include the network. They are means over a handful of calls on one morning, not a load test. The ordering is what matters, and it was the same on both runs.

Prompt caching and follow-up questions

Haiku 5.5 fixes a problem that made prompt caching useless on Haiku 4.5. Anthropic will only cache a prefix above a minimum length, and for Haiku 4.5 that minimum is 4,096 tokens. Our system prompt is 836 tokens on Haiku 4.5, so the cache breakpoint our test placed on it was ignored on every call. Haiku 5.5’s minimum is 512 tokens, and the same system prompt (1,212 tokens on the newer tokenizer) was cached and read back at $0.01 per million on every call after the first.

On our prompt that saves little, about 11 cents per 1,000 replies, because the system prompt is a small slice of each request. It matters more for chatbots with long fixed instructions, and much more for follow-up questions, where the whole earlier conversation can be read from cache. We tested that with automatic caching (a single cache_control field at the top level of the request, which moves the cache breakpoint to the end of the conversation) and one follow-up per conversation: “Is that included in the free plugin, or do I need MxChat Pro for it?”

Per 1,000 replies, as billedFirst questionFollow-upThinking tokens on follow-up
Claude Haiku 4.5$6.15$0.850
Claude Haiku 5.5, default$0.88$0.26264
GPT-6 Luna, reasoning none—$0.070

The first question costs more than in the main test because the whole prompt is written to the cache at 1.25 times the input rate. Every follow-up then reads it back at a tenth of that. Haiku 5.5’s follow-ups were 70% cheaper than its first questions, and still about four times what GPT-6 Luna charged for a follow-up. Most of that difference is thinking: Haiku 5.5 thought on two of its three follow-ups, and at $0.50 per million those thinking tokens were about half of the follow-up bill. With thinking off, expect follow-ups closer to Luna’s. The GPT-6 Luna first-question figure is left out of this table because its prompt was already in OpenAI’s cache from the main test; the $0.53 above is the fair first-question number. For more on OpenAI’s side of caching, see our test of OpenAI’s cache-write charge.

The 100,000-token line

Haiku 5.5 is the only current Claude model whose price depends on prompt length. Anthropic’s pricing page puts any prompt of over 100,000 tokens on a higher price row, and that second row of the rate card is five times the first: $0.50 input, $2.50 output, $0.05 for a cache read. The page does not say whether cached tokens count toward the 100,000, so assume they do. Every other Claude 4.6 and later model is billed at one rate across its full million-token context.

For a retrieval chatbot this is not close. Our requests were about 6,800 tokens, roughly a fifteenth of the line. To see where the line actually falls, we kept adding passages from our own site to a single request and counted tokens on each model:

Knowledge base text in one requestClaude Haiku 5.5 tokensClaude Haiku 4.5 tokens
25,803 characters (12 passages)9,6306,422
85,664 characters (30 passages)29,22919,324
163,442 characters (54 passages)55,15636,358
244,603 characters (78 passages)81,80753,468

On our content Haiku 5.5 counts about three characters per token, so the 100,000-token line sits at roughly 300,000 characters, or about 50,000 words, in a single request. A chatbot that retrieves a few relevant passages per question never gets near it. The ones that can cross it are bots that paste a whole site, a product catalogue or a very long conversation history into every request. If yours does, check the token count before switching: above the line, Haiku 5.5’s input rate ($0.50) is still half Haiku 4.5’s, but its tokenizer counts about 1.5 times as many tokens, so the per-request saving shrinks to around a quarter.

Which cheap model should a WordPress chatbot use?

  • If you are on Claude Haiku 4.5, switch. Haiku 5.5 was 85% cheaper on identical requests at its defaults and answered every question we checked correctly. Set thinking: {"type": "disabled"} unless you have a reason not to, and check whether your plugin sends any thinking or effort settings of its own.
  • If cost per reply is the only thing that matters, GPT-6 Luna is still cheaper. $0.53 against $0.63 for Haiku 5.5 with thinking off, entirely because of the tokenizer. At 10,000 chatbot replies a month the difference is about a dollar, so on most WordPress sites it should not decide the question by itself.
  • If response time matters most, Haiku 5.5 with thinking off was the fastest model in our test at 0.88 seconds.
  • If your bot sends very large prompts, count tokens on Haiku 5.5 before you switch, and keep requests under 100,000 tokens.
  • If you need the bot to follow sales or escalation instructions closely, test that specifically. None of the small models we tried followed our discount instruction as reliably as Claude Sonnet 5.5, which costs roughly 20 times as much per reply.

MxChat lets you choose the model your chatbot uses and connect it with your own API key, so the bill is whatever your provider charges; there is no markup on tokens. The MxChat documentation covers model and API key setup, and MxChat Pro adds the WooCommerce, live-agent and knowledge-base add-ons the test questions were about.

FAQ

How much does Claude Haiku 5.5 cost?

For prompts up to 100,000 tokens: $0.10 per million input tokens, $0.50 per million output tokens (thinking included), $0.125 per million for a 5-minute cache write, $0.20 for a 1-hour cache write and $0.01 for a cache read. Prompts over 100,000 tokens pay five times those rates. The Batch API halves them. Rates as listed by Anthropic on 8 October 2026.

How much cheaper is Haiku 5.5 than Haiku 4.5?

90% per token, and 85% per chatbot reply in our test ($0.75 against $4.98 per 1,000 replies on identical requests). The difference between the two figures is the newer tokenizer, which counted the same request as 1.5 times as many tokens.

Is Claude Haiku 5.5 the same price as GPT-6 Luna?

The list prices are identical: $0.10 input and $0.50 output per million tokens on both. The bills are not. On the same chatbot prompts Haiku 5.5 cost $0.75 per 1,000 replies at its defaults and $0.63 with thinking off, against $0.53 for GPT-6 Luna with reasoning off, because Haiku 5.5 counted 1.67 times as many tokens.

Does Claude Haiku 5.5 use extended thinking by default?

In our test, yes. With no thinking settings in the request, it thought on 4 of 6 replies, 250 to 530 tokens each, billed as output. Setting effort to low shortened the thinking but did not stop it. Sending "thinking": {"type": "disabled"} turned it off, cut the cost by 17% and the mean response time from 2.02 to 0.88 seconds.

What is the 100K-token rule on Haiku 5.5?

A request with a prompt of over 100,000 tokens is billed at the higher rate row: $0.50 input and $2.50 output per million, five times the normal rates. On our content that line is about 300,000 characters of text in one request. A retrieval chatbot sending a few passages per question stays far below it.

Can Claude Haiku 5.5 cache a short system prompt?

Yes. Its minimum cacheable prefix is 512 tokens, against 4,096 on Haiku 4.5. Our 1,212-token system prompt was cached and read back on Haiku 5.5; the same prompt was too short to cache on Haiku 4.5.

Similar Posts