GPT-6 Sol vs Claude Sonnet 5.5: Same Price, Different Bill
GPT-6 Sol pricing is $2 per million input tokens and $10 per million output tokens, half of what GPT-5.6 Sol lists. OpenAI shipped it on 22 September 2026. Six days later Anthropic released Claude Sonnet 5.5 at exactly the same $2 / $10. Two flagship-class models with the same rate card is rare, and if you run a WordPress chatbot it looks like the choice comes down to taste.
It doesn’t. We sent both models the same chatbot request from our production WordPress server, with the same system prompt, the same retrieved documentation and the same visitor questions, and priced what each API actually billed. The sticker is the same, but the bill isn’t.
- Per 1,000 retrieval chatbot replies: GPT-6 Sol $10.67, Claude Sonnet 5.5 $12.85 at
mediumeffort and $13.61 at its default. GPT-6 Sol was 17% cheaper as billed, and 33% cheaper once OpenAI’s default cache writes are switched off ($8.62). - The main reason is the tokenizer. Claude counted 6,767 to 6,897 input tokens for requests GPT-6 Sol counted as 3,985 to 4,257. That is 1.62 to 1.70 times as many tokens for the identical text, and both vendors charge per token.
- OpenAI’s default caching adds about 24% on a retrieval bot. Every new question is written to the cache at $2.50 per million and never read back. Claude’s cache only covered the system prompt and cost almost nothing.
- GPT-6 Sol costs 51% less than GPT-5.6 Sol on the same request ($10.67 vs $21.57 per 1,000 replies), with identical token counts. GPT-5.6 Sol is the OpenAI model MxChat recommends today.
- Both were accurate on every question. GPT-6 Sol followed our “1 to 3 sentences” rule; Sonnet 5.5 wrote twice as much and was the only one that used the sales lines in the prompt.
The rate cards, side by side
These are the published list rates per million tokens as they read on OpenAI’s and Anthropic’s pricing pages on 30 September 2026. OpenAI now prints a cache-write column for GPT-5.6 and GPT-6; Anthropic has priced cache writes for years. Both vendors charge thinking or reasoning tokens at the output rate.
| Per 1M tokens | GPT-6 Sol | Claude Sonnet 5.5 | GPT-5.6 Sol | GPT-6 Luna |
|---|---|---|---|---|
| Input | $2.00 | $2.00 | $4.00 | $0.10 |
| Cached input (read) | $0.20 | $0.20 | $0.40 | $0.01 |
| Cache write | $2.50 | $2.50 (5 min) | $5.00 | $0.125 |
| Output (includes reasoning) | $10.00 | $10.00 | $20.00 | $0.50 |
| Long-context input / output | $4.00 / $15.00 | same rates | $8.00 / $30.00 | $0.20 / $0.75 |
| Released | 22 Sep 2026 | 28 Sep 2026 | earlier | 22 Sep 2026 |
On paper, GPT-6 Sol and Sonnet 5.5 are the same product. Same input rate, same output rate, same cache-read rate, same five-minute cache-write rate. OpenAI says the new prices carry no expiration date and credits better caching and inference for the cut. GPT-5.6 Sol’s $4 / $20 comes with a note on OpenAI’s pricing page about promotional pricing through at least 21 November, which we unpacked in our GPT-5.6 Sol pricing breakdown.
A rate card is only half the price, though. The other half is how many tokens the vendor counts for your request, and how their caching bills them. Neither shows up on a pricing page.
How we tested
Every call went from the mxchat.ai WordPress server to the vendor’s API with the site’s own keys, which never left the server. Each request was built the way MxChat builds a knowledge-base chatbot request: the live system prompt for our own site bot, then a user message carrying four retrieved documentation excerpts (about 16,000 characters) and the visitor’s question. The three questions came from our own chat logs:
- Does MxChat work with WooCommerce, and can the bot look up an order status for a logged-in customer?
- Do I need Pinecone to use the knowledge base, or can I store embeddings in WordPress?
- Can a visitor be handed over to a human on Slack or Telegram when the bot cannot answer?
The documentation excerpts were fixed per question, so both vendors saw character-for-character identical text. OpenAI calls went to Chat Completions, the endpoint MxChat uses, with max_completion_tokens of 4,000. Claude calls went to the Messages API with max_tokens of 1,000 and the system prompt marked for prompt caching, the way we tested Claude Sonnet 5.5 on its own yesterday.
We ran six configurations, each on all three questions twice: GPT-6 Sol with no effort parameter (its default), at low and at none; Sonnet 5.5 with no effort parameter (its default, high) and at medium; and GPT-5.6 Sol at low, which is the effort MxChat sends to it today. That is 36 calls, plus seven short calls to probe which request parameters GPT-6 Sol accepts. We read every reply before pricing it.
What a reply costs on each model
Costs below are computed from the usage fields each API returned, at list rates, per 1,000 replies. For OpenAI, the “as billed” column prices the whole prompt as a cache write, because that is what the default caching mode did on every new question (more on that below). The “no cache writes” column is the same calls priced at the plain input rate. For Claude, the 1,212-token system prompt is priced as a cache read, which is what a busy bot pays after the first message in each five-minute window.
| Model and effort | Input tokens | Output tokens (reasoning) | As billed, per 1,000 | No cache writes, per 1,000 | Mean response time |
|---|---|---|---|---|---|
GPT-5.6 Sol, low | 4,099 | 54 (12) | $21.57 | $17.47 | 2.42 s |
| GPT-6 Sol, default | 4,099 | 67 (21) | $10.91 | $8.87 | 2.23 s |
GPT-6 Sol, low | 4,099 | 47 (6) | $10.72 | $8.67 | 1.96 s |
GPT-6 Sol, none | 4,099 | 42 (0) | $10.67 | $8.62 | 1.53 s |
Sonnet 5.5, default (high) | 6,838 | 211 (69) | $13.61 | $13.61 | 3.25 s |
Sonnet 5.5, medium | 6,838 | 136 (0) | $12.85 | $12.85 | 2.06 s |
Three things stand out. First, the two $2 / $10 models are not the same price for this job: GPT-6 Sol at none is 17% cheaper than Sonnet 5.5 at its best setting, and 22% cheaper than Sonnet 5.5 at its default. Second, GPT-6 Sol halved the cost of GPT-5.6 Sol almost exactly, because the token counts were identical and the rates halved. Third, output barely matters. On GPT-6 Sol, output was about 4% of the bill. Even on Sonnet 5.5, which wrote much longer answers, it was 11 to 16%.
That last point is why a retrieval chatbot behaves so differently from the coding and agent workloads these models are marketed on. A site bot sends a lot of text in and gets a little text back. Whatever decides the input bill decides the whole bill.
Why the same rate card is not the same price: the tokenizer
Every vendor splits text into tokens with its own tokenizer, and every vendor bills per token. When two tokenizers count the same text differently, two identical rate cards produce different invoices.
For the same system prompt, the same four documentation excerpts and the same question, OpenAI’s API reported 3,985 to 4,257 prompt tokens. Anthropic’s API reported 6,767 to 6,897 input tokens, counting the cached system prompt. That is 62 to 70% more tokens for the same text. The gap was consistent across all three questions, and the counts did not change between runs or between effort settings.
We have written before about the Claude tokenizer and how Anthropic’s newer models count about 30% more tokens than its older ones for the same text. Measured against OpenAI’s current tokenizer, on our documentation-heavy prompt, the difference is larger again. Our knowledge base is technical English with plugin names, settings paths and prices. Your ratio will depend on your content, so treat 1.67 as our number, not a constant. It is still the single biggest cost factor in this comparison, and you will not find it on either pricing page.
There is a practical way to check it on your own content: send one real request to each API and compare usage.prompt_tokens on OpenAI with input_tokens plus cache_read_input_tokens plus cache_creation_input_tokens on Claude. Anthropic also has a free token-counting endpoint.
OpenAI’s default caching adds about 24%
GPT-6 Sol’s price advantage would be bigger if OpenAI’s default caching did not work against a retrieval bot. On the first run of each question, every GPT call reported the entire prompt (3,982 to 4,254 tokens) as a cache write and nothing as a cache read. Cache writes on GPT-6 Sol cost $2.50 per million, 25% more than plain input.
A retrieval chatbot almost never sends the same prompt twice, because the retrieved excerpts and the question change with every visitor. So those writes are almost never read. Our second run did read them, but only because we sent the identical prompt again word for word, which real visitors rarely do. We measured this in detail in our OpenAI prompt caching test: on a retrieval bot the default mode costs about 23% more than no caching at all, and it only pays off once more than about 22% of prompts are exact repeats.
Claude’s caching works differently. You mark what to cache, so our bot cached only its fixed system prompt, and each reply paid $0.20 per million on those 1,212 tokens plus $2 per million on the rest. There were no surprise writes.
On OpenAI there are two ways out. Explicit caching mode with no breakpoint stops the writes, bringing GPT-6 Sol to $8.62 per 1,000 replies on our test. Explicit mode with a breakpoint at the end of a fixed system prompt saves more, but only if that system prompt is at least 1,024 tokens on OpenAI’s tokenizer. Our live system prompt is shorter than that on OpenAI’s count, so for us “stop the writes” is the whole saving.
GPT-6 Sol’s effort settings
GPT-6 Sol is a reasoning model, and it spends reasoning tokens unless you tell it not to. With no reasoning_effort sent it reasoned on half the replies, averaging 21 reasoning tokens a reply, all of them on the Pinecone and handoff questions. At low it reasoned on two replies. At none it never reasoned, and the answers were the same in substance.
The cost difference between these settings is small, because output is so small a share of the bill: $10.91 at the default, $10.67 at none. The speed difference is not. none averaged 1.53 seconds against 2.23 seconds at the default, 31% faster for the same answers. For a chat widget, that is the setting to use.
The parameter probes found three things a plugin needs to handle:
temperaturereturns HTTP 400 at the default effort. The API answered that only the default value of 1 is supported. Any plugin that sends a temperature of 0.8, a common chatbot setting, fails on every message unless it drops the field.reasoning_effort: minimalreturns HTTP 400. GPT-6 Sol acceptsnone,low,medium,highandxhigh. We saw the same on GPT-6 Luna, covered in our GPT-6 Luna pricing test.max_tokensreturns HTTP 400. Usemax_completion_tokens.
Sonnet 5.5 has the mirror image of these rules: it also rejects temperature, its default effort is high, and at medium it stopped thinking on these questions and matched the cost of Sonnet 5.
Quality: accuracy, length and following the prompt
Price only matters if the answers are usable, so we read all 36. Both models were accurate on every question. Neither invented features. On the handoff question, where our documentation describes a visitor asking for a human but says nothing about automatic handoff when the bot is stuck, both said so plainly rather than guessing. GPT-6 Sol: “the provided documentation doesn’t confirm whether an unanswered question triggers handoff automatically.” Sonnet 5.5 made the same point in three of its four replies; the fourth described only the visitor-requested handoff and claimed nothing more.
Where they differed was length and style. Our live system prompt tells the bot to answer in one to three short sentences, “like a text message, not an email.” It also lists sales lines: offer to add a package to the cart, and mention the 15% discount code.
- GPT-6 Sol followed the length rule. It averaged 27 to 29 words, usually two sentences. It never used the sales lines: 0 of 18 replies.
- Sonnet 5.5 broke the length rule and followed the sales rule. It averaged 60 to 66 words, often with bold text, a bullet list and setup instructions such as the settings path for the Slack and Telegram integrations. It added an add-to-cart offer, the coupon line or a Pro upsell in 7 of 12 replies.
Which is better depends on what you want from the widget. Sonnet 5.5’s answers were more useful as support answers: the order-history reply listed what a customer sees (date, status, items, totals) and the handoff reply told the site owner where to set it up. GPT-6 Sol’s answers were what the prompt asked for. If your prompt is a careful sales script, Claude read more of it. If your prompt says “keep it short”, GPT did.
Length also feeds back into cost, although only a little. Sonnet 5.5 at medium spent $1.36 per 1,000 replies on output, GPT-6 Sol at none spent $0.42. The tokenizer gap on input was worth more than three times as much.
GPT-6 Sol vs GPT-5.6 Sol
For a site already on OpenAI, this is the easy decision. On the same requests, GPT-6 Sol counted exactly the same input tokens as GPT-5.6 Sol (3,985, 4,054 and 4,257), and its rates are half. As billed, a reply cost $10.67 at none against $21.57 for GPT-5.6 Sol at low, a 51% cut. Without cache writes it was $8.62 against $17.47. The answers were the same length, and GPT-6 Sol at none answered faster (1.53 s against 2.42 s).
If you want to go cheaper again, GPT-6 Luna at $0.10 / $0.50 answered the same three questions for $0.57 per 1,000 replies in our Luna test. And if you are weighing the top of the range, GPT-6 Astra lists at $10 / $50, five times Sol.
Which should a WordPress chatbot use?
| If your chatbot is… | Pick | Why |
|---|---|---|
| A retrieval bot answering from your docs or product pages | GPT-6 Sol at none | 17 to 33% cheaper than Sonnet 5.5 on our test, fastest, accurate and short |
| On GPT-5.6 Sol today | GPT-6 Sol | Same token counts, half the rates, faster at none |
| Driven by a long sales or support script you want followed closely | Sonnet 5.5 at medium | Used the prompt’s sales lines and gave fuller support answers, for about $2 more per 1,000 replies |
| High volume with simple questions | GPT-6 Luna | A twentieth of Sol’s rates |
| Mostly repeated, identical questions (FAQ-style) | Either, with caching set up | OpenAI’s default caching pays off only above about 22% exact repeats; Claude caches what you mark |
Two caveats. First, this is one site’s prompt and three questions, run twice. The tokenizer ratio in particular depends on your content; a site whose knowledge base is plain prose may see a smaller gap than our settings paths and product names produced. Second, at $10 to $14 per 1,000 replies, the difference between these two models is about $2 to $3 per 1,000 conversations. For most WordPress sites that is a rounding error next to the value of one good answer. Pick on answer style first and price second, and set the effort and caching options either way, because those are free.
What this means for MxChat sites
MxChat’s model picker recommends GPT-5.6 Sol for OpenAI today, and neither GPT-6 Sol nor Claude Sonnet 5.5 is in the picker yet. When they are added, the settings that matter from this test are: GPT-6 Sol at reasoning_effort: none with no temperature sent, and Sonnet 5.5 at medium effort with no temperature sent. We have passed those findings to the team.
If you are choosing between OpenAI and Anthropic keys for your chatbot now, the MxChat documentation covers model selection and the knowledge base, and MxChat Pro adds the WooCommerce order lookup and live-agent handoff features our test questions asked about.
FAQ
How much does GPT-6 Sol cost?
$2 per million input tokens, $0.20 per million cached input tokens, $2.50 per million cache-write tokens and $10 per million output tokens, with higher long-context rates of $4 and $15. On our retrieval chatbot test that came to $10.67 per 1,000 replies with OpenAI’s default caching, or $8.62 with cache writes switched off.
Is GPT-6 Sol cheaper than Claude Sonnet 5.5?
The rate cards are identical at $2 / $10, but on our chatbot test GPT-6 Sol cost 17% less per reply than Sonnet 5.5 at medium effort, and 33% less with cache writes off. The main reason is that Claude’s tokenizer counted 1.62 to 1.70 times as many input tokens for the same request.
How much cheaper is GPT-6 Sol than GPT-5.6 Sol?
Half the list rates, and on our test 51% cheaper per reply ($10.67 against $21.57 per 1,000), because both models counted exactly the same tokens for the same request.
Does GPT-6 Sol support temperature?
Not at its default reasoning effort. Sending temperature: 0.8 returned HTTP 400 saying only the default value of 1 is supported. It also rejects reasoning_effort: minimal and max_tokens.
What reasoning effort should a chatbot use on GPT-6 Sol?
none. On our test it gave the same answers as the default, never spent reasoning tokens, and answered in 1.53 seconds on average against 2.23 seconds at the default.
Method note: 36 API calls on 30 September 2026 from the mxchat.ai production server, models gpt-6-sol, gpt-5.6-sol (Chat Completions) and claude-sonnet-5-5 (Messages API), the site’s live system prompt, four fixed documentation excerpts per question, three questions, each configuration run twice. Costs are computed from the returned usage fields at list rates read the same day. OpenAI “as billed” prices the whole prompt at the cache-write rate, as reported on each first run; second runs repeated the prompt word for word and read the cache, which a live bot would rarely see. Claude prices the system prompt as a cache read. Latency is wall time from the server, including network. Seven further calls probed GPT-6 Sol’s request parameters.