Cover card: GPT-6.1 Sol pricing, cheaper cache, slower chatbot. $10.89 vs $10.66 per 1,000 new replies and 2.6x slower than GPT-6 Sol, from 42 real API calls

GPT-6.1 Sol Pricing: Cheaper Cache, Slower Chatbot

GPT-6.1 Sol pricing is $2 per million input tokens and $10 per million output tokens, the same as GPT-6 Sol, which OpenAI released only a week earlier. The one number that moved is cached input: $0.10 per million instead of $0.20. OpenAI shipped GPT-6.1 Sol on 29 September 2026 and describes it as near-Astra quality at a fifth of Astra’s price.

That pitch is aimed at coding agents. We wanted to know what it means for a WordPress chatbot, so we sent GPT-6.1 Sol and GPT-6 Sol the same chatbot requests from our production server and priced what the API billed. In short, the rate card got cheaper but our chatbot’s bill did not.

  • Per 1,000 chatbot replies to new questions: GPT-6.1 Sol $10.89 at its lowest effort, GPT-6 Sol $10.66. The newer model cost 2% more at low and 9% more at its default effort ($11.63).
  • GPT-6.1 Sol has no none effort. The API rejects it. Its floor is low, so it reasons on some replies, and reasoning is billed as output.
  • It was 2.6 times slower. Mean response time was 3.59 seconds at low and 6.18 seconds at default, against 1.39 seconds for GPT-6 Sol at none.
  • The cheaper cache rate only showed up on follow-up turns. When a visitor asked a second question in the same conversation, GPT-6.1 Sol cost $1.00 per 1,000 replies against $1.32, which is 24% less.
  • It refused 3 of 15 questions the sources could answer. GPT-6 Sol answered all 21.

GPT-6.1 Sol pricing, next to the models around it

These are the list rates per million tokens as they read on OpenAI’s pricing page on 1 October 2026.

Per 1M tokensGPT-6.1 SolGPT-6 SolGPT-6 AstraGPT-5.6 Sol
Input$2.00$2.00$10.00$4.00
Cached input (read)$0.10$0.20$1.00$0.40
Cache write$2.50$2.50$12.50$5.00
Output (includes reasoning)$10.00$10.00$50.00$20.00
Cached read as a share of input5%10%10%10%

GPT-6.1 Sol also has the usual alternative tiers. Batch and Flex are half price ($1.00 input, $0.05 cached, $5.00 output). Fast mode, the tier OpenAI used to call priority processing, is double ($4.00 and $20.00). OpenAI’s model page lists a 1,050,000-token context window, a 128,000-token output limit, and higher rates for requests above 272,000 tokens. A site chatbot will not get near that threshold.

So input, output and cache-write rates match GPT-6 Sol to the cent, and “a fifth of Astra’s price” is plain arithmetic: $10 and $50 became $2 and $10. The only new number is the cached read, which dropped from 10% of the input rate to 5%. Whether that matters depends on how much of your traffic is ever read from the cache. For a retrieval chatbot that is not much, as we found in our OpenAI prompt caching test.

OpenAI’s model page describes GPT-6.1 Sol as built for “complex coding, computer use, and professional work”. A support chatbot is none of those, which is the reason to measure rather than assume.

How we tested

Every call went from the mxchat.ai WordPress server to OpenAI’s Chat Completions endpoint with the site’s own API key, which never left the server. Each request was built like a knowledge-base chatbot request: the live system prompt for our own site bot, then a user message carrying four retrieved documentation excerpts (about 16,000 characters) and the visitor’s question. The three questions came from our own chat logs:

  1. Does MxChat work with WooCommerce, and can the bot look up an order status for a logged-in customer?
  2. Do I need Pinecone to use the knowledge base, or can I store embeddings in WordPress?
  3. Can a visitor be handed over to a human on Slack or Telegram when the bot cannot answer?

These are the same requests, character for character, that we used the day before in our GPT-6 Sol vs Claude Sonnet 5.5 comparison, so the GPT-6 Sol figures here are a fresh re-run of that test. They landed within a few cents per 1,000 replies of that test’s numbers.

We ran five configurations, each on all three questions twice: GPT-6.1 Sol with no effort parameter and at low, and GPT-6 Sol with no effort parameter, at low and at none. Then we ran a follow-up test: for each question we sent the first turn, appended the model’s answer, and asked a second question in the same conversation. Add 13 short calls to probe which request parameters the models accept, and the total is 42 chat calls and 13 probes. We read every reply before pricing it.

What GPT-6.1 Sol rejects

The probes came first, because a model that returns HTTP 400 costs nothing and answers nothing. GPT-6.1 Sol is stricter than GPT-6 Sol in one way that matters for chatbots. The GPT-6 Sol column comes from the same probes run a day earlier, plus one more today.

Request parameterGPT-6.1 SolGPT-6 Sol
reasoning_effort: none400, not supportedAccepted
reasoning_effort: minimal400400
reasoning_effort: low, medium, high, xhighAcceptedAccepted
reasoning_effort: max on Chat Completions400400
temperature: 0.8400, only the default of 1400 at default effort
top_p: 0.9400Not tested
max_tokens400, use max_completion_tokens400, use max_completion_tokens

The first row is the important one. On GPT-6 Sol, none was the cheapest and fastest setting in our tests, and it is what we would send from a chatbot. GPT-6.1 Sol does not have it. OpenAI’s model page lists medium as the default, and the lowest value the API accepted was low.

If your plugin or script already sends reasoning_effort: none for GPT-6 Sol and you change only the model name to gpt-6.1-sol, every request fails with a 400 until you change the effort too. The same goes for a custom temperature. Check this before switching a live bot.

What a reply costs on each model

Costs are computed from the usage fields the API returned, at list rates, per 1,000 replies. “New question” prices the request as the API billed it the first time it saw that prompt: almost the whole prompt as a cache write, which is what OpenAI’s default caching does. “No cache writes” is the same call at the plain input rate. “Exact repeat” is the same prompt sent again word for word, read from the cache.

Model and effortOutput tokens (reasoning)New question, per 1,000No cache writes, per 1,000Exact repeat, per 1,000Mean response time
GPT-6.1 Sol, default (medium)139 (87)$11.63$9.58$1.806.18 s
GPT-6.1 Sol, low64 (15)$10.89$8.84$1.063.59 s
GPT-6 Sol, default60 (20)$10.84$8.80$1.421.76 s
GPT-6 Sol, low47 (6)$10.71$8.66$1.291.74 s
GPT-6 Sol, none41 (0)$10.66$8.61$1.241.39 s

Bar chart of cost per 1,000 chatbot replies to new questions: GPT-6.1 Sol default $11.63, GPT-6.1 Sol low $10.89, GPT-6 Sol default $10.84, low $10.71 and none $10.66, with the cache-write premium shown separately

Input tokens were identical on both models: 3,985, 4,054 and 4,257 for the three requests, a mean of 4,099. There is no tokenizer change between GPT-6 Sol and GPT-6.1 Sol. At the same input rate, that makes the input bill identical, and input is 88 to 96% of what a retrieval chatbot pays.

What differs is the output side. GPT-6.1 Sol reasoned on every one of its six default-effort calls, spending 16 to 133 hidden reasoning tokens each time, and on three of six calls at low. GPT-6 Sol at none never reasons, and at its default it reasoned on two of six. Reasoning tokens are billed as output at $10 per million. At GPT-6.1 Sol’s default, a mean of 87 reasoning tokens on top of about 50 visible ones made the output charge more than three times GPT-6 Sol’s.

That is still a small part of the total. GPT-6.1 Sol at low came out 2% above GPT-6 Sol at none, 23 cents per 1,000 replies. At 10,000 replies a month the two models cost $108.90 and $106.60. Nobody will choose between them on that difference.

The speed gap is bigger than the price gap

Bar chart of mean response time for the same chatbot request: GPT-6.1 Sol default 6.18 seconds, GPT-6.1 Sol low 3.59 seconds, GPT-6 Sol default 1.76 seconds, low 1.74 seconds and none 1.39 seconds

GPT-6.1 Sol took a mean of 3.59 seconds at low and 6.18 seconds at its default, measured from our server as the full round trip without streaming. GPT-6 Sol at none took 1.39 seconds. Its slowest single call, 1.70 seconds, was faster than GPT-6.1 Sol’s fastest first-turn call, 2.30 seconds.

For a coding agent that runs for minutes, a few seconds per call is nothing. For a visitor who typed a question into a chat bubble, the difference between one and a half seconds and six is the difference between a bot that feels instant and one that feels stuck. Streaming softens it, since the visitor sees words as they arrive, but reasoning happens before the first visible word, so the wait moves to the start of the reply.

Both figures are means of six calls on one morning. Latency moves with OpenAI’s load and a new model is often slower in its first week, so treat the ratio as a snapshot. The direction matches the token counts, though: the model that thinks more takes longer.

Where the cheaper cache rate shows up

The cached-input cut is real, and we saw it on two kinds of request.

The first is an exact repeat. On the second run of each question the API read the whole prompt from the cache. At $0.10 per million, 4,096 cached tokens cost GPT-6.1 Sol 41 cents per 1,000 replies, against 82 cents on GPT-6 Sol. After output, an exact repeat cost $1.06 per 1,000 on GPT-6.1 Sol at low and $1.24 on GPT-6 Sol at none. Real visitors rarely send a word-for-word repeat of a prompt that includes retrieved documentation, so this case is mostly theoretical for a site bot.

The second is a follow-up turn, and this one happens all the time. A visitor asks a question, reads the answer, and asks “is that in the free version?” If the second request repeats the first one exactly and adds the new messages at the end, the API reads the earlier part from the cache.

Follow-up turn, mean of 3GPT-6.1 Sol, lowGPT-6 Sol, none
Input tokens4,1764,165
Read from cache4,0964,096
Written to cache7766
Output tokens3933
Cost per 1,000 follow-ups$1.00$1.32
Same call with nothing cached$8.74$8.66
Mean response time2.53 s1.29 s

Bar chart of cost per 1,000 replies by request type: new question GPT-6.1 Sol $10.89 vs GPT-6 Sol $10.66, exact repeat $1.06 vs $1.24, follow-up turn $1.00 vs $1.32

On a follow-up, GPT-6.1 Sol was 24% cheaper. The first turn’s 4,000 tokens came back at the cached rate on both models, and GPT-6.1 Sol’s cached rate is half. This is where the new price helps a chatbot: conversations that keep going.

Putting the two together, GPT-6.1 Sol at low costs 23 cents more per 1,000 new questions and 32 cents less per 1,000 follow-ups. It becomes the cheaper model once a bot handles more than about 0.7 follow-up turns per new question. With two follow-ups per question it is 3% cheaper overall, and with three, 5%.

The catch: your follow-ups have to be cacheable

That saving only arrives if the front of the request is identical from one turn to the next. Many retrieval chatbots do not work that way. They search the knowledge base again for every message and put the fresh results near the top of the request, next to the system prompt. When the retrieved text changes, the request stops matching the cache at that point and everything after it is billed as new.

MxChat is in that group today. Version 3.2.22 appends the retrieved content to the system message, ahead of the conversation, so the front of the request changes on every question. Our fixed system prompt on its own is shorter than the 1,024 tokens OpenAI needs before it will cache anything. For our own bot, the cached-read price is a number that almost never applies, on either model. We have logged the request layout as a change for the plugin: retrieved content after the conversation would let follow-up turns read the cache.

To check your own bot, look at usage.prompt_tokens_details in a response to a second message in a conversation. If cached_tokens is zero on a follow-up, a lower cached rate saves you nothing.

Did it answer the questions?

We do not price replies we have not read. Our system prompt tells the bot to answer only from the supplied sources, in one to three short sentences, and to say it lacks the information when the sources do not cover the question.

First-turn repliesAnswered from the sourcesRefusedMean length
GPT-6.1 Sol, default4 of 62 (order status, both runs)30 words
GPT-6.1 Sol, low8 of 91 (human handoff)33 words
GPT-6 Sol, all three efforts21 of 21025 to 27 words

GPT-6.1 Sol said it did not have enough information on three first-turn calls where the retrieved excerpts did contain the answer. Both of its default-effort replies to the WooCommerce order-status question were refusals. The excerpt lists order-status lookups for logged-in customers, and GPT-6 Sol answered that question correctly all seven times it was asked. When GPT-6.1 Sol did answer, the facts matched the sources: Pinecone is optional, the WooCommerce add-on needs a Pro license, Slack and Telegram handoff are in the free plugin.

On the handoff question the sources confirm that a visitor can ask for a human. They do not say the bot escalates by itself. GPT-6 Sol never claimed automatic escalation, and said so explicitly in six of its seven handoff replies. One of GPT-6.1 Sol’s default-effort replies blurred the two, and its other answers kept them apart. All six follow-up answers, three per model, were correct.

Fifteen calls is a small sample, and a refusal is the safe failure for a support bot, far better than an invented answer. But a refusal on a question the knowledge base covers is still a visitor who did not get help. A model tuned for long coding tasks reading a “do not guess” instruction more strictly is a plausible explanation, and we would want more than three questions before calling it a pattern.

What about Gemini 4 Argon?

Google announced Gemini 4 Argon this week, and launch coverage reports an introductory $2 input and $10 output per million tokens, rising to $4 and $20 later, with cached input at a 95% discount. That would make four models listed at $2 / $10 in ten days: GPT-6 Sol, Claude Sonnet 5.5 (measured in our Claude Sonnet 5.5 pricing test), GPT-6.1 Sol and Argon.

We could not measure it. On 1 October the Gemini API listed 61 models for our key and none of them was Argon; reports say access starts with a group of trusted testers. No end date for the introductory rate has been published either. We will run the same requests through it the day it appears in the API. Until then, our most recent Gemini numbers are in the Gemini 3.8 Flash pricing test.

Which OpenAI model should a WordPress chatbot use?

On this evidence, GPT-6 Sol at reasoning_effort: none is still the better OpenAI model for a retrieval chatbot. It costs slightly less per new question, replies in under two seconds, and answered everything the sources covered. GPT-6.1 Sol costs the same or a little more, takes two and a half to four and a half times as long, and was more likely to refuse.

GPT-6.1 Sol makes sense when the work matches what it was built for: long agent runs, code, multi-step tool use, and anything that resends a large unchanged prefix many times. There the halved cached rate applies to most of the tokens and the extra reasoning does useful work.

If cost matters more than model quality, neither Sol is the cheapest option. GPT-6 Luna answered the same kind of request for about 57 cents per 1,000 replies in our GPT-6 Luna test. At the other end, GPT-6 Astra costs five times what either Sol does. And GPT-5.6 Sol, at $4 / $20, now lists at twice the price of both newer Sol models; the background is in our GPT-5.6 Sol pricing breakdown.

For MxChat users: the model picker in MxChat 3.2.22 lists the GPT-5.6 family and does not yet include GPT-6 Sol or GPT-6.1 Sol. We have passed today’s parameter findings to the plugin team so that when they are added, each model is sent an effort value it accepts. Model setup is covered in the MxChat documentation, and the add-ons mentioned in the test questions come with MxChat Pro.

FAQ

How much does GPT-6.1 Sol cost?

$2.00 per million input tokens, $0.10 per million cached input tokens, $2.50 per million cache-write tokens and $10.00 per million output tokens, as listed on 1 October 2026. Batch and Flex are half those rates and Fast mode is double. Reasoning tokens are billed as output.

Is GPT-6.1 Sol cheaper than GPT-6 Sol?

Only on cached input, which is $0.10 instead of $0.20 per million. Input, output and cache-write rates are identical. On our chatbot test GPT-6.1 Sol cost 2 to 9% more per new question because it spends reasoning tokens that GPT-6 Sol at none does not, and 24% less on follow-up turns that were read from the cache.

Does GPT-6.1 Sol support reasoning_effort none?

No. On Chat Completions the API returned HTTP 400 for none and minimal and accepted low, medium, high and xhigh. It also rejected a custom temperature, top_p and max_tokens.

Is GPT-6 Sol being retired now that GPT-6.1 Sol is out?

We found no retirement date. On 1 October 2026 the API’s model record for gpt-6-sol showed no shutdown date, and both models were available side by side.

Is GPT-6.1 Sol a good model for a customer support chatbot?

It works, but in our test GPT-6 Sol was the better fit: about the same cost, 2.6 times faster, and no refusals on questions the knowledge base covered. GPT-6.1 Sol is tuned for coding and long agent tasks, where its cheaper cache and extra reasoning pay off.

Method note: 42 Chat Completions calls and 13 parameter probes from the mxchat.ai server on 1 October 2026, max_completion_tokens 4,000, no streaming, default service tier. Prices are list rates read from OpenAI’s pricing page the same day. “New question” prices 3 tokens at the input rate and the rest of the prompt as a cache write, matching the usage the API reported on every first call. Response times are full round trips from one server on one morning.