OpenAI Ultrafast Pricing: 6x the Cost, 0.3 s Faster
OpenAI now sells three speeds for GPT-6.1 Sol. Standard is the normal rate. Fast mode, the tier that used to be called Priority processing, costs twice as much. Ultrafast, which OpenAI added for GPT-6.1 Sol this week, costs six times as much: $12 per million input tokens and $60 per million output tokens, against $2 and $10 on Standard. OpenAI calls Ultrafast “the fastest service tier in the OpenAI API” and says to use it “when speed justifies the higher cost”. It does not publish latency figures. We wanted to know what OpenAI Ultrafast pricing buys a WordPress chatbot, so on 11 October we sent the same chatbot request through all three tiers from our production server and timed every reply as it streamed.
The short version: for short support answers, Fast mode gave most of the speed for a third of Ultrafast’s extra cost, and Ultrafast did not run at all through the API endpoint most WordPress chatbot plugins use.
- Per 1,000 new questions, GPT-6.1 Sol cost $10.92 on Standard, $21.66 on Fast and $66.18 on Ultrafast. The multipliers on the rate card held almost exactly per reply: 2.0x and 6.1x.
- The first word arrived after 1.43 seconds on Standard, 0.84 s on Fast and 0.99 s on Ultrafast (means of nine streamed calls each). On short replies Ultrafast did not show text any sooner than Fast.
- Ultrafast finished the whole reply first: 1.10 s against 1.39 s on Fast and 2.37 s on Standard. Its advantage is generation speed once text starts. On a 30-word answer that was worth 0.29 seconds over Fast, for an extra $44.52 per 1,000 replies.
- Ultrafast was the most consistent. Its slowest first token was 1.21 s. Fast’s was 1.56 s and Standard’s 2.60 s.
- Chat Completions rejected Ultrafast on 9 of 9 attempts with “Invalid service_tier argument”. It only worked through the Responses API. Fast mode worked on both.
- GPT-6 Sol on Fast mode, through Chat Completions, was the quickest setup we ran: 0.71 s to the first token, 1.05 s to the end, at $21.52 per 1,000 replies.
OpenAI Ultrafast pricing next to Standard and Fast mode
These are the short-context rates per million tokens on OpenAI’s pricing page as we read it on 11 October 2026. Short context means requests of up to 272,000 input tokens, which covers any site chatbot we know of.
| GPT-6.1 Sol, per 1M tokens | Input | Cached input | Cache write | Output | vs Standard |
|---|---|---|---|---|---|
| Standard | $2.00 | $0.10 | $2.50 | $10.00 | 1x |
| Fast mode (formerly Priority) | $4.00 | $0.20 | $5.00 | $20.00 | 2x |
| Ultrafast | $12.00 | $0.60 | $15.00 | $60.00 | 6x |
| Ultrafast, long context (over 272K input) | $24.00 | $1.20 | $30.00 | $90.00 | 6x the long-context Standard rate |
Every line scales together. Cached input, cache writes and output all carry the same multiplier as input, so prompt caching saves the same share on Ultrafast as it does on Standard. In cash terms it saves more, because each cached token was more expensive to begin with.
For comparison, GPT-6 Astra on Ultrafast lists at $60 input and $300 output per million tokens, also six times its Standard rate. GPT-6 Sol, the model in MxChat’s own OpenAI list, has Fast mode at $4 and $20 but no Ultrafast tier: our test request for it came back with the same “Invalid service_tier argument” error. OpenAI’s Fast mode guide notes that Priority processing was renamed Fast mode on 30 July 2026, so older guides and code that say priority describe the same tier. Both values worked in our calls, and the API reported the tier back as fast either way.
How we tested
All calls went from our WordPress server in the US directly to OpenAI on 11 October 2026, using the API key our live chatbot uses. Each request carried our live chatbot system prompt (3,478 characters) and one of three real pre-sales questions:
- Does MxChat work with WooCommerce, and can the bot look up an order status for a logged-in customer?
- Do I need Pinecone to use the knowledge base, or can I store embeddings in WordPress?
- Can a visitor be handed over to a human on Slack or Telegram when the bot cannot answer?
For each question we included four passages from our own documentation and blog posts, about 16,100 characters, as the knowledge base. These are the same prompt, passages and questions we used for our GPT-6.1 Sol pricing test on 1 October, which came to about 4,100 input tokens per request.
We set reasoning effort to low, the lowest value GPT-6.1 Sol accepts (it rejects none) and the value MxChat sends for OpenAI’s GPT-6 models. Every call streamed, so we could record two times: when the first visible word arrived and when the last one did. The first number is what a visitor watching the chat window notices. Each tier ran each question three times, nine calls per tier, through the Responses API. We then ran GPT-6.1 Sol and GPT-6 Sol on Standard and Fast through Chat Completions, nine calls each, plus nine attempts at Ultrafast that failed. That makes 63 calls that returned an answer.
Costs are computed from the usage block the API returned, at the list rates above. A “new question” is priced the way OpenAI billed the first sighting of each prompt: 3 tokens at the input rate and the rest as a cache write. We confirmed that usage shape on all three tiers with a separate long prompt. A “repeat” is the same prompt sent again and read from the cache.
What a reply costs on each tier
| GPT-6.1 Sol, Responses API | Standard | Fast mode | Ultrafast |
|---|---|---|---|
| Per 1,000 new questions | $10.92 | $21.66 | $66.18 |
| Per 1,000, no cache write | $8.87 | $17.56 | $53.90 |
| Per 1,000 exact repeats (read from cache) | $1.05 | $1.98 | $7.41 |
| Mean output tokens (of which reasoning) | 67 (15) | 58 (11) | 79 (32) |
| Mean words in the answer | 34 | 31 | 30 |
| Output as a share of the bill | 7.5% | 6.6% | 8.6% |
| 5,000 new questions a month | $54.60 | $108.30 | $330.90 |
A chatbot bill is mostly input. The retrieved passages and system prompt were about 4,100 tokens and the answer about 60, so output was under a tenth of the cost on every tier. That is why the cost per reply tracks the input multiplier so closely. It also means a faster tier cannot make a reply cheaper by finishing sooner. You pay for the tokens, and the tokens were the same.
Ultrafast replies billed more reasoning tokens on average, 32 against 15 on Standard and 11 on Fast. With nine calls per tier we would not read much into that, but it pushed Ultrafast’s per-reply cost slightly above an exact six times Standard.
How much faster each tier was
| Mean of 9 streamed calls | First token | Slowest first token | Rest of the reply | Complete reply | Per 1,000 new questions |
|---|---|---|---|---|---|
| GPT-6.1 Sol, Standard (Responses) | 1.43 s | 2.60 s | 0.94 s | 2.37 s | $10.92 |
| GPT-6.1 Sol, Fast (Responses) | 0.84 s | 1.56 s | 0.55 s | 1.39 s | $21.66 |
| GPT-6.1 Sol, Ultrafast (Responses) | 0.99 s | 1.21 s | 0.11 s | 1.10 s | $66.18 |
| GPT-6.1 Sol, Standard (Chat Completions) | 1.90 s | 3.74 s | 0.76 s | 2.66 s | $10.99 |
| GPT-6.1 Sol, Fast (Chat Completions) | 1.47 s | 1.85 s | 0.60 s | 2.07 s | $21.98 |
| GPT-6 Sol, Standard (Chat Completions) | 1.17 s | 2.57 s | 0.31 s | 1.48 s | $10.70 |
| GPT-6 Sol, Fast (Chat Completions) | 0.71 s | 1.67 s | 0.33 s | 1.05 s | $21.52 |
Ultrafast is very fast once it starts. It produced the rest of an answer in about a tenth of a second, where Standard took almost a second. What it did not do was start sooner than Fast mode. Most of the wait on a chatbot reply is before the first word, while the model reads 4,000 tokens of context and does its short burst of reasoning. Ultrafast spends that time at about the same pace as Fast.
That matters because chat widgets stream. A visitor sees the answer begin, and from then on the text is arriving faster than they can read it on any tier. So for a 30-word support answer, the number a visitor feels is roughly the first-token time, and on that number Fast mode matched Ultrafast at a third of the price. Ultrafast’s gain would be bigger on long replies. At Standard’s pace a 600-token answer would take several seconds to finish writing, and Ultrafast would cut most of that. Its tighter worst case is also a real benefit if a slow outlier is what you are trying to remove.
Network time is in all of these figures. They are full round trips from one server on one morning, not lab benchmarks, and a different region or time of day will move them. The differences between tiers were consistent across all three questions.
Ultrafast needs the Responses API
This was the finding we did not expect. Sending service_tier: "ultrafast" to the Chat Completions endpoint failed on every attempt with HTTP 400, “Invalid service_tier argument”, for GPT-6.1 Sol and GPT-6 Astra alike. The same request through the Responses API was accepted, and the response confirmed service_tier: "ultrafast". OpenAI’s Ultrafast guide only shows the Responses API and WebSockets, and it recommends WebSockets for agents that make many quick tool calls.
Many WordPress chatbot plugins, MxChat included, send chat through Chat Completions. On those, Ultrafast is not a setting you can switch on. Fast mode is. It accepted fast and priority on both endpoints. According to OpenAI’s Fast mode guide, it can also be set as a project default in your OpenAI settings, so requests that do not name a tier would get it. We did not test the project setting.
MxChat does not send a service tier today, so an MxChat site gets Standard unless the OpenAI project default says otherwise. Its OpenAI model list includes GPT-6 Sol, GPT-6 Astra and GPT-6 Luna. It does not yet list GPT-6.1 Sol.
Things to know before you pay for speed
The cache did not carry across tiers in our runs. The first Fast and Ultrafast calls on each question showed zero cached tokens, even though Standard had sent the identical prompt seconds earlier. The same happened between the Responses API and Chat Completions. If you mix tiers, expect each one to warm its own cache and pay its own cache writes. Our OpenAI prompt caching test covers what those writes cost.
Fast mode can quietly fall back to Standard. OpenAI’s guide says that if traffic ramps too quickly, some Fast requests are served at standard speed and billed at the standard rate, with service_tier: "default" in the response. It suggests growing traffic by no more than 50% every 15 minutes once you pass a million input tokens a minute. A site chatbot will rarely get near that, but it is worth logging the tier the API reports, not the one you asked for. Every one of our calls was served on the tier requested.
Ultrafast has its own rate limits. OpenAI lists 1,000,000 tokens per minute for GPT-6.1 Sol on its Build usage tier, 4,000,000 on Launch and 40,000,000 on Grow. At about 4,100 tokens per chatbot request, the Build limit is roughly 240 requests a minute.
One answer pattern to watch. On the WooCommerce question, Ultrafast replied “I don’t have enough information” on two of its three runs and Fast mode on one of three, where Standard answered all three. Every other answer on every tier was correct against the source passages. Three runs per question is too few to say whether the tier had anything to do with it. Reasoning models at the default temperature do vary from call to call. We are noting it rather than drawing a conclusion.
Which tier makes sense for a WordPress chatbot
| Your situation | Tier we would pick | Why |
|---|---|---|
| A support or pre-sales bot with short answers | Standard, or Fast if replies feel slow | Fast cut the first-token wait by 0.6 s for about $10.74 per 1,000 replies |
| A plugin that uses Chat Completions | Standard or Fast | Ultrafast is rejected on that endpoint |
| Long answers such as product guides or generated drafts | Consider Ultrafast | Its generation speed matters more as answers get longer |
| A voice or real-time agent that cannot tolerate outliers | Consider Ultrafast | It had the tightest worst case in our test |
| Speed matters most and cost matters too | A faster model on Fast mode | GPT-6 Sol on Fast beat GPT-6.1 Sol on Ultrafast for a third of the price |
| Background jobs such as summaries or tagging | Batch or Flex | Half the Standard price if nobody is waiting; see our Batch API cost test |
The last row in our latency table is the one we would point most site owners to. GPT-6 Sol on Fast mode replied in 1.05 seconds end to end, slightly faster than GPT-6.1 Sol on Ultrafast, and cost $21.52 per 1,000 new questions instead of $66.18. Choosing a quicker model gave as much speed as the most expensive tier. If you are deciding between GPT-6 Sol and its rivals, our GPT-6 Sol vs Claude Sonnet 5.5 comparison has the per-reply numbers, and the GPT-6 Astra pricing test covers the top of OpenAI’s range.
MxChat lets you choose the OpenAI model per bot, and the setup steps are in the MxChat documentation. The WooCommerce, live-agent and other add-ons mentioned in the test questions come with MxChat Pro. For request routing that does not need a full reply at all, our Decisions API test found answers in about a tenth of a second.
FAQ
How much does OpenAI Ultrafast cost?
For GPT-6.1 Sol, $12 per million input tokens, $0.60 cached input, $15 per million cache-write tokens and $60 per million output tokens on requests up to 272,000 input tokens, as listed on 11 October 2026. That is six times the Standard rate. GPT-6 Astra on Ultrafast is $60 input and $300 output.
Is Ultrafast faster than Fast mode?
It finished replies faster in our test, 1.10 s against 1.39 s, because it generates text much more quickly once it starts. It did not show the first word sooner: 0.99 s against 0.84 s on Fast. For short streamed chatbot answers the difference a visitor notices was small.
What is the difference between Fast mode and Priority processing?
They are the same tier. OpenAI renamed Priority processing to Fast mode on 30 July 2026. The API accepts service_tier set to fast or priority, and in our calls both were reported back as fast. It costs twice the Standard rate.
Can I use Ultrafast with Chat Completions?
Not in our test. Every Chat Completions request with service_tier: "ultrafast" was rejected with “Invalid service_tier argument”. The Responses API accepted it. OpenAI’s guide documents Ultrafast for the Responses API and WebSockets.
Which models support Ultrafast?
OpenAI’s guide lists GPT-6 Astra and GPT-6.1 Sol. A GPT-6 Sol request with the Ultrafast tier was rejected in our test.
Does a faster tier make each reply cheaper?
No. You pay per token at the tier’s rate, and the token counts were similar on every tier. A chatbot reply is mostly input, so the cost per reply followed the input multiplier: about 2x on Fast and 6x on Ultrafast.
Method: 63 answered API calls plus probes from the mxchat.ai server on 11 October 2026, GPT-6.1 Sol and GPT-6 Sol, reasoning effort low, max output 4,000 tokens, streamed, through the Responses API and Chat Completions. Prices are list rates read from OpenAI’s pricing page the same day. Times are full round trips including network from one US server.