Cover image: OpenAI Ultrafast costs six times the price and was 0.3 seconds faster than Fast mode on chatbot replies; GPT-6.1 Sol cost $10.92 per 1,000 replies on Standard, $21.66 on Fast mode and $66.18 on Ultrafast

OpenAI Ultrafast Pricing: 6x the Cost, 0.3 s Faster

OpenAI now sells three speeds for GPT-6.1 Sol. Standard is the normal rate. Fast mode, the tier that used to be called Priority processing, costs twice as much. Ultrafast, which OpenAI added for GPT-6.1 Sol this week, costs six times as much: $12 per million input tokens and $60 per million output tokens, against $2 and $10 on Standard. OpenAI calls Ultrafast “the fastest service tier in the OpenAI API” and says to use it “when speed justifies the higher cost”. It does not publish latency figures. We wanted to know what OpenAI Ultrafast pricing buys a WordPress chatbot, so on 11 October we sent the same chatbot request through all three tiers from our production server and timed every reply as it streamed.

The short version: for short support answers, Fast mode gave most of the speed for a third of Ultrafast’s extra cost, and Ultrafast did not run at all through the API endpoint most WordPress chatbot plugins use.

  • Per 1,000 new questions, GPT-6.1 Sol cost $10.92 on Standard, $21.66 on Fast and $66.18 on Ultrafast. The multipliers on the rate card held almost exactly per reply: 2.0x and 6.1x.
  • The first word arrived after 1.43 seconds on Standard, 0.84 s on Fast and 0.99 s on Ultrafast (means of nine streamed calls each). On short replies Ultrafast did not show text any sooner than Fast.
  • Ultrafast finished the whole reply first: 1.10 s against 1.39 s on Fast and 2.37 s on Standard. Its advantage is generation speed once text starts. On a 30-word answer that was worth 0.29 seconds over Fast, for an extra $44.52 per 1,000 replies.
  • Ultrafast was the most consistent. Its slowest first token was 1.21 s. Fast’s was 1.56 s and Standard’s 2.60 s.
  • Chat Completions rejected Ultrafast on 9 of 9 attempts with “Invalid service_tier argument”. It only worked through the Responses API. Fast mode worked on both.
  • GPT-6 Sol on Fast mode, through Chat Completions, was the quickest setup we ran: 0.71 s to the first token, 1.05 s to the end, at $21.52 per 1,000 replies.
Bar chart of GPT-6.1 Sol cost per 1,000 chatbot replies by OpenAI service tier on 11 October 2026: Standard $10.92, Fast mode $21.66, Ultrafast $66.18

OpenAI Ultrafast pricing next to Standard and Fast mode

These are the short-context rates per million tokens on OpenAI’s pricing page as we read it on 11 October 2026. Short context means requests of up to 272,000 input tokens, which covers any site chatbot we know of.

GPT-6.1 Sol, per 1M tokensInputCached inputCache writeOutputvs Standard
Standard$2.00$0.10$2.50$10.001x
Fast mode (formerly Priority)$4.00$0.20$5.00$20.002x
Ultrafast$12.00$0.60$15.00$60.006x
Ultrafast, long context (over 272K input)$24.00$1.20$30.00$90.006x the long-context Standard rate

Every line scales together. Cached input, cache writes and output all carry the same multiplier as input, so prompt caching saves the same share on Ultrafast as it does on Standard. In cash terms it saves more, because each cached token was more expensive to begin with.

For comparison, GPT-6 Astra on Ultrafast lists at $60 input and $300 output per million tokens, also six times its Standard rate. GPT-6 Sol, the model in MxChat’s own OpenAI list, has Fast mode at $4 and $20 but no Ultrafast tier: our test request for it came back with the same “Invalid service_tier argument” error. OpenAI’s Fast mode guide notes that Priority processing was renamed Fast mode on 30 July 2026, so older guides and code that say priority describe the same tier. Both values worked in our calls, and the API reported the tier back as fast either way.

How we tested

All calls went from our WordPress server in the US directly to OpenAI on 11 October 2026, using the API key our live chatbot uses. Each request carried our live chatbot system prompt (3,478 characters) and one of three real pre-sales questions:

  • Does MxChat work with WooCommerce, and can the bot look up an order status for a logged-in customer?
  • Do I need Pinecone to use the knowledge base, or can I store embeddings in WordPress?
  • Can a visitor be handed over to a human on Slack or Telegram when the bot cannot answer?

For each question we included four passages from our own documentation and blog posts, about 16,100 characters, as the knowledge base. These are the same prompt, passages and questions we used for our GPT-6.1 Sol pricing test on 1 October, which came to about 4,100 input tokens per request.

We set reasoning effort to low, the lowest value GPT-6.1 Sol accepts (it rejects none) and the value MxChat sends for OpenAI’s GPT-6 models. Every call streamed, so we could record two times: when the first visible word arrived and when the last one did. The first number is what a visitor watching the chat window notices. Each tier ran each question three times, nine calls per tier, through the Responses API. We then ran GPT-6.1 Sol and GPT-6 Sol on Standard and Fast through Chat Completions, nine calls each, plus nine attempts at Ultrafast that failed. That makes 63 calls that returned an answer.

Costs are computed from the usage block the API returned, at the list rates above. A “new question” is priced the way OpenAI billed the first sighting of each prompt: 3 tokens at the input rate and the rest as a cache write. We confirmed that usage shape on all three tiers with a separate long prompt. A “repeat” is the same prompt sent again and read from the cache.

What a reply costs on each tier

GPT-6.1 Sol, Responses APIStandardFast modeUltrafast
Per 1,000 new questions$10.92$21.66$66.18
Per 1,000, no cache write$8.87$17.56$53.90
Per 1,000 exact repeats (read from cache)$1.05$1.98$7.41
Mean output tokens (of which reasoning)67 (15)58 (11)79 (32)
Mean words in the answer343130
Output as a share of the bill7.5%6.6%8.6%
5,000 new questions a month$54.60$108.30$330.90

A chatbot bill is mostly input. The retrieved passages and system prompt were about 4,100 tokens and the answer about 60, so output was under a tenth of the cost on every tier. That is why the cost per reply tracks the input multiplier so closely. It also means a faster tier cannot make a reply cheaper by finishing sooner. You pay for the tokens, and the tokens were the same.

Ultrafast replies billed more reasoning tokens on average, 32 against 15 on Standard and 11 on Fast. With nine calls per tier we would not read much into that, but it pushed Ultrafast’s per-reply cost slightly above an exact six times Standard.

How much faster each tier was

Bar chart of time to first token and total reply time for GPT-6.1 Sol on 11 October 2026: Standard 1.43 s and 2.37 s, Fast mode 0.84 s and 1.39 s, Ultrafast 0.99 s and 1.10 s, GPT-6 Sol on Fast 0.71 s and 1.05 s
Mean of 9 streamed callsFirst tokenSlowest first tokenRest of the replyComplete replyPer 1,000 new questions
GPT-6.1 Sol, Standard (Responses)1.43 s2.60 s0.94 s2.37 s$10.92
GPT-6.1 Sol, Fast (Responses)0.84 s1.56 s0.55 s1.39 s$21.66
GPT-6.1 Sol, Ultrafast (Responses)0.99 s1.21 s0.11 s1.10 s$66.18
GPT-6.1 Sol, Standard (Chat Completions)1.90 s3.74 s0.76 s2.66 s$10.99
GPT-6.1 Sol, Fast (Chat Completions)1.47 s1.85 s0.60 s2.07 s$21.98
GPT-6 Sol, Standard (Chat Completions)1.17 s2.57 s0.31 s1.48 s$10.70
GPT-6 Sol, Fast (Chat Completions)0.71 s1.67 s0.33 s1.05 s$21.52

Ultrafast is very fast once it starts. It produced the rest of an answer in about a tenth of a second, where Standard took almost a second. What it did not do was start sooner than Fast mode. Most of the wait on a chatbot reply is before the first word, while the model reads 4,000 tokens of context and does its short burst of reasoning. Ultrafast spends that time at about the same pace as Fast.

That matters because chat widgets stream. A visitor sees the answer begin, and from then on the text is arriving faster than they can read it on any tier. So for a 30-word support answer, the number a visitor feels is roughly the first-token time, and on that number Fast mode matched Ultrafast at a third of the price. Ultrafast’s gain would be bigger on long replies. At Standard’s pace a 600-token answer would take several seconds to finish writing, and Ultrafast would cut most of that. Its tighter worst case is also a real benefit if a slow outlier is what you are trying to remove.

Network time is in all of these figures. They are full round trips from one server on one morning, not lab benchmarks, and a different region or time of day will move them. The differences between tiers were consistent across all three questions.

Ultrafast needs the Responses API

This was the finding we did not expect. Sending service_tier: "ultrafast" to the Chat Completions endpoint failed on every attempt with HTTP 400, “Invalid service_tier argument”, for GPT-6.1 Sol and GPT-6 Astra alike. The same request through the Responses API was accepted, and the response confirmed service_tier: "ultrafast". OpenAI’s Ultrafast guide only shows the Responses API and WebSockets, and it recommends WebSockets for agents that make many quick tool calls.

Many WordPress chatbot plugins, MxChat included, send chat through Chat Completions. On those, Ultrafast is not a setting you can switch on. Fast mode is. It accepted fast and priority on both endpoints. According to OpenAI’s Fast mode guide, it can also be set as a project default in your OpenAI settings, so requests that do not name a tier would get it. We did not test the project setting.

MxChat does not send a service tier today, so an MxChat site gets Standard unless the OpenAI project default says otherwise. Its OpenAI model list includes GPT-6 Sol, GPT-6 Astra and GPT-6 Luna. It does not yet list GPT-6.1 Sol.

Things to know before you pay for speed

The cache did not carry across tiers in our runs. The first Fast and Ultrafast calls on each question showed zero cached tokens, even though Standard had sent the identical prompt seconds earlier. The same happened between the Responses API and Chat Completions. If you mix tiers, expect each one to warm its own cache and pay its own cache writes. Our OpenAI prompt caching test covers what those writes cost.

Fast mode can quietly fall back to Standard. OpenAI’s guide says that if traffic ramps too quickly, some Fast requests are served at standard speed and billed at the standard rate, with service_tier: "default" in the response. It suggests growing traffic by no more than 50% every 15 minutes once you pass a million input tokens a minute. A site chatbot will rarely get near that, but it is worth logging the tier the API reports, not the one you asked for. Every one of our calls was served on the tier requested.

Ultrafast has its own rate limits. OpenAI lists 1,000,000 tokens per minute for GPT-6.1 Sol on its Build usage tier, 4,000,000 on Launch and 40,000,000 on Grow. At about 4,100 tokens per chatbot request, the Build limit is roughly 240 requests a minute.

One answer pattern to watch. On the WooCommerce question, Ultrafast replied “I don’t have enough information” on two of its three runs and Fast mode on one of three, where Standard answered all three. Every other answer on every tier was correct against the source passages. Three runs per question is too few to say whether the tier had anything to do with it. Reasoning models at the default temperature do vary from call to call. We are noting it rather than drawing a conclusion.

Bar chart of the extra cost of faster OpenAI service tiers per 1,000 chatbot replies on 11 October 2026: Fast mode plus $10.74 for 0.98 s faster, Ultrafast plus $55.26 for 1.27 s faster, Ultrafast over Fast plus $44.52 for 0.29 s

Which tier makes sense for a WordPress chatbot

Your situationTier we would pickWhy
A support or pre-sales bot with short answersStandard, or Fast if replies feel slowFast cut the first-token wait by 0.6 s for about $10.74 per 1,000 replies
A plugin that uses Chat CompletionsStandard or FastUltrafast is rejected on that endpoint
Long answers such as product guides or generated draftsConsider UltrafastIts generation speed matters more as answers get longer
A voice or real-time agent that cannot tolerate outliersConsider UltrafastIt had the tightest worst case in our test
Speed matters most and cost matters tooA faster model on Fast modeGPT-6 Sol on Fast beat GPT-6.1 Sol on Ultrafast for a third of the price
Background jobs such as summaries or taggingBatch or FlexHalf the Standard price if nobody is waiting; see our Batch API cost test

The last row in our latency table is the one we would point most site owners to. GPT-6 Sol on Fast mode replied in 1.05 seconds end to end, slightly faster than GPT-6.1 Sol on Ultrafast, and cost $21.52 per 1,000 new questions instead of $66.18. Choosing a quicker model gave as much speed as the most expensive tier. If you are deciding between GPT-6 Sol and its rivals, our GPT-6 Sol vs Claude Sonnet 5.5 comparison has the per-reply numbers, and the GPT-6 Astra pricing test covers the top of OpenAI’s range.

MxChat lets you choose the OpenAI model per bot, and the setup steps are in the MxChat documentation. The WooCommerce, live-agent and other add-ons mentioned in the test questions come with MxChat Pro. For request routing that does not need a full reply at all, our Decisions API test found answers in about a tenth of a second.

FAQ

How much does OpenAI Ultrafast cost?

For GPT-6.1 Sol, $12 per million input tokens, $0.60 cached input, $15 per million cache-write tokens and $60 per million output tokens on requests up to 272,000 input tokens, as listed on 11 October 2026. That is six times the Standard rate. GPT-6 Astra on Ultrafast is $60 input and $300 output.

Is Ultrafast faster than Fast mode?

It finished replies faster in our test, 1.10 s against 1.39 s, because it generates text much more quickly once it starts. It did not show the first word sooner: 0.99 s against 0.84 s on Fast. For short streamed chatbot answers the difference a visitor notices was small.

What is the difference between Fast mode and Priority processing?

They are the same tier. OpenAI renamed Priority processing to Fast mode on 30 July 2026. The API accepts service_tier set to fast or priority, and in our calls both were reported back as fast. It costs twice the Standard rate.

Can I use Ultrafast with Chat Completions?

Not in our test. Every Chat Completions request with service_tier: "ultrafast" was rejected with “Invalid service_tier argument”. The Responses API accepted it. OpenAI’s guide documents Ultrafast for the Responses API and WebSockets.

Which models support Ultrafast?

OpenAI’s guide lists GPT-6 Astra and GPT-6.1 Sol. A GPT-6 Sol request with the Ultrafast tier was rejected in our test.

Does a faster tier make each reply cheaper?

No. You pay per token at the tier’s rate, and the token counts were similar on every tier. A chatbot reply is mostly input, so the cost per reply followed the input multiplier: about 2x on Fast and 6x on Ultrafast.

Method: 63 answered API calls plus probes from the mxchat.ai server on 11 October 2026, GPT-6.1 Sol and GPT-6 Sol, reasoning effort low, max output 4,000 tokens, streamed, through the Responses API and Chat Completions. Prices are list rates read from OpenAI’s pricing page the same day. Times are full round trips including network from one US server.