OpenAI Decisions API Pricing: Faster, Not Always Cheaper
OpenAI launched the Decisions API in public beta on 6 October 2026. It does one narrow job: you send text (or an image) and a short list of questions, and it returns typed answers, such as a yes/no probability, one choice from a fixed list, or a score on a scale. It runs only on GPT-6 Luna, and the pricing is unusual. You pay $0.10 per million input tokens and nothing for output, because there is no output text to bill.
For a WordPress chatbot, that is the shape of several small jobs that run around every reply: which support queue does this message belong in, does the visitor need a human, can the retrieved help pages answer the question at all. We wanted to know what the OpenAI Decisions API costs for those jobs in practice, and whether it beats doing the same thing with Chat Completions on the same model. So on 7 October 2026 we sent 144 requests from our production WordPress server to both endpoints and priced every one from the usage block OpenAI returned.
- It is fast. The median Decisions request took 0.11 seconds from our server, round trip. The same routing job on Chat Completions took 0.97 seconds at effort none and 1.44 seconds at the default effort. OpenAI’s own processing time was about 50 milliseconds against 850 to 1,300.
- It is not automatically cheaper. Asking one question (which queue?) cost $0.020 per 1,000 messages. Asking two (queue and human handoff) cost $0.038, which is more than Chat Completions at effort none answering both in one JSON reply ($0.030). Each question you add costs about 182 input tokens.
- It does not cache. We sent the same 3,400-token request twice. Chat Completions served the repeat from cache for $0.040 per 1,000. Decisions reported 0 cached tokens and charged $0.350 again.
- It was the most accurate setup we tested. 30 of 30 queues and 30 of 30 handoff calls correct. Chat Completions missed one of each, depending on the effort setting.
- The money is small either way. At 10,000 chats a month, two-question routing costs about $0.38 on Decisions and $0.30 on Chat Completions. The case for the new API is the speed and the probabilities, not the bill.
OpenAI Decisions API pricing: the rate card
OpenAI’s pricing page and the Decisions guide give the rates below. The Decisions rate applies only to requests sent to /v1/decisions. The same model through Chat Completions or the Responses API is billed at its normal rates, which include output tokens and a prompt cache.
| Endpoint and model | Input / 1M | Cached input / 1M | Cache write / 1M | Output / 1M |
|---|---|---|---|---|
| Decisions API, gpt-6-luna | $0.10 | no cache | none | none |
| Chat Completions, gpt-6-luna | $0.10 | $0.01 | $0.125 | $0.50 |
| Chat Completions, gpt-6-sol | $2.00 | $0.20 | $2.50 | $10.00 |
Two things stand out. First, the input price is the same as GPT-6 Luna’s normal input price. Decisions does not make input cheaper; it removes the output charge. Second, OpenAI says long-context multipliers and regional processing premiums also apply to Decisions, as they do elsewhere. Our requests were all short-context and unpinned to a region.
We also checked which models the endpoint accepts. A request naming gpt-6.1-sol returned HTTP 404, “The model `gpt-6.1-sol` does not exist or you do not have access to it.” For now it is GPT-6 Luna or nothing.
How we tested it
Every request went from the server that runs mxchat.ai, using the OpenAI key our own chatbot uses, on the morning of 7 October 2026. We ran two jobs.
Routing. We wrote 30 visitor messages of the kind a WordPress plugin support bot actually gets: a duplicate charge, a missing license key, a blank chat widget after an update, a 500 error in admin-ajax.php, presales questions about WooCommerce and Gemini support, a university asking about 40 sites, a poem request, “lol”, and “Can I talk to a real human please?” Each message had a correct queue (billing, setup, presales, bug, other) and a correct yes or no for “hand this to a human now”. Seven of the 30 needed a human. We sent every message four ways:
- Decisions API with one choice question: which queue.
- Decisions API with two questions in one request: which queue, and a yes/no predicate for human handoff.
- Chat Completions on GPT-6 Luna at
reasoning_effort: none, with a strict JSON schema returning both fields. - The same Chat Completions request with no effort set, so the model’s default (medium) applied.
The queue descriptions and the handoff rule were word for word the same in both APIs. Decisions takes them as structured choices and instructions; Chat Completions got them in the system prompt.
Answerability check. Many chatbots answer from retrieved documents, and a common step is to ask first whether those documents can answer the question, so the bot can hand off or say “I don’t know” instead of guessing. We reused the same retrieved help-page windows (about 3,400 tokens each) and questions as our GPT-6 Luna pricing test, plus three questions the pages could not answer, such as “Do you have a phone number I can call for support?” Each of the six went to Decisions as a single predicate and to Chat Completions (effort none, JSON schema) as a true/false field. We sent all six twice, to see what caching did.
Costs below come from the token counts in each response, priced at the list rates above. OpenAI does not return a dollar figure on either endpoint.
Cost per 1,000 routing decisions
| Setup | Input tokens (mean) | Output tokens (mean) | Cost per 1,000 | Queue correct | Handoff correct |
|---|---|---|---|---|---|
| Decisions, queue only | 201 | 0 | $0.020 | 30 / 30 | not asked |
| Decisions, queue + handoff | 383 | 0 | $0.038 | 30 / 30 | 30 / 30 |
| Chat Completions, effort none | 219 | 17 | $0.030 | 29 / 30 | 30 / 30 |
| Chat Completions, default effort | 219 | 45 (21 reasoning) | $0.044 | 30 / 30 | 29 / 30 |
The input counts explain the result. Decisions bills only input, but the questions themselves are input, and the endpoint wraps each one in more tokens than a plain system prompt does. Our one-question request came to about 201 tokens. Adding the handoff predicate, whose instructions are 32 words long, took it to about 383. Chat Completions carried the same queue list and the same handoff rule in 219 tokens, then spent 17 tokens writing {"intent":"billing","needs_human":true}. At $0.50 per million, those 17 output tokens cost less than the 164 extra input tokens Decisions needed.
So the pricing has a crossover. For one question, Decisions was the cheapest setup we measured. For two questions in one request, Chat Completions at effort none was about 20% cheaper. The more questions you stack into one Decisions call, the more the input overhead adds up, because nothing is saved on output when the output was only 17 tokens to begin with.
The setup Decisions clearly beats is Chat Completions at the default effort. Left at medium, GPT-6 Luna spent an average of 21 reasoning tokens deciding where a message about a license key should go, and up to 68 on a single message. That made it the most expensive and the slowest of the four, and it was no more accurate. We saw the same pattern in our GPT-6 Luna pricing test, where the default effort cost 7% more on full replies and added no quality. A Decisions request has no reasoning tokens to spend, so that choice is made for you.
Accuracy, and what the probabilities are good for
Decisions got every queue and every handoff right on our set. Chat Completions at effort none put “Does the Agency plan renew automatically every year?” in presales rather than billing, which is arguable. At the default effort it flagged the university asking about a custom deal for 40 sites as not needing a human, even though the handoff rule names custom-sales requests. Thirty messages is a small set, so treat this as “Decisions was at least as accurate”, not as a measured accuracy gap.
The bigger difference is what comes back. Chat Completions returns true or false. Decisions returns a probability for every predicate and a probability for every choice, so you can set your own threshold. On our set the handoff probabilities were cleanly separated: the seven messages that needed a human scored 0.94 to 1.00, and the other 23 scored 0.00 to 0.02. On the answerability check, the three answerable questions scored 1.00 and the three unanswerable ones 0.13 to 0.39. Any threshold around 0.5 would have worked.
Be careful with the queue confidence field, though. All 30 queue choices were right, but 19 of them came back with confidence below 0.9, mostly where two queues overlap: “Embeddings are not being created in Pinecone” is a bug report and a setup question, and Decisions picked bug with confidence 0.31. Only two of the 30 fell below 0.5. If you route low-confidence messages to a human, set the cut-off from your own messages, or you will send most of your traffic to the person you were trying to save time for.
One more detail for anyone writing the integration. In a separate probe we asked eight questions in one request, six of them deliberately vague (“Is the message about topic 4?”). Seven came back answered. One came back as "type": "refusal" with no probability, even though nothing in it touched a content policy. Your code needs to handle a refusal on any question, not just on obviously unsafe input.
Answerability checks and the missing cache
This is where the two APIs part ways. A check over retrieved sources is almost all input: about 3,400 tokens of help pages and a one-line question. Both endpoints charge $0.10 per million for that input, so on a first request the costs were close: $0.350 per 1,000 on Decisions and $0.347 on Chat Completions before caching.
Caching changes both sides. Chat Completions on GPT-6 Luna writes long prompts to its cache automatically and bills the write at 1.25 times the input rate. In a separate probe with a fresh 1,800-token prompt, the first call reported 1,797 cache-write tokens and the second and third reported 1,797 cached tokens. On our 3,400-token checks that makes the first request about $0.432 per 1,000, and every exact repeat $0.040, because 3,290 to 3,570 tokens were read from cache at $0.01 per million. We covered how this write charge works in our OpenAI prompt caching test.
Decisions reported cached_tokens: 0 and cache_write_tokens: 0 on all 20 requests that repeated an earlier input, at 1,900 to 3,700 tokens. It never writes a cache entry and never reads one, so the repeat costs exactly what the first request cost.
How much that matters depends on how much of your prompt repeats. In a chatbot, the retrieved sources change with every question, so the expensive part rarely repeats exactly, and both APIs would bill most of it at full price. What does repeat is a long fixed prefix: a system prompt, a policy document, a product catalogue you send with every request. If your check sends the same 2,000-token policy text before each message, Chat Completions bills that prefix at a tenth of the price after the first call. Decisions bills it in full every time.
| Answerability check (~3,400 tokens) | Cost per 1,000 | Median time | Correct |
|---|---|---|---|
| Decisions, first request | $0.350 | 0.11 s | 6 / 6 |
| Decisions, exact repeat | $0.350 | 0.11 s | 6 / 6 |
| Chat Completions, first request (with cache write) | $0.432 | 0.82 s | 6 / 6 |
| Chat Completions, exact repeat (cached) | $0.040 | 0.82 s | 6 / 6 |
One thing worth noticing in the table: an answerability check over the same sources the bot will answer from costs a large share of the answer itself. Our GPT-6 Luna test priced a full retrieval reply at $0.57 per 1,000. A $0.35 check in front of it adds about 61%, because it sends the sources a second time. That is worth paying only if the check saves you replies, for example by handing unanswerable questions straight to a person or a contact form instead of generating a polite non-answer.
Speed: the real reason to use it
OpenAI says Decisions is “10x faster than the Responses API”. From our server, the median was 0.11 seconds for every Decisions setup, including the 3,400-token checks, against 0.82 to 1.44 seconds on Chat Completions. That is a 7x to 13x gap in wall-clock time, and the header OpenAI returns with each response (openai-processing-ms) showed a median of about 50 milliseconds on Decisions against 725 to 1,326 on Chat Completions. The slowest Decisions call took 0.42 seconds. The slowest Chat Completions call took 3.78.
For a chat widget, that is the useful number. Routing and handoff checks run before the bot can reply. A one-second check before a one-second answer makes the visitor wait twice as long. A tenth of a second does not show.
What it costs at real chatbot volumes
| Chats per month | Decisions, queue only | Decisions, queue + handoff | Chat Completions, none | Decisions, sources check |
|---|---|---|---|---|
| 1,000 | $0.02 | $0.04 | $0.03 | $0.35 |
| 10,000 | $0.20 | $0.38 | $0.30 | $3.50 |
| 100,000 | $2.01 | $3.83 | $3.05 | $34.98 |
Routing costs are close to nothing on either API. A site handling 100,000 chats a month would pay under $4 to route them all. The answerability check is a real cost at scale, because it is input-heavy, and that is also the job where losing the prompt cache can hurt if your prefix repeats.
When to use the Decisions API in a WordPress chatbot
Based on what we measured, these are the cases where it makes sense:
- Anything that runs before the reply and has to be quick. Queue routing, “does this need a human”, spam and abuse screening, language detection. Saving 0.8 to 1.3 seconds per message is the main benefit.
- One or two questions per request. One question was the cheapest setup we tested. At two it was slightly more expensive than Chat Completions, and each additional question added about 182 input tokens.
- When you want a threshold, not a yes or no. Probabilities let you send a borderline handoff to a human and a clear one to the bot, and adjust the line later without rewriting a prompt.
And the cases where Chat Completions is still the better choice:
- Long fixed instructions sent with every request. Chat Completions caches them at a tenth of the input price after the first call. Decisions charges full price every time.
- Several questions where speed does not matter. One short JSON reply carries many fields for fewer tokens than the same questions as separate Decisions entries.
- Any model other than GPT-6 Luna. Decisions runs only on Luna for now.
- Anything that needs the answer as text. Decisions returns probabilities and choices, never a sentence, so it cannot write the reply itself.
One practical caution: the API is in beta. OpenAI says general availability is expected in the coming weeks, and the request format could change before then. If you build on it now, keep a Chat Completions fallback for the same job.
What this means for MxChat
MxChat already offers live agent handoff, answers from your own knowledge base, and runs on your own OpenAI key. The Decisions API is a separate endpoint, not a new chat model, so it is not something you pick in a model list. It is a candidate for the small checks that run around the reply. If you want the handoff and knowledge-base features today, see the MxChat documentation or MxChat Pro. For how AI and human support compare in general, see our AI chatbots vs human agents comparison, and for the larger OpenAI model on the same chatbot request, our GPT-6.1 Sol pricing test.
Frequently asked questions
How much does the OpenAI Decisions API cost?
$0.10 per million input tokens on GPT-6 Luna, with no charge for output, according to OpenAI’s Decisions guide on 7 October 2026. Long-context and regional processing multipliers can apply. In our test, one routing question cost $0.020 per 1,000 messages and two questions $0.038.
Is the Decisions API cheaper than Chat Completions?
Not always. The input price is the same as GPT-6 Luna’s normal rate. Decisions saves the output charge, but adds input tokens for each question. With one question it was cheaper than any Chat Completions setup we tried. With two, Chat Completions at effort none was about 20% cheaper. With a long repeated prompt, Chat Completions can be much cheaper because Decisions does not cache.
Does the Decisions API support prompt caching?
We saw no caching. On 20 requests that repeated an earlier input, the response reported 0 cached tokens and 0 cache-write tokens, so each repeat cost the same as the first request. Chat Completions served the same repeats from cache at $0.01 per million tokens.
Which models work with the Decisions API?
Only gpt-6-luna during the beta. A request with gpt-6.1-sol returned HTTP 404 on 7 October 2026.
How fast is the Decisions API?
The median request took 0.11 seconds from our WordPress server, with about 50 milliseconds of processing time reported by OpenAI. The same jobs on Chat Completions took 0.82 to 1.44 seconds.
What kinds of questions can it answer?
Three types: a predicate (a probability that something is true), a choice (one value from a list you supply, with a probability for each), and a score (a position on an ordered scale of levels you define). You can include several in one request, and each is answered independently.