Mistral API Pricing: Medium 3.5 Costs 3x Large, Tested
Mistral’s API price list has an odd shape this autumn. Mistral Medium 3.5 costs $1.50 per million input tokens and $7.50 per million output tokens. Mistral Large 3 costs $0.50 and $1.50. So the model called Medium is three times the price of the model called Large on input and five times on output. If you are choosing a Mistral model for a WordPress chatbot, the name tells you very little about the bill.
We wanted numbers rather than a rate card, so on 5 October 2026 we sent the same three customer-support questions, with the same system prompt and the same retrieved help pages, to five Mistral models from our production WordPress server. This is what Mistral API pricing works out to on a real retrieval chatbot, per 1,000 replies, as billed.
- Mistral Medium 3.5 cost $6.78 per 1,000 chatbot replies. Large 3 cost $2.08, Small 4 $0.61 and Ministral 3 14B $0.45.
- Input was 95% of the Medium 3.5 bill. A support bot sends about 4,300 tokens of instructions and retrieved text to get back a 25-word answer, so the input rate is the price that matters.
- Medium 3.5’s prompt cache did not lower the bill through OpenRouter. The second pass reported about 4,250 cached tokens and was billed at the full $1.50 per million. On Medium 3.1 and Ministral 3 14B, the same cache hits cut the bill by 85%.
- Turning on reasoning added 44% to Medium 3.5 ($9.75 per 1,000 replies) and tripled its response time, for answers that were not better.
- The most expensive model did not give the best answers. Medium 3.5 averaged 24 words a reply and skipped details the cheaper models included. Large 3 was more complete, but it linked customers to a page on our site that does not exist.
Mistral API pricing: the current rate card
These are the per-million-token rates from Mistral’s API pricing page, read on 5 October 2026, and the matching OpenRouter listings. Mistral lists a 90% discount on cached input across its range and half price for batch jobs. Mistral Medium 3.1 is no longer on Mistral’s own price page but is still served on OpenRouter, so we tested it as the reference point for the older Medium price.
| Model | Input / 1M | Output / 1M | Cached input (Mistral’s page) | Cache read on OpenRouter | Context |
|---|---|---|---|---|---|
| Mistral Medium 3.5 | $1.50 | $7.50 | 90% off | Not listed | 256K |
| Mistral Large 3 | $0.50 | $1.50 | 90% off | $0.05 | 256K |
| Mistral Small 4 | $0.15 | $0.60 | 90% off | $0.015 | 256K |
| Ministral 3 14B | $0.20 | $0.20 | Not stated | $0.02 | 256K |
| Ministral 3 8B | $0.15 | $0.15 | Not stated | $0.015 | 256K |
| Ministral 3 3B | $0.10 | $0.10 | Not stated | $0.01 | 128K |
| Mistral Medium 3.1 (older, OpenRouter only) | $0.40 | $2.00 | n/a | $0.04 | 128K |
Two things stand out. Medium 3.5, released at the end of April 2026 as a 128-billion-parameter dense model aimed at coding and agent work, costs 3.75 times what Medium 3.1 cost on input. And the Ministral models charge the same rate for input and output, which is unusual and happens to suit chatbots, where the input is long and the reply is short.
How we tested
The test is the same one we have used for every model in this series, so the numbers line up with our earlier posts on Claude Haiku 4.5, GPT-6 Luna and Gemini Flash-Lite.
- Same request shape as MxChat sends: our production system prompt (3,478 characters) as the system message, then a user message holding four retrieved help pages (about 16,100 characters) and the visitor’s question.
- Three questions: whether the bot can look up a WooCommerce order status, whether you need Pinecone for the knowledge base, and whether a visitor can be handed to a human on Slack or Telegram.
- Five models, each asked every question twice. We also ran Medium 3.5 and Small 4 with reasoning set to
high, and ran two-turn conversations on Medium 3.5 and Large 3 with the follow-up “Is that included in the free plugin, or do I need MxChat Pro for it?” - Through OpenRouter, pinned to Mistral as the provider with fallbacks off, so every call ran on Mistral’s own servers. We used OpenRouter because MxChat supports an OpenRouter key and it is the easiest way for a WordPress site to reach Mistral. Costs are what OpenRouter billed per call, read from the
usage.costfield.
That came to 45 chatbot replies plus 4 probe calls. Many more attempts on Large 3 and Small 4 came back as HTTP 429 “temporarily rate-limited upstream”, which we cover below because it matters for a live chatbot.
Cost per 1,000 chatbot replies
| Model | Per 1,000 replies (billed) | Prompt tokens | Reply tokens | Avg. words | Avg. time |
|---|---|---|---|---|---|
| Mistral Medium 3.5 | $6.78 | 4,315 | 41 | 24 | 1.17 s |
| Medium 3.5, reasoning high | $9.75 | 4,315 | 437 (482 reasoning) | 12 | 3.61 s |
| Mistral Large 3 | $2.08 | 4,268 | 56 | 34 | 1.71 s |
| Mistral Small 4 | $0.61 | 4,315 | 27 | 19 | 0.82 s |
| Small 4, reasoning high | $0.72 | 4,315 | 245 (234 reasoning) | 31 | 2.01 s |
| Ministral 3 14B | $0.45 | 4,303 | 78 | 47 | 1.32 s |
| Mistral Medium 3.1 | $0.26 | 4,315 | 33 | 19 | 0.82 s |
The ranking follows the input rate almost exactly. Every model saw about 4,300 prompt tokens and wrote 20 to 80 tokens back. On Medium 3.5 the input cost $6.47 of the $6.78, or 95%. Output price, which is what most comparison tables lead with, barely moves a support bot’s bill. That is the same pattern we found with DeepSeek V4.1 Flash and the Gemini Flash-Lite models: on retrieval chatbots, input is the price.
Two rows are lower than the rate card suggests. Ministral 3 14B and Medium 3.1 both read most of the prompt from cache on repeat questions and were billed the discounted cache rate. Without that, Ministral 3 14B costs $0.88 and Medium 3.1 $1.79 per 1,000 replies, which is the fairer comparison for a site where most questions are new. Even at $1.79, Medium 3.1 is about a quarter of Medium 3.5’s cost on the same request.
Mistral’s tokenizer was consistent across the line-up. The identical request counted as 4,184 to 4,486 tokens on all five models, depending on the question. For comparison, Claude Sonnet 5.5 counted the same text as 6,767 to 6,897 tokens in our tokenizer test, so Mistral is not paying a token-count penalty here.
Mistral Medium 3.5 vs Large 3: the naming trap
Medium 3.5 is Mistral’s newer, larger dense model, built for agent and coding work. Large 3, from December 2025, is a mixture-of-experts model priced for volume. For a chatbot answering questions from retrieved pages, Large 3 was the better buy in our test on every measure except speed.
| Measure | Mistral Medium 3.5 | Mistral Large 3 |
|---|---|---|
| Rate card (input / output) | $1.50 / $7.50 | $0.50 / $1.50 |
| Per 1,000 replies, billed | $6.78 | $2.08 |
| Average reply length | 24 words | 34 words |
| Average response time | 1.17 s | 1.71 s |
| Cache discount through OpenRouter | None seen | Applied (768 cached tokens billed at $0.05) |
| Rate-limit errors in our run | 0 | Yes, repeatedly |
| Invented links | 0 | 3 to a 404 page, plus 1 guessed /pricing/ URL |
On the answers themselves, Large 3 said in its WooCommerce reply that order lookups need MxChat Pro and that the bot shows order date, items and totals. Medium 3.5 said only that the add-on exists and that the bot can look up orders. On the Pinecone question, Large 3 gave the 50,000-entry guidance from our docs; Medium 3.5 gave one sentence. On the human-handoff question, Large 3 explained how Slack and Telegram each receive the conversation and that both are in the free plugin; Medium 3.5 gave the settings path. Neither was wrong, but Medium 3.5 was the less useful answer at three times the price.
Medium 3.5 also matched Medium 3.1’s answer word for word on the WooCommerce question in one of the two runs. For this kind of short, grounded answer, we could not see what the newer model adds.
The cache that did not discount
Prompt caching is the main lever for chatbot costs, because the system prompt is identical on every request and follow-up turns repeat the whole conversation. Mistral’s price page lists 90% off cached input. Through OpenRouter, that discount showed up on some models and not on Medium 3.5.
We sent the same requests twice. On the second pass, OpenRouter reported about 4,250 of the roughly 4,300 prompt tokens as cached for Medium 3.5, Medium 3.1 and Ministral 3 14B alike. The bills went three different ways:
| Model | No cache hit, per 1,000 | Cache hit, per 1,000 | Change |
|---|---|---|---|
| Mistral Medium 3.5 | $6.79 | $6.74 | None (cost difference is shorter replies) |
| Ministral 3 14B | $0.79 | $0.12 | −85% |
| Mistral Medium 3.1 | $1.79 (rate card) | $0.26 | −86% |
The arithmetic is exact. A Medium 3.5 call with 4,196 prompt tokens and 39 reply tokens was billed $0.0065865, which is 4,196 × $1.50 + 39 × $7.50 per million. It cost the same with zero cached tokens and with 4,096 cached tokens. OpenRouter’s endpoint listing for Medium 3.5 shows no cache-read price at all, while Large 3, Small 4, the Ministral models and Medium 3.1 each list one.
We could not test a direct Mistral key, so we cannot say whether Mistral’s own billing applies the 90% discount to Medium 3.5. If you run Medium 3.5 through OpenRouter, assume no cache discount until the listing shows a cache-read rate. A two-turn conversation on Medium 3.5 cost $6.71 for the first turn and $6.70 for the follow-up, so every turn costs the same as a new question.
Reasoning mode: more cost, shorter answers
Both Medium 3.5 and Small 4 accept a reasoning setting. With reasoning.effort set to high, Medium 3.5 used an average of 482 reasoning tokens per reply, pushing the cost from $6.78 to $9.75 per 1,000 replies and the response time from 1.17 to 3.61 seconds. Its visible answers got shorter, averaging 12 words, and lost the setup paths and links the non-reasoning answers included. One handoff answer was simply “Yes, both Slack and Telegram live agent handoffs are available.”
Small 4 behaved better with reasoning on. Its cost rose only from $0.61 to $0.72, because its output is cheap, and its answers improved: the handoff reply explained how Telegram and Slack each receive the conversation, where the non-reasoning replies were a bare “Yes”. At 2 seconds it is slower, but if you are running Small 4 and find its answers too thin, reasoning is a cheap fix. On Medium 3.5 we would leave it off.
Answer quality across the range
All the answers we read were grounded in the retrieved pages. None claimed a feature that the sources did not support on the first turn. The differences were in completeness and in links:
- Ministral 3 14B gave the longest and most complete answers at 47 words on average: order history limits, the 50,000-entry Pinecone guidance and the free-plugin status of live handoff. For the price, it was the strongest result in the test.
- Mistral Large 3 was close behind on content but invented URLs. It linked to
/add-ons/woocommerce/, which returns 404, in three replies, and to a/pricing/page in a fourth that only works because our site happens to redirect that path. A bot that links customers to a missing page is worse than one that gives no link. - Mistral Small 4 without reasoning was accurate but thin. Its Pinecone answer (“WordPress database storage is supported as an option”) did not tell the visitor whether they need Pinecone.
- Mistral Medium 3.5 and Medium 3.1 gave short, correct answers with links to pages that exist.
One more result is worth knowing if your system prompt includes a sales line. Ours tells the bot to mention a 15% discount. None of the 45 Mistral replies used it. When we ran the same test on Claude Sonnet 5.5, it brought up the discount in 14 of 18 replies. If your bot is meant to sell as well as answer, test that behaviour before you switch models.
Rate limits on shared OpenRouter capacity
Our first pass sent 21 requests to Large 3 and Small 4. Fifteen came back as HTTP 429 with the message that the model was “temporarily rate-limited upstream”, and OpenRouter suggested adding your own Mistral key to get your own limits. Retrying after 15 seconds worked for most, but three requests failed six retries in a row. Medium 3.5, Medium 3.1 and Ministral 3 14B returned no 429s in the same session.
For a chatbot, a 429 is a visitor waiting on a spinner or seeing an error. If you put Large 3 or Small 4 behind a live chat through OpenRouter, add your own Mistral key in OpenRouter’s integration settings or pick a fallback model, and keep an eye on error rates for the first few days.
How Mistral compares with the other models we have priced
Every model in this chart answered the same three questions with the same system prompt and sources, from the same server, between 24 September and 5 October 2026. The non-Mistral figures were billed directly by each vendor.
| Model | Per 1,000 replies | Source |
|---|---|---|
| Ministral 3 14B | $0.45 | This test |
| GPT-6 Luna | $0.57 | GPT-6 Luna test |
| Mistral Small 4 | $0.61 | This test |
| Gemini 3.5 Flash-Lite | $1.35 | Gemini Flash-Lite test |
| DeepSeek V4.1 Flash (peak hours) | $1.53 | DeepSeek Flash test |
| Mistral Large 3 | $2.08 | This test |
| Claude Haiku 4.5 | $4.93 | Claude Haiku 4.5 test |
| Mistral Medium 3.5 | $6.78 | This test |
| GPT-6 Sol (effort none) | $10.67 | GPT-6 Sol vs Sonnet 5.5 |
| Claude Sonnet 5.5 (effort low) | $13.02 | Claude Sonnet 5.5 test |
Ministral 3 14B and Small 4 sit in the budget tier with GPT-6 Luna, and Ministral was the cheapest model we have measured on this request. Large 3 lands between the budget models and Claude Haiku 4.5. Medium 3.5 costs more than Haiku 4.5 and about two-thirds of GPT-6 Sol, which is a lot to pay for 24-word answers.
What it costs a real site
A WordPress store answering 5,000 chatbot questions a month, with prompts the size of ours, would pay roughly:
| Model | 5,000 replies / month | 50,000 replies / month |
|---|---|---|
| Ministral 3 14B (no cache hits) | $4.40 | $44 |
| Mistral Small 4 | $3.05 | $31 |
| Mistral Large 3 | $10.40 | $104 |
| Mistral Medium 3.5 | $33.90 | $339 |
These figures cover the chat model only. They leave out the embedding calls that find the right help pages, and they assume answers of 20 to 80 tokens. A bot that writes long replies, or sends a longer system prompt, will cost more, roughly in proportion to the extra tokens.
Which Mistral model to use for a WordPress chatbot
- Start with Ministral 3 14B. It was the cheapest and gave the most complete answers in our test, and its flat $0.20 rate suits a bot that reads a lot and writes a little.
- Use Small 4 with reasoning on if you want Mistral’s newer small model. It costs $0.72 per 1,000 replies with reasoning and answers better than it does without.
- Use Large 3 if you need more capacity, but check every link it produces and plan for rate limits on shared OpenRouter capacity.
- Skip Medium 3.5 for support chat. It is priced for agent and coding work, its answers here were the shortest, and its cache gave no discount through OpenRouter.
MxChat connects to Mistral models through an OpenRouter key, or through the OpenAI-compatible custom endpoint if you use your own Mistral key. Set it up under MxChat → Settings → API Keys; the MxChat documentation covers model selection, and MxChat Pro adds the WooCommerce order features our first test question asked about. Whichever model you pick, set a spending limit in your provider’s dashboard and run a week of real questions before you trust any rate card, including ours.
Frequently asked questions
How much does the Mistral API cost?
As of 5 October 2026, Mistral Medium 3.5 costs $1.50 per million input tokens and $7.50 per million output tokens, Mistral Large 3 costs $0.50 and $1.50, Mistral Small 4 costs $0.15 and $0.60, and the Ministral 3 models cost $0.10 to $0.20 for both input and output. Mistral lists 90% off cached input and half price for batch jobs.
Why is Mistral Medium more expensive than Mistral Large?
The names reflect product lines, not price order. Medium 3.5 is a newer 128B dense model released in April 2026 for coding and agent work, priced at $1.50 / $7.50. Large 3 is an older mixture-of-experts model priced at $0.50 / $1.50. In our chatbot test Medium 3.5 cost $6.78 per 1,000 replies against $2.08 for Large 3.
What is the cheapest Mistral model for a chatbot?
In our test, Ministral 3 14B at $0.45 per 1,000 replies as billed, or $0.88 with no cache hits. Mistral Small 4 cost $0.61. Ministral 3 3B and 8B are cheaper per token but we did not test their answer quality.
Does Mistral prompt caching work through OpenRouter?
For Large 3, Small 4, the Ministral models and Medium 3.1, yes: cached tokens were billed at the discounted rate and cut repeat-question costs by about 85%. For Medium 3.5, no: OpenRouter reported cached tokens but billed them at the full $1.50 per million, and its listing shows no cache-read price.
Should I turn on reasoning for a Mistral chatbot?
On Small 4, it is worth trying: it added about $0.11 per 1,000 replies and gave fuller answers. On Medium 3.5 it added 44% to the cost and tripled response time while making answers shorter.
Can I use Mistral with MxChat?
Yes. MxChat supports an OpenRouter key, which reaches every Mistral model in this post, and any OpenAI-compatible endpoint. Your provider bills you directly for usage.