Mistral Large 4 Pricing: The Half-Price Preview, Tested
Mistral released Mistral Large 4 as a public preview on 6 October 2026 and priced it at half its list rate: $0.68 per million input tokens and $2.09 per million output tokens, against a list price of $1.36 and $4.18. Mistral’s changelog says the launch pricing is “50% off for 2 weeks”. That points to the discount ending around 20 October, though Mistral has not published an exact end date. Anyone choosing a model for a WordPress chatbot this month is therefore looking at a price that will double within days. We wanted to know what Mistral Large 4 pricing means per reply, both now and after the sale, so we ran it from our production server on 10 October with the same chatbot request we used for our Mistral API pricing test five days earlier.
The rate card is not the main story. Mistral Large 4 reasons before it answers unless you tell it not to, and on a short support question that reasoning was most of the output we paid for.
- At the sale price with reasoning switched off, Large 4 cost $2.97 per 1,000 replies with no cache hits, 31% more than Mistral Large 3 at $2.26 on the same request.
- With default settings it cost $4.87 per 1,000. It billed a mean of 978 output tokens for answers averaging 33 words. Reasoning off billed 66.
- After the sale those figures become $5.93 and $9.75. At default settings that is 4.3 times Large 3, and more than Mistral Medium 3.5 ($6.79) costs on the same request.
- Default reasoning took 11.7 seconds per reply on average and 23.8 seconds at worst. Reasoning off took 2.1 seconds, close to Large 3’s 1.8.
- The answers were no better with reasoning on. All 29 replies were correct against their source pages. The reasoning-off replies were slightly longer and more often included a link.
- Large 4 answered all 22 of our requests on the first attempt. Large 3 returned HTTP 429 (rate limited) on 16 of 21 attempts through the same route, as it did on 5 October.
Mistral Large 4 pricing during and after the preview
These are the per-million-token rates as listed on 10 October 2026. Large 4’s sale prices are what Mistral bills during the preview. OpenRouter lists the same Large 4 endpoint with a 50% discount flag, served by Mistral itself. Large 3 and Medium 3.5 are the prices we read for our 5 October test, and neither has changed since.
| Per 1M tokens | Input | Cached input | Output | Context window |
|---|---|---|---|---|
| Mistral Large 4, preview sale (to about 20 Oct) | $0.68 | $0.07 | $2.09 | 1,048,576 |
| Mistral Large 4, list price | $1.36 | $0.14 | $4.18 | 1,048,576 |
Mistral Large 3 (mistral-large-2512) | $0.50 | $0.05 | $1.50 | 262,144 |
| Mistral Medium 3.5 | $1.50 | not discounted on OpenRouter in our tests | $7.50 | 262,144 |
Even at the sale price, Large 4 costs more per token than the model it replaces: 36% more for input and 39% more for output than Large 3. At list price it is 2.7 times Large 3 on input and 2.8 times on output. Mistral’s pricing page still answers its FAQ with “Mistral Large costs $0.5 /M tokens in and $1.5 /M tokens out”, which is Large 3’s rate. If you have budgeted from that line, it no longer describes the newest Large model.
Mistral describes Large 4 as an open-weight, general-purpose multimodal model with a mixture-of-experts architecture. Its changelog says the open weights are “coming soon”. For now the API is the only way to run it. Once the weights are out, other hosts will be able to sell it at their own prices. Our GLM-5.3 host comparison found more than a tenfold spread in cost per reply between hosts of one open-weight model, so the list price may not stay the only price for long.
How we tested
All calls went from our WordPress server to OpenRouter’s chat completions endpoint, pinned to Mistral as the provider with fallbacks disabled, on 10 October 2026. That is the route an MxChat site uses when it picks a Mistral model from the OpenRouter list. Each request carried our live chatbot system prompt (3,478 characters) and one of three real pre-sales questions:
- Does MxChat work with WooCommerce, and can the bot look up an order status for a logged-in customer?
- Do I need Pinecone to use the knowledge base, or can I store embeddings in WordPress?
- Can a visitor be handed over to a human on Slack or Telegram when the bot cannot answer?
For each question we pulled four passages from our own documentation and blog posts, about 16,100 characters, and sent them as the knowledge base with the question. These are the same system prompt, passages and questions as our 5 October Mistral test, so the figures can be compared directly. Large 4 counted the request as 4,160 input tokens on average and Large 3 as 4,326, so Large 4’s tokenizer counted about 3% fewer tokens for the same text, question by question.
We ran five configurations, each question twice:
- Large 4, default: no reasoning parameter sent, which is what MxChat sends today.
- Large 4, effort high:
reasoning: {effort: "high"}. - Large 4, reasoning off:
reasoning: {effort: "none"}. Sendingenabled: falsehad the same effect in our probes. - Large 3 and Medium 3.5, defaults, as baselines.
Costs come from the usage block OpenRouter returns with each response, and we checked every figure against token counts multiplied by the listed rates. One Large 3 call failed after six attempts because of rate limiting, so Large 3’s figures use five replies; every other configuration has six. Our prompt-cache hits were irregular. Some Large 4 calls read 3,000 to 4,000 tokens from cache and others read none, on identical requests. To compare like with like, the headline figures below price every request as if nothing came from cache, and the billed figures with cache hits are listed separately.
Cost per 1,000 replies
| Per 1,000 replies | Sale price, no cache | List price, no cache | Billed in our test (with cache hits, sale price) | Output tokens per reply (mean) | Words in answer (mean) |
|---|---|---|---|---|---|
| Large 4, default | $4.87 | $9.75 | $4.05 | 978 | 33 |
| Large 4, effort high | $4.76 | $9.52 | $2.59 | 924 | 32 |
| Large 4, reasoning off | $2.97 | $5.93 | $2.87 | 66 | 44 |
| Large 3 | $2.26 | $2.26 | $1.98 | 65 | 40 |
| Medium 3.5 | $6.79 | $6.79 | $6.79 | 42 | 24 |
Three things stand out.
Reasoning is the expensive part. With reasoning off, Large 4 billed 66 output tokens per reply, about the same as Large 3. With default settings it billed 978, ranging from 508 to 1,995, for a visible answer of about 33 words. Nearly all of that output was reasoning the visitor never sees. It is billed at the output rate, so output went from 5% of the bill with reasoning off to 40% at default settings.
“Effort high” was not more expensive than the default. The default already reasons at length on this kind of request. The two settings billed similar output (924 against 978 tokens), and the gap between them in the billed column comes from cache hits, not from the setting. Every effort-high call happened to hit the cache and only half of the default calls did.
After the sale, the default is the most expensive Mistral option we tested. At list price, Large 4 with default reasoning costs $9.75 per 1,000 replies, ahead of Medium 3.5’s $6.79. Medium 3.5’s per-token price is higher, but it answered in 42 tokens without reasoning.
Speed
Reasoning costs time as well as money. Large 4’s default replies took 11.7 seconds on average, and the slowest took 23.8 seconds. Reasoning off brought that down to 2.1 seconds on average and 2.9 at worst, close to Large 3’s 1.8 and Medium 3.5’s 1.9. These are full wall-clock times from our server in the US, including network.
For a chat widget this matters more than the raw figure suggests. MxChat streams replies, but a reasoning model sends nothing visible until it has finished reasoning, so the visitor stares at the typing indicator for the whole reasoning phase. A ten-second pause before the first word is long enough for some visitors to give up.
Are the answers better with reasoning on?
Not on these questions. We read all 29 replies against the source passages. Every configuration answered all three questions correctly: WooCommerce order lookups come with the Pro add-on and need a logged-in customer; Pinecone is optional and embeddings can stay in the WordPress database; Slack and Telegram handoff are in the free core plugin under Settings, Integrations.
The differences were in style. The reasoning-off replies were a little longer (44 words against 33) and were the only Large 4 replies that included links, such as the WooCommerce add-on page. Default and effort-high replies were tidy summaries without links. Details such as the 50,000-entry guidance for Pinecone and the WooCommerce version requirement all came from the source pages.
Pre-sales questions with the answer in the retrieved passages are the easy case for a site chatbot, and also the most common one. Reasoning may earn its cost on multi-step troubleshooting or on questions the knowledge base only half answers. Retrieval questions don’t need it. Our Gemini 3.8 Flash test found the same pattern, with most of the billed output as hidden thinking and no visible difference in the replies.
Rate limits: Large 4 answered first time, Large 3 did not
Every Large 4 request we sent through OpenRouter’s shared Mistral capacity succeeded on the first attempt: 18 test calls and 4 probes. Large 3, on the same route a few minutes later, returned HTTP 429 with “temporarily rate-limited upstream” on 16 of 21 attempts. Our retry loop waited 12 seconds between tries, and one Large 3 request was still failing after six. We saw the same on 5 October, when Large 3 and Small 4 failed 15 of 21 calls.
That is a snapshot of shared capacity on one morning, not a guarantee. OpenRouter’s error message suggests adding your own Mistral key to get your own rate limits, and a direct Mistral API key avoids the shared pool entirely. But if a site runs Large 3 through OpenRouter’s shared key today, visitors will sometimes get no answer. Large 4 is newer and, for now, much less congested.
What Mistral Large 4 costs a real site
The table below scales our per-reply costs to three traffic levels, with no cache hits. A reply here is one bot answer to one visitor message.
| Replies per month | Large 4 off, sale | Large 4 off, list | Large 4 default, sale | Large 4 default, list | Large 3 |
|---|---|---|---|---|---|
| 5,000 | $14.85 | $29.65 | $24.35 | $48.75 | $11.30 |
| 20,000 | $59.40 | $118.60 | $97.40 | $195.00 | $45.20 |
| 100,000 | $297 | $593 | $487 | $975 | $226 |
Cache hits reduce these figures. Large 4’s cached input is 10% of its input price, so when a request did hit the cache its input cost fell sharply. Our billed figures were 3% to 46% below the no-cache figures, depending on how many calls happened to hit. Because the hit rate was irregular on identical back-to-back requests, budget from the no-cache column.
What this means for an MxChat site
MxChat reaches Mistral Large 4 through OpenRouter: enter an OpenRouter key, load the model list and choose mistralai/mistral-large-4-0. You can also call Mistral’s own API through the custom OpenAI-compatible provider. On the OpenRouter route, MxChat currently sends the model, the messages and a temperature, with no reasoning setting. A site that picks Large 4 today therefore gets the default reasoning behaviour, with the cost and delay measured above.
If you are evaluating Large 4 during the sale, price it at list. Two weeks of half price is a good time to test answer quality, but the rate you will pay from about 20 October is $1.36 and $4.18.
If you want Large 4 for a support bot, turn reasoning off where your setup allows it. That cut the cost by 39% and the wait from 11.7 seconds to 2.1 in our test, with answers that were just as accurate. Through a custom provider or your own code, that is reasoning: {effort: "none"} on OpenRouter.
If Large 3 works for you, it is still the cheaper model: $2.26 per 1,000 replies against $5.93 for Large 4 with reasoning off at list. The case for moving is Large 4’s 1M-token context window and, for now, fewer 429s on shared capacity. Our earlier Mistral API pricing comparison covers Small 4 and Ministral 14B, which are cheaper again.
If you are comparing across vendors, Large 4 with reasoning off at list ($5.93) lands between the budget tier and the mid-tier flagships. Claude Haiku 5.5 with thinking off cost $0.63 per 1,000 on our 8 October request in our Haiku 5.5 pricing test, and GPT-6 Luna $0.53. Those figures come from a different request with a different system prompt layout, so treat the comparison as an order of magnitude, not a precise ratio.
MxChat connects to Mistral, OpenAI, Claude, Gemini and other providers with your own API key, and you pay the provider directly with no markup on tokens. The MxChat documentation covers model selection and the OpenRouter setup. MxChat Pro adds the WooCommerce, live-agent and knowledge-base features the test questions were about.
FAQ
How much does Mistral Large 4 cost?
During the public preview, Mistral bills $0.68 per million input tokens, $0.07 per million cached input tokens and $2.09 per million output tokens. That is half the list price of $1.36, $0.14 and $4.18. Mistral’s changelog describes the preview pricing as 50% off for two weeks from the 6 October 2026 launch.
When does the Mistral Large 4 discount end?
Mistral’s changelog says “Launch pricing: 50% off for 2 weeks” in its 6 October 2026 entry. Counting from launch, that is around 20 October 2026. Mistral had not published an exact end date when we checked on 10 October, so check its pricing before you rely on the sale rate.
Is Mistral Large 4 more expensive than Mistral Large 3?
Yes, even at the sale price: input costs 36% more and output 39% more than Large 3’s $0.50 and $1.50. At list price it is about 2.7 to 2.8 times Large 3 per token. On our chatbot request Large 4 with reasoning off cost $2.97 per 1,000 replies at the sale price and $5.93 at list, against $2.26 for Large 3.
Does Mistral Large 4 use reasoning by default?
Yes. With no reasoning parameter it reasoned on every request we sent, including a two-word “Say OK” probe. On our support questions that meant a mean of 978 billed output tokens and 11.7 seconds per reply. Sending reasoning: {effort: "none"} through OpenRouter switched it off, with 66 output tokens and 2.1 seconds.
Is Mistral Large 4 good for a WordPress chatbot?
It answered all three of our pre-sales questions correctly in every configuration. With reasoning off it is fast enough for a chat widget and costs about 31% more than Large 3 during the sale. At default settings it is slow for live chat, and from about 20 October it is the most expensive Mistral option we have measured on our request.
Is Mistral Large 4 open source?
Mistral calls it an open-weight model, but the weights were not yet downloadable on 10 October 2026. Mistral’s changelog says the open weights are coming soon. Until then the API is the only way to run it.