Gemini 3.7 Flash Is Half Price Until January 1
On 13 August Google shipped Gemini 3.7 Flash and cut the price in half. That is the headline everyone ran. The part almost nobody ran is printed on the same page, in the same table, one line down: the cut expires on 31 December 2026, and on 1 January the price doubles.
Both numbers are published today. The second one is not a prediction or a rumour — it is Google’s own listed price for the same model, with a date attached. If you are picking a model for a WordPress chatbot this month on the strength of that $0.75, you are signing up for $1.50 in nineteen weeks.
This is not a Google problem. It is now the normal shape of AI API pricing, and it has quietly become the most under-reported thing about running an AI feature on your site. I fetched all three major vendors’ own pricing pages this morning and wrote down what expires when. The answers are inconsistent enough to be worth a table.
What Google actually shipped on 13 August
Gemini 3.7 Flash landed as Google’s “workhorse” tier — the model aimed at high-volume, everyday jobs rather than hard reasoning. The benchmark story was coding and agentic workflows. The pricing story is simpler:
| Gemini 3.7 Flash | Through 31 Dec 2026 | From 1 Jan 2027 | Change |
|---|---|---|---|
| Input (per 1M tokens) | $0.75 | $1.50 | 2.0× |
| Output (per 1M tokens) | $3.75 | $7.50 | 2.0× |
One detail deserves more attention than it got: the discount was applied retroactively to Gemini 3.6 Flash, which had launched only weeks earlier at the higher price. So the older model got cheaper on the day the newer one arrived, and both now sit on the same 31 December cliff. If you deployed 3.6 Flash in late July at $1.50, you are currently paying $0.75 without having changed anything — and you will be back at $1.50 in January without having changed anything either.

Three vendors, three completely different expiry regimes
Here is where it stops being a Google story. I read Anthropic’s, Google’s and OpenAI’s own pricing pages on 23 August 2026. Every figure below came off the vendor’s page, not off a comparison blog — for reasons that will become obvious in a moment.
| Vendor | Model | Promotional price | Expiry | What happens next |
|---|---|---|---|---|
| Gemini 3.7 / 3.6 Flash | $0.75 / $3.75 | 31 Dec 2026 | Doubles to $1.50 / $7.50, date published | |
| OpenAI | GPT-5.6 Sol | $4 / $20 | “at least through 21 Nov 2026” | No successor price published |
| Anthropic | Claude Sonnet 5 | $2 / $10 | Cancelled | Introductory price became the standard price |

Three vendors, three regimes, and only one of them gives you a number you can actually budget against:
- Google publishes the cliff. You know the date and you know the new price. This is annoying but honest, and it is the only one of the three you can put in a spreadsheet.
- OpenAI publishes a floor, not a ceiling. “At least through 21 November 2026” tells you the price will not rise before that date. It does not tell you what happens on the 22nd. That is a real planning gap, and the hedged phrasing is doing a lot of work.
- Anthropic cancelled its own increase. Claude Sonnet 5 launched at $2/$10 described as introductory pricing through 31 August 2026, with a scheduled rise to $3/$15 on 1 September. That rise is off. The docs now say, verbatim: “The previously scheduled increase to $3/$15 per million input/output tokens on September 1, 2026 will not occur.”
We wrote about that Anthropic reversal when it happened — see Claude Sonnet 5 pricing: the September increase was cancelled, which we had to retitle after the vendor moved the goalposts in the helpful direction for once.
The trackers are already wrong, and you can watch it happen
Search gemini api pricing today and the first result is Google’s own documentation. Good. Below the vendor pages sit a stack of independent pricing trackers and “every model, every tier” explainers, several of which rank well and update often.
At least one of them, as its result appeared in this morning’s search listing, quotes Gemini 3.6 Flash at $1.50 / $7.50 — the January price — as though it were the current one. It is not current. Google’s own page says 3.6 Flash is $0.75 / $3.75 through 31 December, because the retroactive cut caught it too. Another currently-ranking result carries “May 2026” in its title and a January date stamp on the page.
This is not a swipe at anyone in particular. It is a structural problem, and it is the fourth time in three weeks that we have caught a secondary source carrying a stale vendor price:
- Comparison articles quoting Intercom’s Expert seat at $139 when intercom.com said $132.
- Multiple AI-pricing trackers still describing the Sonnet 5 September increase as scheduled, more than a week after Anthropic said in writing that it would not occur — including one dated 19 August.
- Trackers quoting Gemini 3.6 Flash at its pre-cut price after the cut was applied retroactively.
The rule we now run on, and the one worth stealing: for any price you are going to make a decision on, open the vendor’s own page. Model pricing changed at least four times across three vendors in the last month. No third-party tracker updates faster than that reliably, and the SERP does not sort by freshness.
What this actually costs on a WordPress site
Per-million-token prices are useless until you know how many tokens a conversation on your site actually burns. So here are ours, measured rather than estimated.
Our working figure is 7,201 input + 716 output tokens per conversation, measured from this site’s own chat transcript table. The lopsidedness is the important part: input outweighs output roughly ten to one, because every turn re-sends the system prompt and the retrieved context. That ratio means the input price is what governs your bill, and the output price — the one vendors like to lead with — barely matters.
At that token shape, one conversation costs:
| Model | Price per 1M (in / out) | Per conversation | Per 1,000 conversations |
|---|---|---|---|
| Gemini 2.5 Flash-Lite | $0.10 / $0.40 | $0.0010 | $1.01 |
| GPT-5.6 Luna | $0.20 / $1.20 | $0.0023 | $2.30 |
| Gemini 3.1 Flash-Lite | $0.25 / $1.50 | $0.0029 | $2.87 |
| Gemini 3.7 Flash (through 31 Dec) | $0.75 / $3.75 | $0.0081 | $8.09 |
| Claude Haiku 4.5 | $1 / $5 | $0.0108 | $10.78 |
| Gemini 3.7 Flash (from 1 Jan) | $1.50 / $7.50 | $0.0162 | $16.17 |
| Claude Sonnet 5 | $2 / $10 | $0.0216 | $21.56 |
| GPT-5.6 Terra | $2 / $12 | $0.0230 | $22.99 |
| GPT-5.6 Sol | $4 / $20 | $0.0431 | $43.12 |

The honest read on those numbers: at a thousand conversations a month, the entire spread between the cheapest and dearest option is about forty dollars. For most WordPress sites this is not the line item worth agonising over, and anyone selling you a model migration on cost alone is selling you forty dollars.
The January change is worth noticing anyway, for a reason that has nothing to do with the absolute size. Gemini 3.7 Flash crosses Claude Haiku 4.5 on 1 January. Today it is the cheaper of the two ($0.0081 vs $0.0108). In January it is the dearer ($0.0162 vs $0.0108) — and Haiku’s price has no expiry attached to it. If you choose Flash today on price and never revisit, you end up on the wrong side of a comparison you thought you had already made. That is the actual cost of an undated decision.
The line item nobody budgets for: retrieved context
Here is the measurement that surprised us most, and it is the one that generalises.
On this install, the retrieval context attached to answers averages 11,516 characters — roughly 2,879 tokens — per retrieval. The entire human-readable conversation, every message from both sides across every session in the table, averages about 2,826 characters per session.
The context we retrieve and send is roughly four times the size of the conversation itself. Nobody is typing 11,000 characters at your chatbot. Your retrieval layer is, on every single turn, and it is billed at the input rate.
The practical consequence is that how much you retrieve moves your bill considerably harder than which model you picked. Halving your retrieved context is a ~40% cut to a RAG-heavy chatbot’s bill and it works on every vendor at once, with no migration, no re-testing, and no exposure to anyone’s expiry date. Switching from Gemini 3.7 Flash to Flash-Lite saves less than tightening what you send it.
A caveat stated plainly, because the sample deserves it: the transcript table on this install currently holds 31 messages across 5 sessions, with 7 rows carrying retrieval context. That is a small sample and it is pruned on a schedule, so treat the ratio as an order-of-magnitude finding rather than a precise constant. The direction has been stable across every measurement we have taken; the exact multiple has not. We walked through the full token accounting in our measurement of Intercom Fin’s per-resolution pricing against raw token costs, and the per-token method in the self-hosted chatbot cost breakdown.
What to actually do before 1 January
Four things, in order of how much they are worth.
- Find out what model your chatbot is actually calling. Not what you chose at setup — what is configured right now. On this site it is
gpt-5.6-luna, which we can confirm because we went and looked rather than remembering. Plugins get updated, defaults get remapped, and model identifiers get retired out from under you. - Measure your retrieved context before you shop for a cheaper model. If your RAG payload is four times your conversation, that is the lever. Model choice is the smaller number.
- Put 1 January 2027 in the calendar if you are on Gemini 3.7 or 3.6 Flash. Not to panic — to re-run the comparison with the new number, because the ranking changes on that date and the crossover with Haiku 4.5 is real.
- Never budget off a tracker. Open the vendor’s page. It takes thirty seconds and, on current evidence, roughly a third of the ranking pricing content is wrong at any given moment.
And a fifth, which is really the whole point: write the date next to the price. A price without an expiry attached is not a price, it is a quote with the terms missing. Two of the three major vendors are currently shipping exactly that.
Does any of this change which chatbot plugin you should use?
Honestly, no — and it would be convenient for us to pretend otherwise.
What it does change is how much the bring-your-own-key architecture is worth. If your chatbot plugin holds your API key and calls the vendor directly, every one of these price movements reaches you immediately: the Gemini cut on 13 August landed in your bill that day, and so will the January increase. You can switch models in an afternoon when the numbers move. If instead you are on a hosted per-conversation or per-resolution plan, none of this touches you — you pay the vendor’s retail rate regardless of what the underlying token price does, which is fine when prices rise and considerably less fine when they halve.
MxChat is the first kind: you supply the key, you pay the model provider directly at whatever today’s price happens to be, and you change models from a dropdown when a table like the one above shifts. That is a real advantage in a month like this one and a mild inconvenience in a quiet month. You can see the supported providers in the documentation, and the plugin tiers on the MxChat Pro page.
FAQ
Is the Gemini 3.7 Flash price increase definitely happening?
Google has published $1.50 / $7.50 as the price “starting January 1, 2027” on its own pricing page. Publishing a future price is not the same as committing to it — Anthropic published a September increase and then cancelled it. Plan for the published number and be pleasantly surprised if it moves.
Does the increase affect Gemini 3.6 Flash too?
Yes. The August cut was applied retroactively to 3.6 Flash, and both models are listed at $0.75 / $3.75 through 31 December and $1.50 / $7.50 afterwards.
Which model is cheapest for a WordPress chatbot right now?
On our measured token shape, Gemini 2.5 Flash-Lite at $0.0010 per conversation, then GPT-5.6 Luna at $0.0023. But the spread across the whole table is around $40 per thousand conversations, so answer-quality and latency should probably decide this, not price.
Did Claude Sonnet 5 get more expensive on 1 September 2026?
No. The scheduled increase to $3/$15 was cancelled and $2/$10 is now the standard price.
What is the single biggest lever on chatbot API cost?
For a retrieval-backed chatbot, how much context you retrieve per turn. On this install it outweighs the entire conversation by roughly four to one, and it is billed at the input rate on every turn.
The short version
Gemini 3.7 Flash is genuinely cheap right now and genuinely temporary. Google has told you the date and the new number, which is more than OpenAI has done for GPT-5.6 Sol. Anthropic cancelled its increase outright and turned an introductory price into a permanent one.
The money, though, is not in the model column. It is in how much you send. We measured our retrieval payload at four times the size of the conversation it supports, and no amount of model-shopping beats sending less.
Every price in this article was read from the vendor’s own pricing page on 23 August 2026. Google’s page is dated 13 August 2026. If you are reading this after 1 January 2027, the Gemini figures above have changed — go and check, which is rather the point.