Gemini Flash-Lite Pricing: 2.5 Closed, 3.5 Costs 3x More
Gemini 2.5 Flash-Lite is still the cheapest model on Google’s Gemini API price list, at $0.10 per million input tokens and $0.40 per million output tokens. You probably cannot use it. When we called it on 3 October 2026 with an API key that had never used it before, the API returned HTTP 404 with this message: “This model models/gemini-2.5-flash-lite is no longer available to new users.” The same message points you to Gemini 3.5 Flash-Lite, which costs $0.30 input and $2.50 output, three times as much on input and more than six times as much on output.
Gemini Flash-Lite pricing is now a three-model question: 2.5 Flash-Lite for accounts that already use it, 3.1 Flash-Lite at $0.25 and $1.50, and 3.5 Flash-Lite at $0.30 and $2.50. We sent all of them the same WordPress chatbot requests from our production server, 45 calls in all, and priced what the API reported.
- Gemini 2.5 Flash-Lite is closed to new API keys. Five calls and a token count all returned 404 on our key. Google’s deprecations page, read the same day, lists no shutdown date for it, only that access is limited to users who have used it before.
- Gemini 3.5 Flash-Lite cost $1.35 per 1,000 chatbot replies. The same tokens on 2.5 Flash-Lite’s rate card come to $0.43, so the move Google suggests roughly triples the bill.
- Gemini 3.1 Flash-Lite was 17% cheaper than 3.5 at $1.12 per 1,000 replies, with the same speed (0.66 seconds) and answers we could not fault. It has a published shutdown date of 7 May 2027.
- Input was 93% of every Flash-Lite bill. The input rate, not the output rate, decides what a retrieval chatbot pays.
- Turning thinking on broke 3.5 Flash-Lite. At
thinkingLevel: highit cost 2.5 times as much and 3 of 6 replies came back cut off or empty.
Gemini Flash-Lite pricing, model by model
These are the standard paid-tier rates per million tokens as Google’s Gemini API pricing page listed them on 3 October 2026, for text input. Output prices include thinking tokens.
| Per 1M tokens | Gemini 2.5 Flash-Lite | Gemini 3.1 Flash-Lite | Gemini 3.5 Flash-Lite | Gemini 3.8 Flash |
|---|---|---|---|---|
| Input (text, image, video) | $0.10 | $0.25 | $0.30 | $0.75 ($1.50 from 1 Jan 2027) |
| Output (includes thinking) | $0.40 | $1.50 | $2.50 | $3.75 ($7.50 from 1 Jan 2027) |
| Context caching | $0.01 | $0.025 | $0.03 | $0.075 |
| Batch input / output | $0.05 / $0.20 | $0.125 / $0.75 | $0.15 / $1.25 | $0.375 / $1.875 |
| Audio input | $0.30 | $0.50 | $0.30 | see Google’s page |
| Callable by a new API key on 3 Oct 2026 | No (HTTP 404) | Yes | Yes | Yes |
| Shutdown date on Google’s deprecations page | None listed | 7 May 2027 | None listed | None listed |
Two things stand out. First, each newer Flash-Lite costs more than the last, so “upgrade to the latest Flash-Lite” is a price increase every time. Second, 3.5 Flash-Lite and the older 2.5 Flash (not Lite) now sit on the same card: $0.30 input and $2.50 output. The Lite tier has moved up to where the mid tier used to be.
Gemini 3.1 Flash-Lite also has Flex and Priority tiers. Flex is half price and Priority costs 1.8 times Standard, the same pattern we described for 3.8 Flash in our Gemini 3.8 Flash pricing write-up. Like Batch, Flex is for work that can wait, so it does not help a visitor waiting for a reply.
What “no longer available to new users” means in practice
Google has not published a retirement date for gemini-2.5-flash-lite. Its deprecations page says access to the 2.5 models is being limited to users who have actively used them in the past. Some third-party price trackers give 16 October 2026 as a shutdown date; we could not find that date on any Google page, so treat it as unconfirmed.
What we could check is what the API does. Our key belongs to a WordPress site that runs Gemini models every day but had not called 2.5 Flash-Lite. Every request to it failed:
| Request to gemini-2.5-flash-lite | Result |
|---|---|
| generateContent, no thinking setting | 404 NOT_FOUND, “no longer available to new users” |
generateContent, thinkingBudget: 0 | 404, same message |
generateContent, thinkingLevel: low and high | 404, same message |
| countTokens | 404, same message |
Model list (models.list) | Still listed |
That last row is the trap. The model still appears in the model list your plugin or code may fetch to build a dropdown, so it looks available until the first real request fails. If you set up a new site, a new Google Cloud project or a new API key today and pick 2.5 Flash-Lite because it is the cheapest name on the list, your chatbot will return errors from the first message.
If your site already calls 2.5 Flash-Lite with an established key, it should keep working for now. That is the only group still paying $0.10 and $0.40. Everyone else is choosing between 3.1 and 3.5.
We also called gemini-flash-lite-latest, Google’s alias for the newest Flash-Lite. On 3 October it was served by gemini-3.5-flash-lite. Code that uses the alias moved to the $0.30 / $2.50 rate card without anyone changing a setting.
How we tested
Every call went from the mxchat.ai WordPress server to Google’s generateContent endpoint with the site’s own Gemini key, which never left the server. Each request was built like a knowledge-base chatbot request: the live system prompt for our own site bot as the system instruction, then a user message carrying four retrieved documentation excerpts (about 16,000 characters) and the visitor’s question. The three questions came from our own chat logs:
- Does MxChat work with WooCommerce, and can the bot look up an order status for a logged-in customer?
- Do I need Pinecone to use the knowledge base, or can I store embeddings in WordPress?
- Can a visitor be handed over to a human on Slack or Telegram when the bot cannot answer?
These are the same three questions we used for our GPT-6 Luna test and our DeepSeek V4.1 Flash test, so the results line up with those.
We ran 3.1 Flash-Lite and 3.5 Flash-Lite with no thinking setting, which is how Google serves them by default, and 3.5 Flash-Lite again at thinkingLevel: high. As a reference we ran Gemini 3.8 Flash at its default and at thinkingLevel: low. Each configuration answered all three questions twice. We then ran three two-turn conversations on each Flash-Lite model, with the follow-up “Is that included in the free plugin, or do I need MxChat Pro for it?”. maxOutputTokens was 1,000 on every call, with no streaming.
Costs come from the token counts in each response’s usageMetadata, priced at the list rates above. Google’s countTokens returned 3,256 tokens for the first question’s user message on 3.1 Flash-Lite, 3.5 Flash-Lite and 3.8 Flash alike, and every model billed the full requests at 4,019, 4,105 and 4,371 input tokens. The tokenizer is the same across these models, so price differences come from the rate card and the output length.
Cost per 1,000 chatbot replies
| Model and thinking setting | Per 1,000 replies | Input tokens | Output tokens | Thinking tokens | Input share of bill |
|---|---|---|---|---|---|
| Gemini 2.5 Flash-Lite (rate card only, not callable) | $0.43 | 4,165 | 44 | 0 | 96% |
| Gemini 3.1 Flash-Lite, default | $1.12 | 4,165 | 50 | 0 | 93% |
| Gemini 3.5 Flash-Lite, default | $1.35 | 4,165 | 40 | 0 | 93% |
| Gemini 3.5 Flash-Lite, thinking high | $3.31 | 4,165 | 35 | 788 | 38% |
| Gemini 3.8 Flash, thinking low | $3.39 | 4,165 | 71 | 0 | 92% |
| Gemini 3.8 Flash, default | $5.08 | 4,165 | 66 | 455 | 62% |
Token figures are means per reply. The 2.5 Flash-Lite row prices the mean token counts of the 15 default Flash-Lite replies on 2.5 Flash-Lite’s rate card, since the API would not answer us. It is what an established account would pay for the same work, not a measured call.
Here is the arithmetic behind the 3.5 versus 3.1 gap. A reply carried 4,165 input tokens. At $0.30 per million that is $1.25 per 1,000 replies on 3.5 Flash-Lite, against $1.04 at $0.25 on 3.1 Flash-Lite. Output added $0.10 and $0.08. So of the $0.23 difference, about $0.21 is the input rate. 3.5 Flash-Lite’s much higher output price barely registers, because it wrote only about 40 tokens per reply.
That matches every retrieval chatbot we have priced this year. The knowledge-base excerpts are resent with every question, so a reply reads about 100 tokens for each one it writes. If you compare cheap models for a chatbot, compare input prices first.
Next to other low-cost models
On the same three questions, GPT-6 Luna came to $0.57 per 1,000 replies in our test on 28 September, and DeepSeek V4.1 Flash to $0.77 off-peak and $1.53 at peak on 27 September. Each vendor counts tokens differently, so treat those as rough neighbours rather than an exact ranking. Even so, Gemini 3.1 Flash-Lite at $1.12 and 3.5 Flash-Lite at $1.35 are no longer the cheapest option for a retrieval chatbot. 2.5 Flash-Lite was the cheapest of the three vendors, which is why losing access to it matters.
What it costs per month
| Replies per month | 2.5 Flash-Lite (existing users) | 3.1 Flash-Lite | 3.5 Flash-Lite | 3.8 Flash, thinking low |
|---|---|---|---|---|
| 1,000 | $0.43 | $1.12 | $1.35 | $3.39 |
| 10,000 | $4.34 | $11.16 | $13.50 | $33.90 |
| 100,000 | $43.40 | $111.60 | $135.00 | $339.00 |
| 100,000, from 1 Jan 2027 | $43.40 | $111.60 | $135.00 | $678.00 |
For most WordPress sites these are small numbers. A shop answering 10,000 questions a month pays about $9 more on 3.5 Flash-Lite than it would have on 2.5. The ratio matters more at volume, and when you set a monthly spend limit: a cap sized for 2.5 Flash-Lite will run out about three times as fast on its replacement.
The last row is a reminder that 3.8 Flash’s introductory price ends on 31 December 2026, when its input and output rates double. Flash-Lite has no announced change. We covered what the earlier Flash price expiry did to bills in our Gemini 3.7 Flash price expiry post.
Follow-up turns cost as much as first turns
Google lists a context caching price for every Flash-Lite model, and Gemini 2.5 and later models also cache repeated prompt openings automatically (“implicit caching”), billing the cached part at the lower rate. None of our 45 calls reported any cached tokens. That includes six follow-up turns whose first 4,000 or so tokens repeated, word for word, a request we had sent two seconds earlier.
| Turn (mean of 3 conversations) | 3.1 Flash-Lite input / cached | 3.1 Flash-Lite per 1,000 | 3.5 Flash-Lite input / cached | 3.5 Flash-Lite per 1,000 |
|---|---|---|---|---|
| First question | 4,165 / 0 | $1.15 | 4,165 / 0 | $1.35 |
| Follow-up | 4,256 / 0 | $1.15 | 4,227 / 0 | $1.37 |
So on these models, at this traffic, a follow-up costs as much as a new question. Implicit caching is applied at Google’s discretion and is not guaranteed, so do not count on it in a budget. If you need a cache discount, Google’s explicit context caching API creates a cache you control, at $0.03 per million cached tokens for 3.5 Flash-Lite plus $1.00 per million tokens per hour of storage. For a 1,200-token system prompt that storage costs about $0.03 a day, which only pays off with steady traffic. We measured how caching changes chatbot bills on other vendors in our OpenAI prompt caching test.
Speed
Both Flash-Lite models answered in about two-thirds of a second, full round trip from our server with no streaming: 0.66 seconds on 3.1 (0.52 to 0.89) and 0.67 seconds on 3.5 (0.56 to 0.74). 3.8 Flash at thinkingLevel: low took 1.35 seconds, and at its default, where it thinks before answering, 3.36 seconds, with one call at 6.9 seconds.
For a chat widget, anything under a second feels instant. Both Flash-Lite models are there, and the difference between them is noise.
Answer quality
We read all 45 replies. Every default Flash-Lite reply was based on the retrieved excerpts and got the facts right: WooCommerce order lookup comes from the WooCommerce add-on, embeddings can be stored in WordPress or in Pinecone, and live-agent handoff works through Slack and Telegram.
The differences were in how directly they answered:
- Flash-Lite replies were short. 3.5 Flash-Lite averaged 27 words and 3.1 Flash-Lite 37, against about 49 for 3.8 Flash.
- Flash-Lite did not give a straight “no”. Asked “Do I need Pinecone?”, 0 of 7 default Flash-Lite replies said no outright. They said you can store embeddings in WordPress or use Pinecone, which is accurate but leaves the visitor to work out the answer. All 4 replies from 3.8 Flash opened with “You do not need Pinecone.”
- The follow-up was handled well. Both Flash-Lite models correctly said the WooCommerce add-on needs a Pro licence and that the knowledge base and live-agent handoff are in the free plugin.
- Both followed the system prompt’s sales instructions, sometimes too eagerly, adding an upgrade line to answers about free features. That is a prompt-tuning issue, not a model one, and 3.8 Flash did it too.
For a knowledge-base support bot answering simple questions, either Flash-Lite model is good enough. If your visitors ask yes-or-no questions and you want the answer in the first three words, 3.8 Flash at thinkingLevel: low was noticeably better at that, for about 2.5 times the price.
Do not turn on thinking for 3.5 Flash-Lite
3.5 Flash-Lite does not think unless you ask it to. When we set thinkingLevel: high, it spent 491 to 961 thinking tokens on each reply. Thinking is billed as output at $2.50 per million, so the cost per reply rose from $1.35 to $3.31 per 1,000, about the same as 3.8 Flash at low thinking. Response time went from 0.67 to 2.99 seconds.
Worse, the thinking used up the output allowance. With maxOutputTokens at 1,000, two of the six replies stopped with MAX_TOKENS partway through a sentence, and one returned MALFORMED_RESPONSE with no answer at all, after 906 thinking tokens. A visitor would have seen half an answer or an error. The three replies that did complete were no better than the default ones.
The parameters each model accepted, from our probes on 3 October:
| Setting | 3.1 Flash-Lite | 3.5 Flash-Lite | 3.8 Flash |
|---|---|---|---|
| No thinking setting | OK, no thinking | OK, no thinking | OK, thinks |
thinkingBudget: 0 | OK | 400 error | OK, no thinking |
thinkingLevel: minimal | OK, no thinking | OK, no thinking | 400 error |
thinkingLevel: low | OK, thinks | OK | OK, no thinking on our requests |
thinkingLevel: high | OK, thinks | OK, thinks | OK, thinks |
The safest setting for both Flash-Lite models is no thinking setting at all. That is also the only setting that works on every model in the table. In particular, code written for 2.5-era models that sends thinkingBudget: 0 to switch thinking off will get a 400 error from 3.5 Flash-Lite.
Which Flash-Lite should a WordPress chatbot use?
| Your situation | Our suggestion |
|---|---|
| Your site already runs 2.5 Flash-Lite with an existing key | Keep it while it works, but test 3.1 Flash-Lite now so you are not switching under pressure when access ends. |
| New site or new API key, cost comes first | Gemini 3.1 Flash-Lite. 17% cheaper than 3.5 on our requests, just as fast, with a published shutdown date of 7 May 2027 so you can plan around it. |
| You want the newest Flash-Lite and its longer support window | Gemini 3.5 Flash-Lite, with no thinking setting. Budget about 3 times what 2.5 Flash-Lite cost. |
| Visitors ask yes-or-no product questions and you want direct answers | Gemini 3.8 Flash at thinkingLevel: low. Budget for its price doubling on 1 January 2027. |
Your code uses gemini-flash-lite-latest | It already serves 3.5 Flash-Lite. Pin a specific model if you want to control when your price changes. |
In MxChat, the model is a setting under the chatbot’s AI provider, and you pay Google directly at these list rates with your own key. MxChat sends no thinking setting to Flash-Lite models and thinkingLevel: low to other Gemini 3 models, which matches what worked best in this test. Our model list still includes 2.5 Flash-Lite for sites that already use it; if you are setting up a new key, pick 3.1 or 3.5 Flash-Lite instead. Setup is covered in the MxChat documentation, and MxChat Pro is a one-time licence that adds the WooCommerce, live-agent and other add-ons mentioned in this test.
If you are weighing Gemini against Anthropic for the same job, our Claude Sonnet 5.5 pricing test used the same three questions.
FAQ
How much does Gemini 2.5 Flash-Lite cost?
$0.10 per million input tokens for text, images and video, $0.30 for audio, and $0.40 per million output tokens, as listed on Google’s Gemini API pricing page on 3 October 2026. Context caching is $0.01 per million tokens and the Batch API halves input and output. On our chatbot requests that works out to about $0.43 per 1,000 replies.
Is Gemini 2.5 Flash-Lite being retired?
Google has not published a shutdown date. Its deprecations page says access to the 2.5 models is limited to users who have actively used them before, and on 3 October 2026 the API returned HTTP 404 “no longer available to new users” to a key that had not. Existing users can still call it. Some trackers cite 16 October 2026; we found no Google source for that date.
How much does Gemini 3.5 Flash-Lite cost?
$0.30 per million input tokens (text, image, video and audio) and $2.50 per million output tokens, including any thinking tokens. Context caching is $0.03 per million plus $1.00 per million tokens per hour of storage, and Batch is half price. It came to $1.35 per 1,000 chatbot replies in our test.
Is Gemini 3.1 Flash-Lite cheaper than 3.5 Flash-Lite?
Yes. $0.25 input and $1.50 output against $0.30 and $2.50. On a retrieval chatbot, where input is over 90% of the bill, that made 3.1 Flash-Lite 17% cheaper per reply ($1.12 against $1.35 per 1,000). Note that 3.1 Flash-Lite is scheduled to shut down on 7 May 2027.
Does Gemini Flash-Lite support thinking?
Yes, if you ask for it, but it is off by default. 3.1 and 3.5 Flash-Lite accept thinkingLevel low and high. For a chatbot we would leave it off: on 3.5 Flash-Lite, high thinking cost 2.5 times as much, took 4.5 times as long, and cut off or broke half of the replies at a 1,000-token output limit. 3.5 Flash-Lite rejects thinkingBudget: 0 with a 400 error.
What does gemini-flash-lite-latest point to?
On 3 October 2026 the alias was served by gemini-3.5-flash-lite, according to the modelVersion field in each response. Aliases move when Google releases a new model, and the price moves with them.
Method note: 45 generateContent calls, 20 parameter probes, 4 countTokens calls and 1 models.list call from the mxchat.ai server on 3 October 2026, maxOutputTokens 1,000, no streaming, no explicit cache. Prices are standard paid-tier list rates read from Google’s Gemini API pricing page the same day. Costs price the promptTokenCount, candidatesTokenCount and thoughtsTokenCount each response reported. The 2.5 Flash-Lite figure is priced, not measured, because the API refused our key. Response times are full round trips from one server on one morning.