Gemini API pricing was checked on August 12, 2026 against Google’s official Gemini API pricing page. The live table above shows Standard token rates from our current dataset; the tier details below cover official charges the base schema cannot fully represent.
For the broader catalog, use the Google AI pricing page, full AI API pricing table, and token cost calculator. You can also compare the current model ladders in Gemini vs GPT pricing.
Gemini API Pricing: Quick Reference
Gemini 3.6 Flash and Gemini 3.5 Flash-Lite GA are the two practical starting points. Both appear on the current official rate card without a Preview label; Google explicitly describes 3.5 Flash-Lite as GA. Gemini 3.1 Pro Preview and Gemini 3 Flash Preview should be treated as preview models.
The headline table uses Standard token rates in USD per 1 million tokens. Do not mix those figures with Batch, Flex, or Priority pricing. Google advertises Batch API processing as a 50% cost reduction, but each model’s official table is authoritative for the exact execution-class rate.
Pro pricing changes when a prompt crosses 200K tokens:
| Model and prompt size | Standard input | Standard cached input | Standard output |
|---|---|---|---|
| Gemini 3.1 Pro Preview, at or below 200K | $2.00 | $0.20 | $12.00 |
| Gemini 3.1 Pro Preview, above 200K | $4.00 | $0.40 | $18.00 |
| Gemini 2.5 Pro, at or below 200K | $1.25 | $0.125 | $10.00 |
| Gemini 2.5 Pro, above 200K | $2.50 | $0.25 | $15.00 |
Gemini 3.5 Flash remains listed on Google’s official rate card. An AI Pricing Guru internal lifecycle label such as “legacy” does not mean Google has removed the model; the same is true of the Gemini 2.5 models still present on the checked rate card.
Which Gemini Model Should You Use?
Gemini 3.5 Flash-Lite GA for high-volume utility work
Start with 3.5 Flash-Lite for classification, routing, extraction, short summaries, moderation pre-checks, and lightweight RAG. It leads the low-cost end of the current shortlist and is the better first test when answers can be mechanically checked or escalated.
Gemini 3.6 Flash for the production default
Use 3.6 Flash for user-visible support, product Q&A, multimodal app flows, routine RAG, and extraction that also needs a useful explanation. Route difficult or high-value requests upward only after Flash misses your quality target.
Gemini 2.5 Pro or 3.1 Pro Preview for premium work
Test 2.5 Pro when you need stronger synthesis, coding, document analysis, or multimodal reasoning and can stay within its lower context tier. Test 3.1 Pro Preview for the newest premium Google behavior, but budget both Pro models against the prompt’s 200K threshold. The input side of the rate changes for the entire prompt once that threshold is crossed.
Caching, Audio, and Grounding Charges
Cached-token processing and cache storage are separate line items. For the shortlisted Standard text/image/video rates, cached input can be 90% below Standard input, but that is not a universal discount across every modality or execution class.
The official Standard card lists cache storage at $1 per 1M cached tokens per hour for the lead Flash models and $4.50 per 1M cached tokens per hour for Gemini 3.1 Pro Preview and Gemini 2.5 Pro. Include storage duration in the estimate; a low cached-input processing rate does not eliminate the hourly storage charge. See cached tokens explained for the mechanics.
Audio also needs its own check. Gemini 3 Flash Preview, Gemini 3.1 Flash-Lite, and Gemini 2.5 Flash charge higher Standard input and cached-input rates for audio than for text, image, or video. Gemini 3.5 Flash-Lite’s current Standard input row covers text, image, video, and audio at the same listed rate. Use the official modality row rather than applying the live table’s base input figure blindly.
Google Search and Google Maps grounding have request or query charges separate from token charges. The token chart and calculator on this page exclude grounding, so add those charges when evaluating a grounded search, travel, local, or research workload.
Free Tier and Production Planning
Free availability is model- and execution-class-specific, subject to Google’s rate limits and terms. Do not assume that Gemini Pro access is universally gone: the current Gemini 3.1 Pro Preview table shows no free-tier rate, while the current Gemini 2.5 Pro table shows free-of-charge token use.
Use the official model row to confirm eligibility before testing, then forecast production with paid rates. A free-tier label does not guarantee enough quota or the execution class your deployment needs.
Recommended Gemini Routing Setup
| Workload | Start with | Escalate when |
|---|---|---|
| Classification, routing, extraction | Gemini 3.5 Flash-Lite GA | Validation fails or nuance matters |
| Support, product Q&A, routine RAG | Gemini 3.6 Flash | Quality tests miss the target |
| Premium RAG and document analysis | Gemini 2.5 Pro | Newer preview behavior is required |
| Frontier reasoning or multimodal analysis | Gemini 3.1 Pro Preview | A stable model is required instead |
| Long-context review | Either Pro model after tier modeling | Retrieval can reduce the prompt below 200K |
| Grounded search or local answers | Gemini 3.6 Flash plus grounding budget | Higher reasoning quality justifies Pro |
Measure cost per accepted answer, not only cost per token. A useful routing sequence is Flash-Lite for verifiable utility calls, 3.6 Flash for normal user-facing requests, and Pro only for tasks that fail the cheaper model’s evaluation. Use the inline estimator above or the full token calculator with real prompt and output sizes.
FAQ
What is the best low-cost Gemini API model?
Gemini 3.5 Flash-Lite GA is the low-cost lead model in this checked shortlist. Start there for high-volume work that can be validated, and use Gemini 3.6 Flash when the user sees the answer and quality needs to be higher.
How much does Gemini 3.1 Pro Preview cost above 200K tokens?
For prompts above 200K tokens, Standard pricing is $4 input, $0.40 cached input, and $18 output per 1M tokens. At or below 200K, the corresponding rates are $2, $0.20, and $12.
Does Gemini cached input always cost 90% less?
No. The 90% reduction applies to the shortlisted Standard text/image/video processing rates, not every modality or execution class, and separate hourly cache-storage fees still apply.
Does Gemini Pro still have a free tier?
It depends on the model. The checked rate card shows no free-tier rate for Gemini 3.1 Pro Preview, while Gemini 2.5 Pro shows free-of-charge token use subject to Google’s limits and terms.
Are grounding charges included in Gemini token prices?
No. Google Search and Maps grounding request or query charges are separate from token charges and are excluded from the examples and calculator on this page.
Bottom Line
For a new Gemini deployment, benchmark Gemini 3.5 Flash-Lite GA as the utility layer and Gemini 3.6 Flash as the user-facing default. Move to Gemini 2.5 Pro or Gemini 3.1 Pro Preview only when evaluations justify it, and model the 200K context threshold explicitly.
Before launch, account for cache storage, audio modality rates, and grounding separately. Recheck the official Google rate card, then use our Google pricing page and full comparison table for the latest normalized token data.