Gemini 3.8 Flash Release: Performance, Price, and Is It Free
Gemini 3.8 Flash shipped on September 2, 2026. Flash is known as the workhorse tier that sits below Pro, Google's smartest grade, handling high volumes of work cheaply and quickly rather than the hardest problems. This 3.8 is a follow-up to 3.7 Flash from three weeks earlier, and Google's third Flash release in six weeks.
What made the announcement stand out is that this budget workhorse tier caught up with far more expensive top-tier models on several tests. The price stays exactly where 3.7 had it, with only the performance going up, so anyone already on 3.7 gets a better model at the same cost.
Read the official scorecard to the end, though, and the picture shifts. There are tests it wins big and tests where it lands under half the top score, so the verdict flips depending on which row you look at. There is something left to check before switching on the strength of a headline.
This post covers the Gemini tier system, performance, pricing, and the free-access conditions as of September 2026, and ends with a judgment on whether switching from 3.7 is actually worth it.
- Gemini 3.8 Flash
A general-purpose AI model Google built by refining 3.7 Flash. It is designed to handle long coding tasks, multi-step automation, and complex analysis at low cost.
The Gemini tiers: Pro on top, Flash as the workhorse
Google's official model list splits Gemini into three tiers. Complex work that needs deep reasoning goes to Pro, the smartest tier. Work where response speed and volume matter goes to Flash, the best price-to-performance tier. Below that sits Flash-Lite, cheaper and lighter still.
There is one thing to watch when reading version numbers. Flash has reached 3.8, but as of September 3, 2026 the newest Pro on the list is Gemini 3.1 Pro. The two tiers are numbered independently, so a bigger number does not mean a higher tier. 3.8 Flash simply came out later than 3.1 Pro; it is not the smarter model. For reference, as of September 4, 2026, 3.1 Pro still carries a Preview label while 3.8 Flash is Stable.
What improved in 3.8 Flash: coding and agentic work
The official announcement claims three improvements over 3.7: software work, multi-step agentic tasks, and complex knowledge work. As the model card confirms, this is not a new family but a refinement built on 3.7.
The basic specs: the model ID is gemini-3.8-flash, it takes up to 1 million tokens of input, and it outputs up to 64K tokens. Beyond text it accepts images, audio, and video, so you can feed it a whole document or have it analyze a screen recording together with the audio.
The knowledge cutoff is March 2026, though the model card adds that in some domains the knowledge may only reach January 2025. For questions where recency matters, hand it search results or source material. The effort setting, which controls how much work goes into an answer, is still supported.
The same announcement carries a security-specialized sibling, Gemini 3.8 Flash Cyber. Despite the similar name it is for a completely different audience, covered separately below.
Same price as 3.7, and the API has a free tier
The Gemini API pricing page lists 3.8 Flash and 3.7 Flash at identical rates. The performance went up while the per-token price stayed put.
| Period | Input | Output |
|---|---|---|
| Through December 31, 2026 | $0.75 / 1M tokens | $3.75 / 1M tokens |
| From January 1, 2027 | $1.50 / 1M tokens | $7.50 / 1M tokens |
The Gemini API also has a free tier. As of September 3, 2026, both input and output are free, and thinking tokens count as free output too. The free tier does come with per-minute request and token limits and a daily request cap; the exact numbers only show in the AI Studio rate-limit screen for your signed-in account, so check them before pointing a high-frequency automation at the free tier. One more thing: per the pricing page, prompts and responses on the free tier can be used to improve Google's products, while the paid tier's are not. If you handle sensitive material like client data, weigh that condition before using the free tier.
Also worth noting: identical rates do not guarantee identical bills. Official materials say 3.8 Flash takes more reasoning steps and calls tools more often on complex work, and higher effort settings burn more tokens. Measure actual tokens per task rather than reading the rate card, and for simple work check whether a low effort setting already gives you the quality you need.
Results vary by task: tests it wins and tests it loses
The official model card lists a test called Terminal-bench in two versions. The same name makes them look like one test, but the results point in opposite directions.
| Official benchmark | 3.8 Flash | 3.7 Flash | Opus 5 | Sonnet 5 | GPT-5.6 Sol | GPT-5.6 Terra |
|---|---|---|---|---|---|---|
| Terminal-bench 2.1 | 89.4% | 85.8% | 89.1% | 80.4% | 88.8% | 87.4% |
| Terminal-bench 4.0 | 19.1% | 11.2% | 51.8% | 12.4% | 37.3% | 23.6% |
Look only at the 2.1 row and 3.8 Flash narrowly beats Claude Opus 5. Pull that row alone and you get the headline "Flash beat Opus". In the 4.0 row, the ranking reverses and Opus 5 leads by roughly 2.7 times.
The reason is the scoring method. According to the evaluation methodology, 2.1 gives every model the same tooling, while 4.0 takes each model's best-scoring configuration from the leaderboard. They share a name, but they are effectively different tests. Google hid nothing; both rows sit side by side in the same table. The reader just has to notice which row they are looking at.
The remaining tests split into wins and losses the same way. The full numbers are in the table below; compare models within a row only. Lower is better for the two price rows, higher for everything else.
| Item | 3.8 Flash | 3.7 Flash | Opus 5 | GPT-5.6 Sol |
|---|---|---|---|---|
| Intro input price ($/1M tokens) | 0.75 | 0.75 | 5.00 | 4.00 |
| Intro output price ($/1M tokens) | 3.75 | 3.75 | 25.00 | 20.00 |
| DeepSWE v1.1 | 73.7 | 65.3 | 74.0 | 72.7 |
| GDPVal-AA v2 (Elo) | 1545 | 1482 | 1824 | 1710 |
| Vals Finance Agent v2 | 61.4 | 59.0 | 58.6 | 53.8 |
| Harvey's Legal Agent | 10.0 | 8.8 | 6.7 | 2.5 |
| GDP.PDF | 35.0 | 34.0 | 37.0 | 40.0 |
| CharXiv Reasoning | 86.2 | 84.5 | 83.7 | 85.8 |
| LVBench | 87.8 | 85.4 | 75.4 | 82.1 |
| HLE-Verified | 54.9 | 53.6 | 54.4 | 54.5 |
| OSWorld-2.0 | 59.0 | 50.6 | 75.4 | 62.6 |
| BioMysteryBench (easy) | 88.8 | 87.1 | 90.1 | 79.5 |
| BioMysteryBench (hard) | 56.5 | 43.5 | 49.4 | 44.7 |
| LABBench2 | 86.2 | 82.1 | 84.2 | 82.1 |
The table in one line: on coding, charts, and long video it matches or beats top-tier models at a seventh of the price, while on computer control and broad knowledge work Opus 5 still leads by a wide margin.
To unpack that a little: on long coding work (DeepSWE) it rose sharply from 3.7 to land at effectively the same score as Opus 5, and on chart reading and long-video understanding it tops every model in the comparison. One caveat: the long-video (LVBench) row is not measured under identical conditions across models, so treat that row as a reference. On the losing side, the gaps in computer control (OSWorld) and knowledge work (GDPVal) are large, and specialist PDF reading barely moved from 3.7. It is genuinely strong for the price, but it is not a model that wins at everything.
Reactions: praise for performance, worry about release pace
The launch made noise. As of September 3, 2026 the release thread on Hacker News, the overseas developer community, held 1,058 points and 598 comments, and a benchmark roundup on Reddit's Gemini board drew 275 upvotes. These are live boards, so the numbers shift by the hour, and everything below is user opinion rather than official verification.
The praise centered on disbelief that a workhorse tier posted these scores. A Reddit post arguing that 3.8 Flash effectively plays at the level of the upper Pro tier drew 112 upvotes, and under it a user reported that on a TypeScript task 3.8 Flash easily fixed bugs that 3.1 Pro had introduced four months earlier. Others said it simply feels far more capable than 3.7 in daily use.
The skeptics led with hands-on failures. One user's 1 hour 45 minute meeting recording got truncated mid-transcription and the model refused to finish; others said it cannot sustain heavy reasoning like long reports or calculations, and that its translation work trails 3.1 Pro. On Hacker News, a user running token-heavy coding tasks reported it in no way compares to Claude or GPT-5.6 Sol, needing more prompts for weaker output. The benchmarks themselves drew suspicion too. Several comments looked at the Terminal-bench 4.0 drop and asked whether this is a model trained to the benchmark, with a counterpoint that 4.0 is simply a harder test built after the earlier version was mostly solved. The community, in other words, spotted exactly the split described in the previous section.
Korean Threads showed the same temperature difference. One post shared a screen of the Gemini app answering that its current model is 3.8 Flash, confirming the rollout, while a negative review with over thirty thousand views collected comments like "the same lie every version", pointing at the gap between announced scores and lived experience.
There was also concern about the release pace itself. Three Flash releases in six weeks means the next version lands before a team finishes evaluating the last one. Models do not vanish overnight, though. As of September 3, 2026, 3.5 Flash and 3.6 Flash are still Stable, and the only retirements are the 2.0 family and models that carried a Preview label. The announcement itself says to keep using 3.7 Flash where compute cost matters. The thing to watch is tools like Antigravity: API models linger, but a tool's model picker changes faster, so if you run automation, know which of the two you depend on.
Access: apps need a Pro or Ultra subscription, the API has a free tier
Pulling together the paths named in the announcement and the model card:
| Use | Path |
|---|---|
| General use (Google AI Pro or Ultra subscribers) | Gemini app, AI Mode in Google Search, Google Sheets |
| Development | Google AI Studio, Gemini API, Antigravity, Android Studio, Stitch |
| Enterprise | Gemini Enterprise Agent Platform |
The first row is the one to read carefully. To use 3.8 Flash in the everyday surfaces, the Gemini app, AI Mode in Google Search, and Google Sheets, all three paths require a Google AI Pro or Ultra subscription. The free route is the developer-side API free tier covered earlier. The materials say nothing about per-country rollout timing, so check each service's screen to see whether it has reached your account.
The model card lists weaknesses too: it can make things up, it occasionally slows down or drops responses, and its jailbreak resistance is still being worked on. Read that as a signal to trial it on small tasks before wiring it into long-running automation.
The Cyber variant is not for regular users
The name suggests a hardened security edition of 3.8 Flash, but Gemini 3.8 Flash Cyber is a model for an entirely different audience. It is limited to participants in the Fairwind Program, aimed at government authorities, critical infrastructure operators, and software maintainers, and most individual users do not qualify. No price has been published.
Its strength is finding and fixing vulnerabilities. Per the announcement, it passed 70% on a real-vulnerability discovery evaluation spanning 20 programming languages, and on CWE-Bench, which measures automated patching, its first-try pass rate of 47.2% nearly matched a leading frontier model's 47.8%. The Chrome security team reported it produced 2.6 times more correct patches than the best commercial model, and the security firm Wiz measured 7.5 to 9.7 percent higher recall at 2.3 to 5.2 times lower cost.
Why gate it? The regular 3.8 Flash ships with safeguards against misuse in chemical, biological, radiological, and nuclear domains and in cyber offense. Cyber loosens some of those safeguards so security work can go deep. The ability to find vulnerabilities well is the same ability that attacks with them, which is why it opens only to trusted defenders.
Both models also made a significant jump in prompt injection resistance as measured by Gray Swan. Prompt injection is an attack where instructions hidden in an external document or webpage displace your original command, and the risk grows the longer an agent works across files and the web.
My take: start testing 3.8 on whatever ran on 3.7
For new work I would try 3.8 Flash before 3.7. The API rate is identical while the scores for coding, charts, long video, and expert reasoning went up, so anyone on 3.7 can check for better results at zero added cost.
I would not move everything at once, though. Official numbers show Opus 5 far ahead on computer control and broad agentic work. For automation that matters, pick one representative task, run it through both models with the same input, the same effort, and the same tools, then compare success and token burn before switching.
If you mostly feed it specialist PDFs, there is no rush: that score barely moved from 3.7 and GPT-5.6 Sol is higher. Starting free follows the same order. Run a short trial on your own material in the API free tier, and widen the scope once the results hold. In the end you choose not by who won the announcement, but by your own success rate, token cost, and whether it stays up.
- Flash is the tier below Pro: cheap and fast, while Pro handles hard reasoning better. The tiers are numbered independently, so 3.8 Flash does not outrank 3.1 Pro.
- Pricing matches 3.7, and the Gemini API has a free tier. The consumer apps require a Google AI Pro or Ultra subscription.
- On Terminal-bench 2.1, 3.8 Flash edges out Opus 5; on 4.0, Opus 5 leads by about 2.7 times. The scoring conditions differ.
- It is strong for the price on coding, charts, and long video, while Opus 5 leads on computer control and broad knowledge work.
- The Cyber variant is limited to Fairwind Program participants. Regular users should look at 3.8 Flash.
Frequently asked questions
Can I use Gemini 3.8 Flash for free?
It depends on the path. The consumer surfaces, the Gemini app, AI Mode in Google Search, and Google Sheets, require a Google AI Pro or Ultra subscription, so they are not free. The developer-side Gemini API has a free tier: as of September 3, 2026, input and output are free, including thinking tokens. The free tier has per-minute and daily request limits that vary by model and tier, and the exact numbers show in AI Studio for your signed-in account. Note that free-tier prompts and responses can be used to improve Google's products.
What is the difference between Gemini Flash and Pro?
The official model list describes Flash as the best price-performance tier for fast, high-volume work and Pro as the upper tier for complex work needing deep reasoning. Watch the numbering: as of September 3, 2026 the newest Pro is Gemini 3.1 Pro while Flash has reached 3.8. A bigger number does not mean a higher tier.
Is it worth switching from Gemini 3.7 Flash?
The API price is the same and the coding and agentic scores went up, so it is worth testing on new work. Specialist PDF reading barely improved, and higher effort settings can burn more tokens, so compare on a representative task first.
Where can I use Gemini 3.8 Flash?
The Gemini app, AI Mode in Google Search, and Google Sheets serve Google AI Pro and Ultra subscribers. Developers get Google AI Studio, the Gemini API, Antigravity, Android Studio, and Stitch, and enterprises use the Gemini Enterprise Agent Platform.
Sources (11)Expand to see all sources
- Google official blog, Introducing Gemini 3.8 Flash and 3.8 Flash Cyber, release, pricing, Cyber performance and access (checked 2026-09-03)
- Google DeepMind, Gemini 3.8 Flash Model Card PDF, specs, benchmarks, knowledge cutoff and limitations (checked 2026-09-03)
- Google DeepMind, Gemini 3.8 Flash evaluation methodology, Terminal-bench scoring conditions by version (checked 2026-09-04)
- Google AI for Developers, Gemini Developer API pricing, free tier, intro and standard rates, data-use conditions (checked 2026-09-04)
- Google AI for Developers, Gemini API Models, tier descriptions, model IDs, support status (checked 2026-09-04)
- Google AI for Developers, Rate limits, free tier per-minute and daily limits (checked 2026-09-03)
- Hacker News, Gemini 3.8 Flash and 3.8 Flash Cyber discussion, user reactions after launch (checked 2026-09-04)
- Reddit r/Bard, Check out The Benchmarks, benchmark reactions and the benchmaxxing debate (checked 2026-09-04)
- Reddit r/Bard, hands-on impressions thread, coding success, transcription failure, translation comparison (checked 2026-09-04)
- Threads, post confirming 3.8 Flash active in the Gemini app (checked 2026-09-04)
- Threads, Korean user review of Gemini 3.8 Flash (checked 2026-09-04)