Gemini 3.8 Flash Leads on Reasoning and Vision at $0.75 In, and Ships a Cyber Variant

Gemini 3.8 Flash Leads on Reasoning and Vision at $0.75 In, and Ships a Cyber Variant

6 min readSeptember 4, 2026

Quick verdict

Gemini 3.8 Flash is the clearest sign yet that the cheap tier is eating the expensive tier. The circulating benchmark table shows it leading or tying frontier models on agentic, reasoning, legal and finance, vision, and bio tasks while holding Flash pricing at roughly $0.75 input and $3.75 output per million tokens. Claude Opus 5 still wins the heavy coding and long-horizon agent benchmarks, so this is not a clean sweep. But if your bill is search, reasoning, vision, or document work rather than deep coding loops, 3.8 Flash is the model to route to this week. Google also shipped a security-focused sibling, 3.8 Flash Cyber, at the same speed and price.

What actually shipped

Two releases landed together. The main one is Gemini 3.8 Flash, positioned as the fast, cheap workhorse but with benchmark numbers that reach into frontier territory. The table making the rounds compares it against Gemini 3.7 Flash, Claude Opus 5 and Sonnet 5, and the GPT-5.6 variants, and 3.8 Flash tops or matches the field on several agentic, reasoning, legal and finance, vision, and bio benchmarks. The price stayed at the Flash tier, around $0.75 input and $3.75 output per million tokens, which is the same range 3.7 Flash launched into.

The gap that remains is coding. Claude Opus 5 still leads the heavier coding and agent benchmarks, DeepSWE, Terminal-bench 4.0, and OSWorld-2.0 among them. So the honest read is a split decision: 3.8 Flash pulls ahead on breadth and price, Opus 5 keeps the deepest coding and long-horizon work. Early testers also say 3.8 Flash feels quicker than 3.7 Flash and shines on search-heavy tasks, where Google's own search integration returns faster than the Claude search experience.

The second release is Gemini 3.8 Flash Cyber, which Sundar Pichai called Google's most capable cybersecurity model while keeping Flash speed and pricing. The reported numbers:

BenchmarkGemini 3.8 Flash Cyber
CyberGym86.2%
CWE-Bench (patching)47.2%
Internal vulnerability discovery (20 languages)70%+

A dedicated security model at Flash economics is a real shift. Vulnerability discovery and patch generation are exactly the kind of high-volume, run-it-in-a-loop work where per-token cost decides whether a team can afford to run it across a whole codebase or only on the risky parts.

Why it matters

The pattern here is bigger than one model. A year ago the Flash tier was where you sent the cheap, low-stakes traffic and kept anything hard on a frontier model. 3.8 Flash breaks that split on everything except deep coding. When a model at $0.75 input leads on reasoning, vision, and legal or finance work, the reason to pay frontier output rates for those jobs mostly disappears. That is the same math that made Gemini 3.7 Flash worth testing, and 3.8 pushes it further.

The catch is not the model, it is Google's developer surface. Practitioners flagged rough ergonomics around harnesses and third-party integration, and a sharper worry: aggressive account bans tied to a core Google identity. One warning noted the blast radius can reach Google Cloud accounts linked to the same login, not just Gmail or Workspace. If you are building on Gemini at volume, isolate the identity you use for API access from anything you cannot afford to lose. Benchmarks are one thing, account risk is a separate cost that does not show up on the pricing page.

Video: Gemini 3.8 Flash first look

A hands-on walkthrough of the speed, the benchmark claims, and where 3.8 Flash lands against the frontier.

How to think about the choice

If you pay per token and your traffic is a mix, the right move is routing, not switching. Send reasoning, vision, search, and document work to 3.8 Flash and keep the hard coding loops on Opus 5 or your current frontier model. You get most of the savings without giving up the one thing Flash still trails on. Our guide to cutting coding-agent costs with model routing covers how that split works in practice, and running several models from one place makes the switch cheap to test. For where the coding leaders actually stand, see the best AI models for coding in 2026.

FAQ

How much does Gemini 3.8 Flash cost?

It sits at the Flash tier, roughly $0.75 input and $3.75 output per million tokens, the same range 3.7 Flash launched into. That is well under frontier output rates, which is the whole point of the release.

Is Gemini 3.8 Flash better than Claude Opus 5?

It depends on the task. The benchmark table shows 3.8 Flash leading or tying on agentic, reasoning, legal and finance, vision, and bio work, while Opus 5 still leads the heavy coding and long-horizon agent benchmarks like DeepSWE, Terminal-bench 4.0, and OSWorld-2.0. For a cross-lab view, see our Grok vs Claude vs Gemini comparison.

What is Gemini 3.8 Flash Cyber?

It is a security-focused variant Google calls its most capable cybersecurity model, running at Flash speed and pricing. Reported results include 86.2% on CyberGym, 47.2% on CWE-Bench for patching, and 70% or better on an internal vulnerability-discovery benchmark across 20 programming languages.

Is it safe to build on Gemini for production?

The model is fine, the account risk is the concern. Developers have flagged aggressive Google account bans that can extend to Google Cloud accounts tied to the same identity. Use an isolated identity for API access rather than your primary Google login.

Sources

Further reading

Try all the models mentioned in this article

Admix gives you GPT-5, Claude, Gemini, and 350+ AI models in one app. Compare responses side by side. Free to start.

Start free on Admix

Related articles