Gemini 3.7 Flash Lands Three Weeks After 3.6, at Half the Price

Gemini 3.7 Flash Lands Three Weeks After 3.6, at Half the Price

6 min readAugust 14, 2026

Quick verdict

Gemini 3.7 Flash is a real upgrade to the mid-tier workhorse, and Google shipped it just three weeks after 3.6 Flash. The coding and agent numbers jumped hard, DeepSWE went from 49.0% to 65.3%, and there is a 50% introductory price cut through the end of the year at $0.75/$3.75 per million tokens. If you run coding or agentic work at volume and pay per token, this is the model to test against your own tasks this week. It does not top the frontier, but it lands near it at a price that changes the math.

What actually shipped

Google positioned 3.7 Flash as its default for coding, web development, knowledge work, and agentic workflows. The launch numbers are not a routine bump. The measured deltas over 3.6 Flash:

BenchmarkGemini 3.7 FlashGemini 3.6 Flash
DeepSWE65.3%49.0%
FrontierCode43.6%34.4%
AutomationBench30.4%17.0%
Code Arena (Elo)15881506
Intelligence Index (Artificial Analysis)5652

Pricing runs $0.75 input and $3.75 output per million tokens during the introductory window, then rises to $1.50/$7.50 later. Artificial Analysis measured the high-reasoning variant at 56 on its Intelligence Index, a four-point gain, with roughly 340 output tokens per second and a 1M-token context window, and placed it on both the intelligence-versus-cost and intelligence-versus-speed frontiers. Arena data moved it to #8 on the WebDev Code Arena and #9 on the Text Arena, with a clear jump on the Agent Arena.

The rollout was immediate. 3.7 Flash landed across the Gemini API, AI Studio, Android Studio, Antigravity, Gemini Enterprise, Managed Agents, and Gemini Spark on day one, and propagated straight into third-party tools like Cline, Devin, and VS Code Agents. Philipp Schmid, who works on the Gemini developer side, pointed to better discipline in tool loops: more exploration, more test-running, fewer wasted turns. That maps to why the agentic benchmarks moved more than the raw reasoning score.

Why it matters

Two weeks ago the story was Google slipping off the frontier while xAI and OpenAI traded the top of the Intelligence Index. This release resets that read. Google is not chasing the single highest score, it is competing on the slot that actually moves usage: a fast, cheap model that is good enough at coding and agents to run all day. A four-point index gain matters less than the DeepSWE and AutomationBench jumps, because those are the tasks people pay a per-token bill to run in loops.

At $0.75 input, 3.7 Flash undercuts most frontier output by a wide margin while landing close enough on agentic work that the gap stops mattering for a lot of real jobs. That is the same logic that made Grok 4.6 worth a look at $2/$6, and Google just went cheaper. If your workload is agentic coding, the right move is to run it against your own repo, not the leaderboard, because Code Arena and your codebase are not the same test.

Video: Gemini 3.7 Flash first look

A quick hands-on with the speed and coding behavior, including how the price cut lands against the rest of the mid-tier.

How to think about the choice

If you already pay for a frontier model and your bill is mostly coding loops, 3.7 Flash is worth a routing test. You do not have to drop Opus or Sol to benefit, you just have to send the cheap, high-volume work to the cheap model and keep the hard planning on the expensive one. Our guide to cutting coding-agent costs with model routing covers how that split works, and running several models from one place makes the switch cheap. For where 3.7 Flash sits among coding options, see the best AI models for coding in 2026.

FAQ

How much does Gemini 3.7 Flash cost?

It is $0.75 input and $3.75 output per million tokens during the introductory window through the end of the year, then $1.50/$7.50 after that. Even the later price stays well under most frontier output rates.

Is Gemini 3.7 Flash better than 3.6 Flash?

Yes, on independent benchmarks. DeepSWE moved from 49.0% to 65.3%, AutomationBench nearly doubled to 30.4%, and Artificial Analysis measured a four-point Intelligence Index gain to 56. It is a straight upgrade, and at launch it is cheaper than 3.6 was.

Does it beat Claude Opus or GPT-5.6 Sol?

No. At 56 on the Intelligence Index it sits below the top-tier models on raw score. The pitch is price and speed for coding and agent work, not topping the chart. For a cross-lab view, see our Grok vs Claude vs Gemini comparison.

Where can I use it right now?

It shipped across the Gemini API, AI Studio, Android Studio, and Gemini Enterprise on launch day, and third-party tools including Cline, Devin, and VS Code Agents picked it up the same day.

Sources

Further reading

Try all the models mentioned in this article

Admix gives you GPT-5, Claude, Gemini, and 350+ AI models in one app. Compare responses side by side. Free to start.

Start free on Admix

Related articles