
Gemini 4 Argon Leads 13 of 19 Benchmarks, and Almost No One Can Use It Yet
Quick verdict
Google is back at the frontier, at least on paper. Gemini 4 Argon takes first place on 13 of 19 benchmarks Google published against GPT-6 Astra and Claude Opus 5.5, matches Astra on Artificial Analysis's Intelligence Index, and tops both the Vals Index and Text Arena. It also raises the output ceiling to a claimed one million tokens and prices in at $2/$10 per million on an introductory discount. The catch is big enough to lead with: you cannot use it. Access starts with government users and vetted cyber defenders through Google's Fairwind Program, and everyone else waits while Google hardens the guardrails. So this is a real return to form for a lab that spent most of the year written off, wrapped in a launch almost nobody can touch.
What actually shipped
Argon is aimed at three jobs: long-horizon coding, enterprise knowledge work, and cyber defense. The headline specs:
- Pricing: $4/$20 per million input/output tokens standard, cut to $2/$10 with a 50% intro discount that has no announced end date. Cached input gets a 95% discount.
- Output limit: Google cites a 1M-token output ceiling, up from 64K. Read the fine print, though. Vals lists the real max output at 262K, and Artificial Analysis only reached 1M through a new API feature called Long Decode Continuation, which pauses a long response and resumes it across calls rather than generating a million tokens in one shot.
- Availability: trusted testers only. Government users and cyber defenders get in first through the Fairwind Program, with developer, enterprise, and consumer access gated behind more safety work on cyber and CBRN misuse, prompt injection, and agent sandboxing.
Google also leaned on internal dogfooding to make its case. It says Argon agents freed more than 300 TiB of data-center memory and are migrating over 800,000 lines of C and C++ kernel code to Rust, including a video decoder where agents replaced 32,000 lines of hand-tuned SIMD with safe Rust that runs 2.7x faster with identical output. The team also claims internal agent loops built on Argon helped finish the CK conjecture. Treat the research claim as a press line until the paper lands, but the Rust migration is the kind of grinding, verifiable work these models are genuinely good at.
Where it lands against Astra and Opus 5.5
The independent evals are the part worth reading, because Google's own slide is a sales pitch. Three outside labs ran it, and they mostly agree it is a top-tier model, with caveats on each axis.
Artificial Analysis put Argon at 53 on its Intelligence Index, level with GPT-6 Astra and one point above GPT-6.1 Sol. On cost per task it comes in cheaper than Astra, but the reason matters:
| Model | Intelligence Index | Cost per task | Avg output tokens/task |
|---|---|---|---|
| Gemini 4 Argon (intro price) | 53 | $1.99 | 62K |
| Gemini 4 Argon (standard price) | 53 | $3.98 | 62K |
| GPT-6 Astra | 53 | $3.26 | 27K |
| GPT-6.1 Sol | 52 | — | — |
The cheaper per-task number comes from the 50% discount, not from efficiency. Argon burns about 62K output tokens per task against Astra's 27K, so once the intro pricing lapses it costs more than Astra at the same intelligence score. That is the single most important line in the whole launch for anyone budgeting around it.
The other labs filled in the rest. Vals ranked Argon number one on its index at 68.9%, and on its coding suite it built 30 Vibe Code Bench apps perfectly against 25 for Opus 5 and 24 for Astra. Its Terminal-Bench 4.0 score jumped from 19.0% to 57.6%. On Google's own card, Argon posts 77.9% on DeepSWE v1.1 (versus 74.2% for Opus 5.5 and 74.1% for Astra), 91.7% on LVBench, and 68% on CWE-bench. Arena has it first in Text Arena at 1525 and first for steerability in the preliminary Agent Arena.
Where it is not first is as telling. On Artificial Analysis's Terminal Bench 4 it lands at 57%, behind Sonnet 5.5, Opus 5.5, and Astra. Its hallucination rate is a standout 15% on AA-Omniscience against Astra's 51%, but the tradeoff is lower raw accuracy, 50% versus Astra's 63%, because it abstains more often. A cautious model that says "I don't know" is the right call for enterprise and defense work, which is exactly who gets it first.
The skeptics have a point
Not everyone bought the benchmark card. On Harvey's legal benchmark, Argon's reported 19.6% actually trails Muse Spark 1.2's listed 25.42%, so the "leads everywhere" framing breaks down the moment you leave the benchmarks Google chose to publish. Others flagged the usual risk with a first-party launch: possible preference-data benchmaxxing, and some specific numbers, DeepSWE among them, that looked too clean. One recurring note from reviewers was that DeepSWE is close to saturated, so topping it says less than it used to, and the real signal is Argon trading blows with Astra on Terminal Bench.
None of this makes Argon a weak model. It makes it a normal frontier launch: genuinely strong, oversold on the vendor's own slide, and best judged once outsiders can run it on their own tasks.
What it means if you pay for AI
For now, nothing you can act on directly, because the model is locked up. But the shape of the launch tells you where the market is heading. The three-way race at the top is real again. Google, OpenAI, and Anthropic all have a credible frontier model within a point or two of each other on general intelligence, and they are increasingly separated by price, speed, and quirks like hallucination rate rather than by a clear capability gap. For a side-by-side of how the three labs stack up, see our OpenAI vs Anthropic vs Google breakdown.
The pricing story is the one to watch. Argon's intro discount makes it look cheap, but the token-hungry reality means the sticker can mislead. The same week, GPT-6.1 Sol shipped at the same $2/$10 while actually using fewer tokens per task, and OpenAI re-tiered ChatGPT Pro so the old $200 plan buys about half what it used to. The pattern across all three labs is the same: the headline per-token price keeps falling, but the real cost depends on how many tokens a model spends to finish your job, and flat subscription tiers keep quietly losing value. The only sane response is to pick the cheapest model that clears the bar for each task instead of marrying one vendor's pricing, which is the whole reason an AI aggregator exists. When Argon finally opens to developers, the question will not be "is it good" (it is) but "is it worth 62K tokens a task when Sol does it in a fraction," and you want to be able to switch between them to find out.
Video: Gemini 4 Argon, early tests
A hands-on look at what Argon can do and where it fits against the rest of the frontier, for anyone who wants more than the benchmark card.
FAQ
Can I use Gemini 4 Argon right now?
Not unless you are a government user or a vetted cyber defender in Google's Fairwind Program. Google says developer, enterprise, and consumer access will follow once it finishes hardening the safety guardrails, but it has not given a date.
How much does Gemini 4 Argon cost?
Standard pricing is $4 per million input tokens and $20 per million output. A 50% introductory discount, with no announced end date, brings that to $2/$10, and cached input gets a 95% discount. Watch the real bill, though: Argon averages about 62K output tokens per task versus 27K for GPT-6 Astra, so once the discount ends it can cost more than Astra at the same intelligence score.
Is Gemini 4 Argon better than Claude Opus 5.5 or GPT-6 Astra?
It leads 13 of 19 benchmarks Google published and matches Astra on Artificial Analysis's Intelligence Index, so it is firmly in the same tier. It trails on a few evals, including Terminal Bench 4 and Harvey's legal benchmark, and the lead on some benchmarks drew benchmaxxing concerns. For how the labs compare overall, see our OpenAI vs Anthropic vs Google guide.
Is the 1M-token output limit real?
Partly. Google cites a 1M-token output ceiling, up from 64K, but Vals measured the real single-response max at 262K. Artificial Analysis only reached 1M using Long Decode Continuation, a new feature that pauses and resumes a long response across API calls rather than producing it in one pass.
Sources
- Google DeepMind - introducing Gemini 4 Argon
- Google - Gemini 4 Argon, our next era of frontier intelligence
- Demis Hassabis - Fairwind trusted-tester rollout
- Philipp Schmid - $2/$10 intro pricing and 95% cache discount
- Artificial Analysis - Argon matches Astra at 53 on the Intelligence Index
- Vals - Argon #1 on the Vals Index at 68.9%
- Arena - Argon #1 in Text Arena at 1525
- Argon agents freed 300 TiB of memory and drive Rust migrations
- teortaxes - benchmaxxing and figure concerns
- r/GeminiAI - Gemini 4 Argon launch discussion
Further reading
Try all the models mentioned in this article
Admix gives you GPT-5, Claude, Gemini, and 350+ AI models in one app. Compare responses side by side. Free to start.
Start free on Admix