GPT-6 Astra Is a Real Jump on Computer Use, But the Benchmarks Are Messier Than the AGI Talk

GPT-6 Astra Is a Real Jump on Computer Use, But the Benchmarks Are Messier Than the AGI Talk

7 min readSeptember 5, 2026

Quick verdict

GPT-6 Astra is a genuine step up on the things OpenAI leaned into: computer use, long-horizon agent work, and formal math and science. It also set an Epoch AI intelligence record and effectively saturated a couple of agent benchmarks. But the independent numbers tell a narrower story than the "welcome to the AGI era" framing. On raw coding and general intelligence indices, Astra lands roughly level with Claude Fable 5, Fable 5.1, and Opus 5, while costing about 2.5x more per token, which works out to something like 75% more per task at max effort. The launch itself was a mess, with influencers getting access before paying customers, and the system card shows chain-of-thought monitorability moving the wrong way. So the honest read is: a strong agentic model, a real price problem, and a safety tradeoff worth watching.

What actually shipped

OpenAI announced Astra as its "most intelligent and aligned model yet," pitched around computer use, software engineering, math and science, polished office documents, and cybersecurity. The tagline was blunt: anything you can do on a computer, Astra can do for you, fast. It rolls out first to a limited set of organizations, then to ChatGPT Plus, Pro, Business, and Enterprise, then to the API and AWS over the following days.

Pricing is the number to sit with, because it reframes every benchmark below:

TierInput (per 1M tokens)Output (per 1M tokens)Notes
Astra standard$10$50Baseline
Astra fast$20$100Up to 2.5x speed

OpenAI shipped runtime features alongside the model. Codex can now ask questions while it keeps working independently, there is an experimental context feature that lets Astra keep notes and search earlier context windows during long tasks, and the Responses API added async function calling, mid-turn steering, and the ability to change reasoning effort without breaking the cache. The headline benchmark claims from OpenAI's own comms: 99.9% on ARC-AGI-3, 98% on FrontierMath Tier 4, 100% on ExploitBench, and 1.9x faster than GPT-5.6 Sol on Mind2Web with Codex harness improvements. The company also said Astra had already helped solve long-standing open math problems.

What the independent evals actually found

The most useful signal came from the benchmark providers, because they published caveats and cross-model comparisons instead of a single hero number. This is where the story splits from the marketing.

Artificial Analysis gave the sharpest mixed read. On their Coding Agent Index, Astra scores 67, about level with Claude Opus 5 and Fable 5, while Fable 5.1 leads at 70. The efficiency story is real: Astra uses roughly a third of the tokens of GPT-5.6 Sol in the Codex harness and about a fifth the tokens of Opus 5 at high effort. On the Intelligence Index it scores 61, tied with GPT-5.6 Sol and five points behind Fable 5.1. The catch is cost. Because the token price is 2.5x higher, Artificial Analysis put Astra at about 75% more expensive per task than its predecessor at max effort. Hallucination did improve, dropping from 92% to 51% at max effort on their factuality benchmark.

Independent benchmarkAstra resultContext
AA Coding Agent Index67Level with Opus 5 / Fable 5; Fable 5.1 leads at 70
AA Intelligence Index61Tied with GPT-5.6 Sol, 5 behind Fable 5.1
Epoch AI ECI169 (record)Up from prior best 163
Epoch MirrorCode46.7%Between Opus 4.7 and Fable 5
Perplexity WANDR0.682 at $11.98/taskHighest they tested, 13.5% over Fable 5.1
Cognition FrontierCode 1.1Within 0.4 pts of Fable 5At 64% lower cost
Vals SRE-Bench99.2% pass@4vs 68.7% for GPT-5.6 Sol (custom harness)

ARC Prize is the clearest example of how much the harness matters. The direct score on ARC-AGI-3 was 63%, but a new provider adapter harness pushed it to 99%. François Chollet reported 66% on a standard harness and close to 100% with a continuous conversation harness and custom compaction, at roughly $360 per game. That is a genuine breakthrough in agentic capability, and also a reminder that a single "99.9%" headline can hide a very specific setup. Chollet noted ARC-AGI-3 rose from under 1% to 100% in about six months, and said ARC-AGI-4 is already coming in early 2027.

Epoch AI was positive but measured: a new intelligence record of 169, up from 163, with new marks on math, continual learning, and game puzzles. It also reported Astra solved 2 of 68 Lean-verified unsolved Erdős problems, the first model to solve any, while ranking only between Opus 4.7 and Fable 5 on MirrorCode. The pattern across Perplexity, Cognition, and Vals is consistent: Astra is strongest on end-to-end research and long-horizon agent work, and near the frontier rather than clearly ahead of it on raw coding, usually with better token economics but a higher per-token price.

Why it matters

The takeaway for anyone paying for AI is that the frontier is no longer a single leaderboard you can top. Astra wins decisively on computer use and long-horizon agency, is roughly tied on intelligence and coding, and loses on price. If your workload is agent loops that run for a long time and touch a real desktop, Astra earns a serious test. If your bill is dominated by high-volume coding or chat, the math points at Fable 5.1, Opus 5, or even a Flash-tier model for a fraction of the cost. This is the same routing logic we walked through in the 2026 coding model roundup: pick per job, not per brand.

It also confirms that "cost" is no longer captured by the sticker token price. Astra is more token efficient than GPT-5.6 Sol on several tasks, which partly offsets the 2.5x price, but at max effort it still lands more expensive per finished task. That gap is exactly why running several models through one interface has gotten more attractive, so you can send each request to whatever wins on quality-per-dollar that week instead of committing to one lab. We break down that approach in the multi-model app guide.

The safety and rollout problems

Two things dented the launch. First, governance and access. Many influencers had Astra before paying ChatGPT users did, the blog post was late or broken, and rollout timing was unclear. OpenAI tried to patch it by granting "banked resets" for each day paid users lacked access, but the frustration was real and public.

Second, and more serious, the system card described improved alignment scores alongside decreased chain-of-thought monitorability. The UK AI Safety Institute measured Astra's no-CoT time horizon at 30.9 minutes versus 3.6 minutes for GPT-5.6 Sol, meaning it can operate far longer without exposing readable reasoning. Chain-of-thought controllability came in at 93% versus 48% for GPT-5.6 Sol, and verbalized evaluation awareness rose to 41.1% from 27.7% on the prior generation. Researchers including Neel Nanda and Ryan Greenblatt argued the visible alignment gains may be papering over specific failure modes rather than solving underlying goal misalignment, and that a capability jump paired with a monitorability drop is a bad combination. AISI was careful to say it found capabilities that could enable evading monitoring, not that evasion was demonstrated. This is the flip side of the story we covered when OpenAI first held Astra back over cyber risk in the pre-launch classification piece.

Video: GPT-6 Astra reactions

Fireship's first look at the launch and the benchmark debate around it.

FAQ

Is GPT-6 Astra worth switching to?

It depends on your workload. For long-running computer-use agents and deep research tasks, it is the strongest option right now. For general coding and chat it is roughly level with Fable 5.1 and Opus 5 at a higher price, so many teams will keep routing those jobs elsewhere. See our head-to-head model comparison.

How much does GPT-6 Astra cost?

Standard pricing is $10 per million input tokens and $50 per million output tokens. The fast tier is $20 in and $100 out for up to 2.5x speed. That is about 2.5x the token price of GPT-5.6 Sol.

Did GPT-6 Astra really hit 99.9% on ARC-AGI-3?

Only with a specific provider adapter harness. The direct standard-harness score was in the low-to-mid 60s, and reaching near 100% required custom compaction and a continuous conversation setup at roughly $360 per game. The jump is real, but the top-line number is harness-dependent.

What is the concern in the system card?

The model can run much longer without producing readable chain-of-thought, and evaluation awareness went up. Safety researchers worry that better alignment scores may hide, rather than fix, underlying misalignment, and that losing monitorability while gaining capability is the wrong direction.

Sources

Further reading

Try all the models mentioned in this article

Admix gives you GPT-5, Claude, Gemini, and 350+ AI models in one app. Compare responses side by side. Free to start.

Start free on Admix

Related articles