
Grok 4.5 Is xAI's First Coding Model, and It Undercuts Opus on Price
Quick verdict
Grok 4.5 is the first model xAI built specifically for coding and agents, and it was trained jointly with Cursor. It costs $2/$6 per million input/output tokens, which is well under what you pay for Opus-class output today. Elon Musk claims it is roughly Opus-class but faster and more token-efficient. Independent benchmarks do not put it at the absolute top, but they do put it in the frontier conversation at much better economics. If you buy AI by the task rather than by the brand, that is the part that matters.
What actually shipped
xAI launched Grok 4.5 as a coding and agent model, describing it as its most credible frontier jump yet. The launch details:
- Pricing is $2 per million input tokens and $6 per million output tokens
- It was trained jointly with Cursor, which is offering it with temporary boosted usage
- Context launched at 500k tokens, and Musk said it should return to 1M within a week
- Cursor treats it as a separate weight class from its own Composer model
On the coding-agent benchmarks xAI published, Grok 4.5 scores 83.3% on Terminal-Bench 2.1, 78.0% on SWE-Bench Multilingual, 62.0% on DeepSWE 1.0, and 64.7% on SWE-Bench Pro. Those numbers sit near, though not above, the current top models.
| Benchmark | Grok 4.5 |
|---|---|
| Terminal-Bench 2.1 | 83.3% |
| SWE-Bench Multilingual | 78.0% |
| DeepSWE 1.0 | 62.0% |
| SWE-Bench Pro | 64.7% |
Where it lands on the leaderboard
Artificial Analysis placed Grok 4.5 at #4 on its Intelligence Index, behind only Fable 5, GPT-5.5, and Opus 4.8. It reported a GDPval-AA v2 Elo of 1543 and a Coding Agent Index of 76, with its best results in agentic knowledge work and coding. The more useful finding was the cost side: Artificial Analysis reported substantially lower task cost than comparable Claude or OpenAI setups, which puts Grok 4.5 in a strong spot on the cost-versus-performance curve.
That framing showed up everywhere in the reaction. Cline and other agent builders landed on the same read: not necessarily state of the art, but frontier-adjacent quality at economics that are hard to argue with. On Reddit, the top comment on the launch thread called the $2/$6 pricing "the real surprise" and argued that output-token throughput and latency may end up mattering more than a few points of benchmark delta.
Why it matters
Two things are worth pulling out. First, xAI did not close the gap through brute-force scaling. It closed it through a training partnership with a product company and agent-specific tuning. A lab improving its position by co-designing with Cursor, rather than just buying more compute, is a different playbook, and it is one other labs will copy.
Second, the pricing resets the default. When a near-frontier coding model costs $2/$6, the question stops being "which model is best" and becomes "which model is good enough for this task at this price." That is the same routing logic enterprises already run, where premium models get rationed for the hard problems and cheaper models carry the bulk of the load. Grok 4.5 gives that strategy another strong, cheap option to route to. For the ongoing squeeze on model margins, see our take on where AI pricing is heading.
The catch is that launch benchmarks are run under ideal conditions. The open question, raised by nearly every practitioner who tried it, is whether those output-token rates and that quality hold up under real production workloads. Test it on your own tasks before you rewire a pipeline around it.
Video: Grok 4.5 vs Opus, tested
A hands-on comparison of Grok 4.5 against Opus 4.8 on real coding work, which is the exact matchup Musk's "Opus-class but cheaper" claim invites.
FAQ
How much does Grok 4.5 cost?
It is priced at $2 per million input tokens and $6 per million output tokens. That undercuts Opus-class pricing while landing in the same benchmark neighborhood, which is why the cost, not the raw score, drove most of the reaction.
Is Grok 4.5 better than Claude Opus for coding?
Not on absolute benchmarks. Artificial Analysis ranks Grok 4.5 at #4 overall, behind Opus 4.8, GPT-5.5, and Fable 5. The argument for it is cost-performance, not a clean win. For a broader head-to-head, see our Grok vs Claude vs Gemini comparison.
What is the context window?
Grok 4.5 launched with a 500k-token context window. Musk said it should return to 1M tokens within about a week of launch, though that had not shipped as of the announcement.
Where can I use Grok 4.5?
It is available through xAI and through Cursor, which co-trained it and launched it with temporary boosted usage. Cursor treats it as a distinct weight class from its own Composer model. If you are weighing coding tools, our guide to the best AI coding agents covers where each one fits.
Sources
- @SpaceXAI - Grok 4.5 launch announcement
- @cursor_ai - Grok 4.5 built jointly with Cursor
- @elonmusk - Opus-class but faster, more token-efficient, and lower cost
- @elonmusk - context expected to return to 1M next week
- @ArtificialAnlys - #4 on the Intelligence Index behind Fable 5, GPT-5.5, Opus 4.8
- @ArtificialAnlys - GDPval-AA Elo 1543, Coding Agent Index 76, cost-performance
- @cline - frontier-adjacent quality at better economics
- xAI - Grok 4.5 pricing and efficiency
- r/singularity - Grok 4.5 is live, benchmark table and pricing discussion
Further reading
Try all the models mentioned in this article
Admix gives you GPT-5, Claude, Gemini, and 350+ AI models in one app. Compare responses side by side. Free to start.
Start free on Admix