
Cognition Trained Its Own Coding Model, SWE-2, and Says It Matches the Frontier at 70% Less Cost
Quick verdict
Cognition, the company behind the Devin coding agent, has stopped wrapping other labs' models and started training its own. SWE-2 is that model, and Cognition says it reaches parity with leading coding models on the standard evals while costing up to 70% less to run. The company built the algorithm, the infrastructure, and the training data in-house, and says it scaled reinforcement learning to multiple trillions of parameters to get there. Alongside the model it bought Dioxus Labs, a Rust tooling team, to harden Devin's sandbox and testing, and it launched Devin Voice so you can talk to the agent. If the price claim holds up in real work, the interesting part is not the benchmark. It is that a coding agent company decided the cheapest path to a better product was to own the model.
What actually shipped
SWE-2 is a model tuned specifically for software engineering, not a general chat model that happens to code. In its launch thread, Cognition framed the pitch around cost: parity on leading coding evals at up to 70% lower cost than the models people currently point Devin at. The company said it built the whole stack itself, from the training algorithm to the serving infrastructure to the data pipeline, rather than fine-tuning someone else's weights.
The number Cognition kept repeating was the scale of the reinforcement learning run, which it put at multiple trillions of parameters. One detail the team shared is the kind of thing that only matters if you train models: a simple linear length penalty was enough to keep the shape of the training-time Pareto curve intact across different effort levels, so the model stays efficient whether you ask it for a quick fix or a long agent run. Cognition's Silas Alberti and others on the team walked through the training work in follow-up posts.
Why a coding agent company trained its own model
Cognition spent its first two years as a harness. Devin ran on top of frontier models from Anthropic and OpenAI, and Cognition's value was the scaffolding around them: the planning loop, the sandbox, the memory, the integrations. That works until the model bill becomes the product's biggest cost and you have no control over it. Every price change from a lab lands directly on your margin, and every capability you want has to wait for someone else to ship it.
Training SWE-2 flips that. If Cognition can serve a model that is roughly as good on coding tasks for a fraction of the token cost, it controls its own pricing and can pass the savings to users or keep them. That is the same logic pushing every serious agent company toward owning more of the stack, and it explains why the launch led with a cost multiple instead of a leaderboard screenshot. For a company whose earlier FrontierCode benchmark argued that real-world coding is far harder than the headline eval scores suggest, betting on a cheaper model it can iterate on directly is a consistent read of where the value is.
The Dioxus Labs acquisition
On the same day, Cognition said Dioxus Labs is joining the company to work on Devin's virtual machine, computer use, and testing, while continuing to support the Dioxus framework and related Rust open-source projects. Dioxus is a Rust toolkit for building cross-platform apps, and the team behind it knows how to make fast, reliable developer tooling.
The fit is about the environment the agent runs in, not the model. A coding agent is only as good as the sandbox it works inside: how fast it can spin up a VM, how faithfully it can drive a real computer, how well it can run and read tests. Those are unglamorous systems problems, and they are where a lot of agents quietly fail. Buying a Rust systems team to own that layer is a bet that the next round of gains comes from the harness getting more reliable, not just the model getting smarter.
Devin Voice
Cognition also launched Devin Voice, which it says runs on GPT-Live plus SWE-2. GPT-Live is OpenAI's full-duplex voice model, the one that can listen and speak at the same time, which we covered in our write-up of GPT-Live-1. Pairing that with SWE-2 means you can talk to Devin about a task while it works, instead of typing every instruction.
Voice on a coding agent sounds like a gimmick until you picture the actual use: kicking off a task from your phone, checking on a long run without opening a laptop, or steering an agent mid-job by describing what you want changed. Whether it sticks depends on latency and how well the agent handles being interrupted, but it is a clear sign that Cognition sees Devin as something you supervise from anywhere, not a tool you babysit in a terminal.
Video: SWE-2 tested inside Devin
A hands-on run putting SWE-2 through two dozen coding prompts inside Devin, which is the fastest way to see whether the parity claim survives contact with real tasks.
What to watch before you trust the numbers
A first-party parity claim is a starting point, not a verdict. Cognition is grading its own model, and "parity on leading coding evals" leaves room for cherry-picking which evals and which competitors. The real test is the one the video above starts on: does SWE-2 hold up across a spread of everyday tasks, or does it match the frontier on the benchmarks and slip on the messy stuff. Early outside commentary, including from Yben Pan, was cautiously positive but wanted more independent numbers before calling it settled. Treat the 70% cost figure the same way: it is a claim about Cognition's serving cost, and what reaches you depends on how they price it.
What it means if you pay for coding tools
The direction of travel matters more than any single score. When the company that makes Devin decides the way to compete is to train a cheaper coding model and own the sandbox around it, that is pressure on everyone charging premium rates for agent runs. If SWE-2 is even close to as good for most tasks, the going rate for autonomous coding drops, and the case for locking into one expensive tool gets weaker. That is the same reason a lot of developers now route work through an AI aggregator instead of committing to a single subscription, and it is worth watching where SWE-2 lands in the next roundup of the best AI coding agents once independent testers have had time with it.
FAQ
What is SWE-2 and who made it?
SWE-2 is a coding model built by Cognition, the company behind the Devin AI software engineer. It is Cognition's first model trained in-house rather than a wrapper over another lab's model, and it is tuned specifically for software engineering tasks.
Is SWE-2 really 70% cheaper than other coding models?
That is Cognition's own claim: parity on leading coding evals at up to 70% lower cost. It is a first-party number about their serving cost, so wait for independent testing before treating it as settled. Our guide to the best AI coding agents tracks how these tools compare in practice.
What is Devin Voice?
Devin Voice lets you talk to the Devin agent instead of typing. Cognition says it runs on OpenAI's full-duplex GPT-Live model plus SWE-2, so you can give the agent instructions by voice while it works. See our GPT-Live-1 write-up for the voice model underneath it.
Sources
- Cognition - SWE-2 launch, parity at up to 70% lower cost
- Cognition - Dioxus Labs joining to work on Devin's VM, computer use, and testing
- Cognition - Devin Voice, powered by GPT-Live and SWE-2
- Silas Alberti - notes on the SWE-2 training run
- Yben Pan - early outside commentary on SWE-2
Further reading
Try all the models mentioned in this article
Admix gives you GPT-5, Claude, Gemini, and 350+ AI models in one app. Compare responses side by side. Free to start.
Start free on Admix