
OpenAI Is Holding Back Its Next Model, Astra, Over Cyber Risk
Quick verdict
OpenAI says its next model, Astra, got good enough at agentic coding and cybersecurity that it cannot rule out a "Critical" cyber capability level under its own Preparedness Framework. So it is doing something labs almost never do out loud: pausing parts of the program, tightening network and tool access, hardening how the model weights are stored, and adding monitoring before it ships Astra more widely. The stated goal is still to get the model to defenders, just under stricter controls. If you were waiting on Astra, expect it later and more locked down.
What actually shipped
Nothing shipped, and that is the news. OpenAI put out a statement saying evaluations of the upcoming Astra model show "significant advancements in agentic coding and cybersecurity," strong enough that the lab cannot rule out that Astra reaches the Critical capability level on the cyber track of its Preparedness Framework. Critical is the top rung. It is the level meant to describe a model that could meaningfully help someone carry out a serious cyberattack.
The Preparedness Framework is OpenAI's own scale for measuring dangerous capabilities across categories like cyber, bio, and self-improvement. Each category has thresholds, and once a model crosses one, the framework says the lab has to put safeguards in place before release. This is the first time OpenAI has publicly treated one of its models as a possible Critical cyber case, which is why people paid attention on an otherwise slow news day.
In response, OpenAI listed the steps it is taking. It is pausing internal activities that do not meet its strengthened controls. It is tightening network and tool access so the model has fewer ways to reach outside systems. It is hardening weight security, meaning the raw model files are harder to steal. And it is expanding monitoring during training and evaluation. Sam Altman and Greg Brockman both framed the move as caution rather than alarm, and OpenAI's Boaz Barak, who works on the preparedness side, pointed to the same posture: slow down, add guardrails, still aim to get the tool into the hands of defenders.
| Item | Detail |
|---|---|
| Model | Astra (upcoming, not yet released) |
| Trigger | Strong agentic coding and cybersecurity eval results |
| Classification | Cannot rule out "Critical" cyber under the Preparedness Framework |
| Response | Pause some internal work, tighten network/tool access, harden weight security, expand monitoring |
| Stated goal | Still ship to defenders, under stricter controls |
Why it matters
Set aside the marketing question of whether Astra is "safe." The more useful read is what this says about where the frontier is. A model good enough at offensive and defensive cyber work that its own maker will not rule out the top risk tier is a real capability jump, not a spec-sheet number. Agentic coding and cybersecurity are close cousins, so a model that can drive a long chain of tool calls to fix a codebase can, in principle, drive a similar chain to break into one.
It also sets a reference point for the rest of the industry. OpenAI is not the only lab running a preparedness-style framework, and every serious lab now has to answer the same question: at what point do you delay your own launch over cyber risk? Doing it in public, before a release rather than after an incident, is a different posture than we have usually seen. It reads partly as genuine caution and partly as a signal to regulators and rivals that OpenAI takes this seriously. Both can be true at once.
For you as a user, the near-term effect is simple. Astra arrives later than a no-caution timeline would suggest, and when it does, expect tighter access controls, more monitoring, and probably a slower path to the full-power version through the API. That is not a bad trade. It is worth remembering why the caution exists, though. OpenAI has already had a model reach outside its sandbox during testing once this year, so the concern here is not hypothetical.
Video: how OpenAI's Preparedness Framework works
Background on the framework OpenAI is using to classify Astra, and why the cyber category matters.
FAQ
What is Astra?
Astra is an upcoming OpenAI model that has not been released yet. OpenAI has said its evaluations show strong agentic coding and cybersecurity performance, enough that the lab is treating it as a possible "Critical" cyber case under its safety framework.
What does "Critical" cyber capability mean?
It is the top risk tier on the cyber track of OpenAI's Preparedness Framework. It describes a model that could meaningfully help carry out a serious cyberattack. Crossing that line, or being unable to rule it out, obligates OpenAI to add safeguards before a wider release.
Does this mean Astra is dangerous?
It means OpenAI cannot yet rule out that it is dangerous in the cyber category, which is why it is slowing down and adding controls. The lab still plans to release it, with the stated aim of getting it to defenders under stricter conditions.
When will Astra come out?
OpenAI has not given a date. The whole point of this announcement is that the timeline is now gated on strengthened controls, so a later, more locked-down release is the safe assumption.
Sources
- OpenAI - statement on Astra's cyber classification and controls
- @gdb - Greg Brockman on the decision to slow down
- @sama - Sam Altman on the Astra posture
- @boazbaraktcs - OpenAI preparedness researcher on the classification
- @kimmonismus - Axios summary of the announcement
Further reading
Try all the models mentioned in this article
Admix gives you GPT-5, Claude, Gemini, and 350+ AI models in one app. Compare responses side by side. Free to start.
Start free on Admix