OpenAI has spent years telling us how powerful its AI models are. Now, once again, it is warning us that one of its own creations may be getting dangerously capable.
The company says an upcoming model called Astra has performed well enough in recent cybersecurity and agentic coding evaluations that OpenAI can no longer rule out the possibility that it has reached the Critical cybersecurity threshold under its Preparedness Framework. That is a pretty serious statement, folsk.
Previous OpenAI models, including GPT-5.6-Sol, were classified at the High level for cybersecurity capability. Astra could potentially move beyond that.
Under OpenAI’s framework, a model reaches the Critical cybersecurity level if it can independently identify and develop functional zero-day exploits against many hardened real-world systems, including critical targets. A model could also qualify if it can devise and execute new end-to-end cyberattack strategies after being given little more than a high-level goal.
OpenAI is not saying Astra has definitely reached that point. The company says its preliminary results are strong enough that it cannot rule it out.
Additionally, the company says it has already tightened internal security around Astra. That includes isolated testing environments, restricted network and tool access, stronger protection and encryption for model weights, more monitoring, and sandboxed execution.
OpenAI has also paused internal Astra-related work that does not yet meet those higher security standards.
Additionally, the company says it has deployed universal monitoring across agentic uses of Astra, including during training and evaluation. These monitoring systems inspect the model’s chain of thought for risky behavior and can trigger reviews or interrupt activity deemed dangerous.
OpenAI also plans to work with government agencies, AI safety organizations, and outside testing partners to further evaluate the model. Third parties performing higher-risk testing will receive recommended security controls from OpenAI.
One detail the company went out of its way to clarify is that Astra was not involved in the exploitation of Hugging Face.
OpenAI says the ultimate goal is to put advanced cyber capabilities in the hands of defenders, allowing them to identify vulnerabilities before attackers can exploit them. That is obviously the outcome everyone should want.
And to be clear, I am happy OpenAI is taking safety seriously. I would much rather see the company raise alarms internally, strengthen safeguards, and talk publicly about these risks than pretend everything is fine. Still, this routine is starting to get tiring.
OpenAI builds increasingly powerful AI systems, tells the world how impressive they are, and then explains that those same systems may also be dangerous enough to require extraordinary precautions. At some point, it becomes difficult to ignore the strange contradiction of a company repeatedly warning us about threats created by products it is actively racing to develop.
Maybe that is simply the reality of frontier AI development. Powerful tools can be useful and dangerous at the same time.
But if every major leap in AI capability is followed by another warning about cyberattacks, biological threats, or some other potentially catastrophic capability, the public may eventually start asking a more uncomfortable question…
Just because these models can be built, should companies keep pushing them this far?
Support independent tech journalism
NERDS.xyz is independently owned and operated. If you enjoy my coverage of Linux, AI, hardware, cybersecurity, and tech culture, consider supporting the site on Ko-fi.
Support NERDS.xyz