OpenAI today said it is “pausing” activities involving its upcoming AI model Astra, because its cyber capabilities are potentially too dangerous. OpenAI says its newest internal evaluations show “significant advancements in agentic coding and cybersecurity,” and it cannot rule out “critical cyber capabilities.” Prior OpenAI models, including GPT–5.6 Sol, were labeled as “High.”
Astra triggers stricter guidelines in OpenAI’s “Preparedness Framework.” The guidelines call for caution when developing frontier AI capabilities that create risks of severe harm, and the cybersecurity portion of the framework says OpenAI will implement extra safeguards for models that “create new risks of scaled cyberattacks and vulnerability exploitation.”
The “Critical” threshold Astra may have hit is defined by an ability to identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention, or devise and execute end-to-end novel strategies for cyberattacks.
OpenAI says it is increasing its safeguards and security controls before deploying Astra, including limiting work on the model until new safeguards are in place. The company plans to use isolated testing environments with restricted network and tool access, along with adding sandboxed execution and more monitoring capabilities. OpenAI says it will work with relevant government agencies and AI safety organizations to test Astra.
“We’re committed to working alongside governments, safety institutes, and civil society to ensure that the frontier capabilities of models like Astra, and those that follow, are deployed responsibly and broadly for the benefit of all humanity,” writes OpenAI.
Astra wasn’t formally announced, but OpenAI shared details on its next major model in a recent post outlining its mathematical advancements. Astra solved 10 open problems in math and theoretical computer science for around $2,000 (in Sol API rates).
Advancements in AI are changing cybersecurity for major tech companies like Apple by unearthing an unprecedented number of bugs. Apple recently limited its bug bounty program submissions because it is having trouble handling the volume.
Models like Claude Mythos are able to suss out critical vulnerabilities, and Apple is one of Anthropic’s Mythos partners. Mythos is limited to select companies because in addition to finding vulnerabilities, it has the potential to exploit them.
OpenAI made headlines in July because GPT–5.6 Sol and a “more capable pre-release model” (not Astra) autonomously hacked Hugging Face during internal benchmark testing. Anthropic found Claude had done something similar. Meta this week said it too had an AI model hack another company during a cybersecurity evaluation.


