
ZDNET’s key takeaways
- OpenAI released its previously paused Astra model.
- Researchers are worried the powerful new model could endanger security at scale.
- OpenAI said Astra triggered new internal safety guardrails.
OpenAI’s latest model, named Astra, is finally loose — sort of.
The company halted development on Astra after July’s now-infamous Hugging Face incident — in which an OpenAI agent hacked the developer site, and several others — as it reevaluated internal security standards. As of today, the model is live for a select group.
OpenAI emphasized that Astra sets a new standard for computer use, taking “about 47% less time per task than GPT-5.6 Sol, scoring 72.6% at roughly 40 minutes per task.” In addition to its engineering and browser use capabilities, the company mentioned the model’s cybersecurity prowess and ability to advance science, citing new discoveries in math and reasoning scores for biology.
According to OpenAI, Astra achieves what we’d expect from increasingly capable frontier models: better, more trustworthy automation of routine tasks like filling out forms, organizing calendars, conducting research, and analyzing data.
Also: How OpenAI’s agent escaped: Sprung by humans in a series of preventable events
Among other notable benchmark test scores, Astra scored 59.3% on Agent’s Last Exam, which measures agents’ ability to perform human-level professional work. That score beats Anthropic’s Fable 5 score of 48.7% and Claude Opus 5’s score of 52.7%. In a briefing for the model, OpenAI VP of research Aidan Clark added that Astra is the first model for which training was significantly supported by other models.
(Disclosure: Ziff Davis, ZDNET’s parent company, filed an April 2025 lawsuit against OpenAI, alleging it infringed Ziff Davis copyrights in training and operating its AI systems.)
“If we fast forward a couple of years and we look back and say, ‘When was it really that AGI was created?’ I think it’s going to be about this time, and […] about this model,” OpenAI president Greg Brockman said in the briefing, referring to artificial general intelligence (AGI). But when pressed to elaborate, he hedged without fully defining Astra as AGI, a term that still lacks an agreed-upon, industry-wide definition.
Security concerns
On Tuesday, OpenAI said in a post that the company had “strengthened and tested protections against cyber misuse” while Astra was paused — admitting that, in previous testing, the model had “discovered previously unknown vulnerabilities and turned them into working exploit chains.”
Since the Hugging Face incident — or even further back, since Anthropic’s Mythos 5 release — safety researchers have been much more concerned about the accelerating pace of frontier models and the accompanying cybersecurity capabilities. Additionally, according to a source who spoke with The Information, Astra’s specific model architecture obscures chain of thought more than typical models, making it hard for researchers to interpret the model’s internal decision making. Given how capable this model supposedly is, safety experts are concerned.
Ryan Greenblatt, chief scientist at Redwood Research, was one of a few researchers OpenAI allowed to dig into the Hugging Face incident. In a post on X responding to OpenAI’s Tuesday blog post, he said this architecture choice indicated a “race to the bottom” that “could be catastrophic for our ability to oversee/monitor AIs.”
“In our investigation of the OpenAI/Hugging Face incident, we were heavily reliant on chain-of-thought,” Greenblatt explained. “If the AIs we were investigating had instead been reasoning in latent space, this would have greatly undermined our investigation. Getting a good understanding of the behavior of this many agents was tricky enough even with the use of chain-of-thought!”
That fear of losing insight into a model’s “thinking” has been brewing for a while, though. In July 2025, researchers from nearly every major AI lab, including OpenAI, published a paper expressing concern that this “fragile” yet vital method of observing AIs would be deprioritized in favor of quicker, more capable development over time.
Still, OpenAI said in the release that Astra is its “most aligned model” — meaning it adheres closely to the guidelines developers create for it — “with substantial improvements in understanding user intent and model behavior.”
Also: OpenAI’s attack agent did exactly what it was told – just more relentlessly than expected
“As one way that we test this, we built a new evaluation informed by the Hugging Face incident that evaluates whether a model facing a difficult or impossible task will go beyond its intended scope. Compared to GPT-5.6 Sol, which, without production safeguards, went beyond the authorized target 48.2% of the time, GPT-6 Astra did this in 0% of cases,” the release said.
That’s complex to interpret, though. As ZDNET contributor and security expert David Berlind said, the Hugging Face hack was an example of agents doing exactly what they were designed to, which is to find a way to complete their assigned task. It just happened much faster than we anticipated. Alignment may not solve this.
“Progress in intelligence does not guarantee progress in alignment,” chief scientist Jakub Pachocki said in the briefing. He emphasized that OpenAI would scale down development if it didn’t feel confident in the safety of its releases. But when pressed on how the company could specifically create more safety initiatives in the AI industry, he reiterated that these priorities were being handled internally.
Brockman said that Astra was tested in collaboration with the US government, though he did not specify whether this was within the voluntary framework the government launched earlier this summer.
Availability
For now, OpenAI is following Anthropic’s approach for high-capability cybersecurity models. Starting today, Astra is only available to enterprise customers with access to Daybreak, its somewhat answer to Anthropic’s exclusive Glasswing coalition of cybersecurity users. Similarly, Anthropic recently made its Mythos 5 model available in Claude Security, rather than making it generally available (though its latest Fable upgrade is).
Also: AI a ‘force multiplier’ for low-skilled threat actors: 4 ways organizations should respond
“Over the coming days, [Astra] will become available to all Plus, Pro, Business, and Enterprise users, as well as through the OpenAI API and AWS,” the company said.

