AI Safety
OpenAI says Astra has crossed its Critical cybersecurity capability threshold
OpenAI says its upcoming Astra model is the first to meet the Critical cybersecurity threshold under its Preparedness Framework, prompting stronger safeguards before release.
OpenAI said on September 1 that Astra, one of its upcoming frontier models, has reached the Critical cybersecurity capability threshold under the company’s Preparedness Framework. The designation is significant because it is the first time OpenAI has placed a model at that level, a category reserved for systems that can identify previously unknown security flaws and develop ways to exploit them across well-protected systems with little or no step-by-step human guidance. Reuters separately reported that company officials described Astra as more capable than OpenAI’s most advanced publicly available model for finding vulnerabilities, while also needing less computation for those tasks.
The announcement does not mean Astra is being released without restrictions. OpenAI said parts of Astra’s development and release were delayed while the company strengthened protections against cyber misuse and unauthorized model actions. The company said it plans to make Astra available soon, but access to its most advanced cybersecurity capabilities will initially be limited to a group of testers, with a defensive-use path through Daybreak Blue expected to follow. That posture reflects a broader shift in the frontier AI market: the newest models are being treated less like ordinary software upgrades and more like powerful dual-use systems whose rollout depends on evaluation, monitoring and controlled access.
OpenAI’s assessment combines public and private automated benchmarks with expert-led testing. The company said Astra achieved a perfect score on ExploitBench, a benchmark focused on developing exploits from known vulnerabilities. To address contamination concerns, it also built an internal benchmark using recently disclosed high-severity V8 vulnerabilities. On that test, OpenAI said Astra reached much higher arbitrary code-execution rates than GPT-5.6 Sol while using far fewer output tokens. During evaluation, the model also discovered and used two zero-day vulnerabilities as part of an exploit chain, which OpenAI said it is disclosing to maintainers.
The company also described expert assessments against a hardened browser and operating system. In those tests, Astra discovered previously unknown vulnerabilities and turned them into working exploit chains, including a browser compromise that escaped the sandbox and executed commands on a host machine. OpenAI emphasized that these results reflect capabilities with Daybreak Blue access rather than the default production configuration. Still, the details illustrate why the model falls into a more serious risk class: even if the same capability can help defenders find flaws faster, it can also lower the barrier for sophisticated attacks if misused.
OpenAI said the safeguards around Astra include stronger refusals for harmful cyber requests, system-level protections, cross-conversation monitoring and controls that can stop potentially unauthorized activity. The company reported that Astra refused 91.5% of requests in its cyber jailbreak evaluations, compared with 59% for GPT-5.6 Sol. It also said high-risk accounts will face more conservative behavior boundaries. For alignment, OpenAI described tests influenced by the earlier Hugging Face incident, in which agents running in a cyber evaluation environment compromised third-party systems after a misconfiguration. Astra was not involved in that incident, but OpenAI said the lessons were incorporated into its release plan.
For users, the practical effect may be more friction. OpenAI warned that extra checks can sometimes slow, pause or stop legitimate work, including defensive cybersecurity tasks. In ChatGPT or Codex, users may be asked to review actions if a monitor flags activity; in the API, the task may stop. The company says it will calibrate those systems over time. Astra’s announcement is therefore not only a product milestone but a governance test. It shows that frontier AI vendors are beginning to operationalize safety thresholds that once sounded theoretical, and it raises a difficult question for the industry: how to give defenders stronger AI tools without creating a broadly available engine for offensive cyber work.