toolcall.
PolicyAug 7, 2026, 20:24 UTC

OpenAI says Astra may reach critical cyber capability

The company says internal evaluations were strong enough that it cannot rule out the highest cyber-risk tier in its Preparedness Framework.

OpenAI says internal evaluations of Astra, an upcoming model, showed enough progress in agentic coding and cybersecurity that the company cannot rule out Critical cyber capabilities under its Preparedness Framework.

In OpenAI’s framework, that threshold means a model may be able to find and develop functional zero-day exploits across many hardened real-world critical systems without human help, or carry out novel end-to-end cyberattack strategies against hardened targets from a high-level goal. OpenAI says Astra has not been fully benchmarked yet, but its preliminary results are strong enough to trigger the stricter posture.

The company says it is increasing robustness testing for safeguards and security controls, adding isolated testing environments, restricting network and tool access, strengthening model-weight protections and encryption, and expanding monitoring. OpenAI is also pausing internal Astra activity that does not yet meet those controls, and says it will work with government agencies and selected AI safety organizations to test the model.

One important boundary: OpenAI says Astra was not involved in exploiting Hugging Face. The point of the disclosure is broader: frontier models are moving from coding assistance toward cyber operations, and labs are beginning to treat some unreleased systems as potential critical security assets before they ship.

Sources

Mentioned

ai-securitycybersecurityfrontier-modelsopenaisafety