Next upHack for Humanity: San Francisco (powered by Google Gemini)
News

OpenAI pauses Astra model after it may have crossed 'Critical' cyber threshold

OpenAI said it cannot rule out that its upcoming Astra model has reached the 'Critical' cybersecurity threshold, the first such designation for any of its models.

D
Aug 7, 2026 · 1 min read

OpenAI said on Aug. 7, 2026 that it cannot rule out its upcoming Astra model has crossed the “Critical” cybersecurity threshold under its Preparedness Framework — the first such designation for any of its models. The company is pausing internal work on Astra that lacks enhanced safeguards while it runs further tests.

Under OpenAI’s Preparedness Framework, the risk-evaluation system it uses to grade frontier models, “Critical” is the top cybersecurity tier. It applies to a model that can independently find and build working zero-day exploits against hardened real-world systems, or plan and run an end-to-end cyberattack from a high-level goal alone, without a person in the loop.

In its disclosure on responding to next-frontier critical cyber capabilities, OpenAI said it is moving Astra into isolated testing environments, adding universal monitoring of the model’s agentic actions, and working with government agencies and outside AI safety organizations to probe what it can do before any release.

OpenAI has not confirmed that Astra has definitively passed the threshold, and said testing is continuing. It also said Astra played no part in an earlier security-evaluation incident involving Hugging Face. The Critical label is precautionary: the company says it cannot yet rule the capability out, not that it has demonstrated it against live systems.

The decision is an early test of whether a frontier lab will hold back a finished model on safety grounds rather than ship it. OpenAI framed the pause as temporary, pending results from the added safeguards and third-party testing.

More news