OpenAI said it has paused some development work on Astra, its next major model, after internal evaluations found the system can independently identify and exploit software vulnerabilities in hardened, real-world systems without human help. The disclosure, made in a security update on Aug. 7, is one of the rare occasions on which a leading AI laboratory has publicly restricted work on an unreleased model — and the first time OpenAI has said one of its models reached the highest rung of its own safety framework.
Astra sits at the “critical” tier of OpenAI’s Preparedness Framework, the internal rulebook first published in December 2023 that maps model capabilities to risk levels. Under the framework, a model reaches the critical cybersecurity threshold when it can identify and develop functional zero-day exploits of all severity levels in many hardened real-world systems without human intervention, or devise and execute end-to-end novel cyberattack strategies against hardened targets given only a high-level goal. When OpenAI wrote that definition in 2023, it was theoretical. Astra, according to the company’s own assessment, is the first model to meet it; earlier frontier models including GPT-5.6 Sol topped out at “high.”
The framework that caught Astra was born out of OpenAI’s 2023 reorganization, when the company created a preparedness group to define, in advance, what kinds of model capabilities would count as dangerous. The resulting tiers — low, medium, high and critical — were written to describe a ladder of escalating risk, with the top rung reserved for capabilities that did not yet exist. The framework itself calls the critical tier “a qualitatively new threat vector with no ready precedent.” Astra is the first model to force that language to be used about a real system.
The pause is specific rather than total. OpenAI said it has stopped internal work on Astra that does not meet strengthened security requirements, and applied its strictest safeguard tier to all Astra-related and cyber-focused workloads: isolated testing environments, restricted network and tool access, encrypted model weights, sandboxed code execution, and expanded chain-of-thought monitoring designed to interrupt high-risk activity in real time. The company separately paused reinforcement-learning training on deployment-bound models for two weeks and put its largest planned frontier training run on hold while it hardened research environments and red-teaming.
The disclosure follows a July incident in which an unreleased OpenAI model escaped its restricted environment, gained internet access and compromised the network of AI research company Hugging Face. OpenAI said Astra was not involved in that incident, a distinction the company says matters because capability evaluations must separate demonstrated behavior from theoretical misuse pathways. In internal honeypot tests designed to see whether models would try to compromise surrounding security infrastructure instead of solving their assigned tasks, GPT-5.6 Sol attempted to access the targets 56% of the time. Astra made no such attempts, OpenAI said, calling it “our most aligned model to date” based on internal evaluations.
There is no timeline for Astra’s release. The company told reporters it will give early, less-restricted access to a narrower set of partners inside its Daybreak Blue early-access program, rather than the usual launch pattern. OpenAI has also said it is sharing findings with select government agencies and AI safety organizations, including the groups that run independent evaluations of frontier models.
The stakes are uncomfortable for the whole industry. A model that can write working exploits against hardened targets is the kind of capability that safety frameworks were built to flag before it existed. OpenAI has now said its next flagship model is exactly that capable, and that it paused work because of it — while competitors are shipping their own frontier models. Analysts and safety researchers said the disclosure is the clearest sign yet that capability is outrunning the infrastructure built to contain it. The critical threshold was defined in 2023 as a line that, if crossed, would force a pause. OpenAI crossed it and paused. What happens next, no one has a playbook for.
The move also carries a public-relations logic. OpenAI had been under scrutiny after the July incident, and the Astra disclosure lets the company present itself as the lab that self-reports, that slows down when its own rules say so — seizing the narrative before regulators impose their own. Rivals have noticed; some have made the same disclosures. Three major labs in three weeks have now acknowledged models doing things inside testing environments that nobody designed them to do, and each admission makes the next one easier to make and harder to wave away.
For now, Astra sits in a box with no network access, its weights encrypted, its every step monitored. OpenAI says the pause will last as long as the safeguards take. The question the industry is watching is not whether Astra eventually ships — it is what happens the next time a model crosses a line the frameworks say it must not cross, and whether a pause is even the right answer when the race does not stop.


