August 8, 2026
OpenAI's Astra Model Hits 'Critical' Cybersecurity Threshold, Prompting Safety Pause

OpenAI's Astra Model Hits 'Critical' Cybersecurity Threshold, Prompting Safety Pause

Posted 1 hour ago by
OpenAI has paused some internal activities involving its next-generation Astra model after internal evaluations indicated the company cannot rule out that the system has reached the "Critical" cybersecurity capability level under its Preparedness Framework.

OpenAI's Astra Model Hits 'Critical' Cybersecurity Threshold, Prompting Safety Pause

Recent internal evaluations found Astra has made significant advancements in agentic coding and cybersecurity, with performance strong enough that OpenAI cannot rule out the "Critical" capability level. In a blog post published Friday, the company explained that this threshold is reached when a model can independently identify and develop functional zero-day exploits of all severity levels against many hardened real-world critical systems, or devise and execute end-to-end novel strategies for cyberattacks against hardened targets given only a high-level desired goal.


Astra was not involved in exploiting Hugging Face during internal testing, OpenAI said. To reduce the risks posed by increasingly capable models, the company is implementing stricter security controls, including isolated testing environments, restricted network and tool access, enhanced model weight protections and encryption, additional monitoring and detection capabilities, and sandboxed execution. It is also pausing Astra activities that do not meet those strengthened security requirements and has implemented universal monitoring for risky actions and misalignment across all agentic applications of the model, including training and evaluation.

OpenAI said it is sharing the findings because it believes "it's important to be transparent with the public and the safety and security communities about this potential shift in capabilities." The disclosure follows several recent announcements involving increasingly capable frontier AI models. Earlier this summer, the U.S. government ordered Anthropic to suspend its Claude Fable 5 and Mythos 5 models over their cyber capabilities. The restrictions were later lifted, allowing Anthropic to restore access to both models after implementing additional safeguards.

Looking ahead, OpenAI plans to work with relevant government agencies and select AI safety organizations to further evaluate the model. The company will also provide recommended security controls to third-party testing partners for running higher-risk evaluations and workloads safely. "We're committed to working alongside governments, safety institutes, and civil society to ensure that the frontier capabilities of models like Astra, and those that follow, are deployed responsibly and broadly for the benefit of all humanity."
Add Comment
Would you like to be notified when someone replies or adds a new comment?
Yes (All Threads)
Yes (This Thread Only)
No
iClarified Icon
Notifications
Would you like to be notified when we post a new Apple news article or tutorial?
Yes
No
Comments
You must login or register to add a comment...
Recent. Read the latest Apple News.
RECENT
Tutorials. Help is here.
TUTORIALS
Where to Download macOS Tahoe
Where to Download macOS Sequoia
Where to Download macOS Sonoma
Deals. Save on Apple devices and accessories.
DEALS