OpenAI Halts New Model Rollout Due to Security Worries
OpenAIis pausing internal activity related to a new artificial intelligence model due to security concerns
The AI startupannouncedthis decision on its blog Friday (Aug. 7) after an in-house evaluation of its Astra model caused the company to realize it could not “rule out critical cyber capabilities” under its Preparedness Framework, which outlines how the company responds to possible AI risks
Under theframework, an AI model reaches the “critical cybersecurity threshold if it can identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention,” the blog post said
Models also reach the threshold if they can “devise and execute end-to-end novel strategies for cyberattacks against hardened targets given only a high level desired goal,” OpenAI added
According to the blog post, OpenAI’s recent evaluations found that Astra made significant coding and cybersecurity progress that pushed the model closer to the “critical” threshold
With that in mind, OpenAI says it is “pausing internal activities involving Astra that do not yet meet these strengthened security control requirements,” and that it has “implemented universal monitoring for risky actions and misalignment across all agentic applications of Astra, including training and evaluation.”
The company also said it is working with government agencies and “select” AI safety groups to test the model’s capabilities
The news follows a series of cybersecurity incidents involving advanced AI models. Last month, a pair of OpenAI modelsbroke loosefrom their testing environment and hacked open-
Two days later, Anthropic said that a review of its evaluation history, prompted by OpenAI’s findings,found three incidentssince April in which its Claude models had accessed the systems of three different organizations
AndMetasaid last week that one of its AI modelshacked another companyduring cybersecurity testing. The Facebook owner said an unintentional misconfiguration by a company that performs cybersecurity evaluations, Irregular, provided the model access to the internet during testing
Areporton the Astra incident by The Wall Street Journal includes comments fromJeffrey Ladish, executive director ofPalisade Research, a nonprofit AI lab that examines AI capabilities to determine risks.He said OpenAI should have halted work on Astra after finding out about the Hugging Face hack.
“It’s definitely late,” he said. “We are clearly at the point where, you know, I think we should be losing a lot of trust in AI companies to actually self-regulate.”
For all PYMNTS AI coverage, subscribe to the dailyAI Newsletter

