The most powerful AI models keep going awry,according to the companies building them
The disclosures come as frontier AI models get more powerful and more capable of acting autonomously. They also highlight a growing challenge for the companies building them —the systems designed to test increasingly capable models can have weaknesses of their own
Over the past few weeks, multiple frontier AI models have accessed real systems during cybersecurity testing
Researchers on Friday said China’s popular new Kimi K3 model, made by Moonshot AI, circumvented restrictions in its test environment. Anthropic and Meta also said recently that their own latest models have done things they aren’t supposed to. OpenAI kicked it all off last month when its models went to great — and worrisome — lengths to hack into another company
Amid heightened concern, OpenAI said Friday that its as-yet-unreleased model, Astra, is demonstrating cyber capabilities so advanced that the company can no longer rule out assigning it the highest-risk designation
As a result, OpenAI said it is pausing work on Astra that doesn’t meet new safeguards, and said it will work with government agencies and AI safety groups to further test the model
“astra is a powerful model and we are working to make it generally available,” OpenAI CEO Sam Altman wrote on X on Friday. “given its cyber capabilities, we need a little big longer to do do this safely.”
Explore BI Games
Take a smarter break in your day – and see how far you get
Play now
The security lapses during testing are also amping up pressure on the industry and the White House to find ways to regulate AI systems across the board
There is, of course, also a not small contingent of observers out there who suspect these announcements are just elaborate marketing to hype new models and show antsy investors progress toward the ultimate goal: artificial general intelligence
You can judge for yourself. Here’s how Anthropic, OpenAI, Meta, and researchers testing China’s Kimi K3 say the latest models have gone off the rails
OpenAI’s models find a way
OpenAI researchers revealed eyebrow-raising new details this week about a recent incident in which AI agents escaped the company’s internal testing environment and eventually hacked into Hugging Face’s systems in search of answers
The company said the agents created their own internal message board — even after OpenAI tried to shut it down. One agent reacted to discovering unexpected access by thinking, “Holy shit reader is ADMIN?” Another wrote, “We can communicate now!”
OpenAI researcher Eric Wallace said the agents realized they could accomplish more by working together. “They start to launch these collective attacks on third-party and internal services,” he said. The agents eventually turned their attention to Hugging Face
OpenAI has called the Hugging Face attack an “unprecedented cyber incident.” The episode has taken on new significance as OpenAI tests Astra, its unreleased model that the company says may have reached its highest cybersecurity risk level
“Our latest internal evaluations of Astra, one of our upcoming models, over the past few days indicate significant advancements in agentic coding and cybersecurity,” OpenAI said
OpenAI is imposing stricter security controls on Astra, including sandboxed execution, restricted network access, and stronger protections around model weights. It has also paused internal Astra work that does not meet the heightened requirements
Anthropic’s Claude can do it too
Anthropic said it reviewed more than 141,000 AI tests and found three cases, dating back to April, in which Claude models accessed live systems belonging to real organizations without authorization
“In all cases, Anthropic’s evaluation prompt specified to Claude that its environment was a simulation and that it had no internet access,” the company said in a blog post. “Due to a misunderstanding between us and our evaluation partner, this was not the case, and internet access was available,” it added, referring to Irregular, an AI security startup
The incidents involved Claude Opus 4.7, Mythos 5, and an internal research model. Anthropic said it contacted the organizations involved and that two had not known they had been hacked
The episodes raised questions about whether the bigger failure was the models themselves or the environments containing them. Anthropic said it was discussing a third-party review of the incidents
Muse Spark exploits a third-party vulnerability
Meta also disclosed a cybersecurity-testing mishap this week. The company said its Muse Spark model “exploited a security vulnerability in a third-party service” during an evaluation
A Meta spokesperson told Business Insider that the incident stemmed from a misconfiguration by Irregular, which allowed the model to access the internet during testing
Meta said Irregular notified it about the incident and that the company is investigating. It plans to release more details once that review is complete
Kimi K3 escapes its sandbox
Researchers at the cybersecurity firm Frontier Security said Kimi K3, a popular new model from the Chinese AI company Moonshot, also bypassed restrictions in a cybersecurity testing environment
Researchers said the sandbox — a controlled, isolated environment where an AI model can run code — had been improperly configured
The environment blocked certain web traffic, but Kimi bypassed those restrictions using command-line tools
The researchers said the incident suggested that some cybersecurity evaluations contain weaknesses that capable models can exploit
“This suggests that some of the evaluations on cybersecurity that the community uses are susceptible to security vulnerabilities and allow models to cheat,” they wrote in a report

