Anthropic Reports Claude Breached Three Organizations During Cybersecurity Trials

On Thursday, Anthropic revealed that its AI models gained unauthorized access to the systems of three unnamed organizations during cybersecurity evaluations. The company indicated that Claude reached the internet “from within or while interacting” with a third-party testing environment. This announcement follows OpenAI’s disclosure over a week ago that one of its AI agents compromised Hugging Face during a different cybersecurity test.
The revelation was made after Anthropic initiated “a large-scale retrospective review of our own cybersecurity evaluations” in response to the OpenAI incident, as stated in a blog post released by the company. Anthropic reported identifying 141,006 tests where Claude could have potentially accessed the internet. Subsequently, they found that three different Claude models had accessed the internet in evaluations conducted by the third-party AI testing firm Irregular, leading to hacks into the production systems of three organizations.
Anthropic identified the models involved as Opus 4.7, Mythos 5, and an internal research test model. The earliest occurrences took place in April, indicating they likely went unnoticed by the public for several months. Similar to the OpenAI case, Anthropic had intentionally disabled safeguards intended to limit the AI models’ capabilities and prevent misuse. This means these were not the versions available to the public.
“In all three incidents, Claude had been assigned a capture-the-flag challenge, which is one method we use to evaluate a model’s cybersecurity skills,” stated Anthropic in its blog post. The company noted that in each case, “Anthropic’s evaluation prompt instructed Claude that its environment was a simulation and that it had no internet access.” The oversight was attributed to a “miscommunication” between Anthropic and Irregular.
Despite Claude being prohibited from accessing the internet, Anthropic explained that Irregular had incorrectly set up the machines used to test Claude, inadvertently granting the AI models internet access. “Neither we nor our evaluation partner realized this misconfiguration until it was detected through our additional evaluation monitoring last week,” Anthropic mentioned in the blog post.
“We now have evidence that both of the two major AI laboratories have not only failed to contain their agents but also to recognize their jailbreaks in real time,” remarked Jake Williams, vice president of research and development at Hunter Strategy. “It’s evident that immediate regulation and government oversight for AI testing is essential.”
Both Irregular and Anthropic did not respond immediately to requests for comments.
In contrast to the OpenAI situation, Anthropic clarified that Claude did not uncover or exploit any sophisticated vulnerabilities. Instead, it used basic methods, “like exploiting weak passwords and unauthenticated endpoints.”
OpenAI noted that its AI agent reached the internet by exploiting a zero-day vulnerability. However, it accessed the systems of various third-party organizations using the same range of common cybersecurity weaknesses as Anthropic’s models, specifically finding credentials exposed online.
Anthropic admitted that enhanced “defense-in-depth” measures implemented by both the AI lab and its testing partner could have mitigated or even prevented these incidents, mirroring OpenAI’s response to growing criticism following its own occurrence.
“I don’t see how any of these AI labs can dismiss this as ‘just a routine occurrence,’” Williams asserts. “It’s not—it’s negligence.”
The AI lab emphasized that the models were directed to believe they had no access to the open internet, and for the most part, Claude misinterpreted the organizations it accessed as part of the testing environment. In other words, the models largely failed to recognize that they had escaped containment in the first place.
