AI

How Anthropic AI was able to hack three companies

31 July 2026
3 minutes
Anthropic reveals Claude AI models were able to breach three companies during testing, days after OpenAI found an AI breach that hacked Hugging Face.
Image credit: Anthropic (CC BY 2.0 / Attribution 2.0 Generic / Deed)
Image credit: Anthropic (CC BY 2.0 / Attribution 2.0 Generic / Deed)

Anthropic said on Thursday that some of its Claude AI models had hacked into the systems of three companies during tests.

The company disclosed this just days after OpenAI revealed that one of its AI agents went rogue during an evaluation and attacked AI company Hugging Face. Anthropic stated online that after the OpenAI incident, it launched a “large-scale retrospective review” of its cybersecurity evaluations.

As a result, it found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment. The model then gained unauthorised access to the real systems of three different organisations.

“In all cases, Anthropic’s evaluation prompt specified to Claude that its environment was a simulation and that it had no internet access. Due to a misunderstanding between us and our evaluation partner, this was not the case, and internet access was available,” Anthropic explained via its online statement. “Because of this, when Claude’s search led it to real systems on the open internet, it treated them as part of the exercise.”

The incidents, which Anthropic referred to as an “operational failure”, were tasked with ‘capture-the-flag’ challenges. These are fictional scenarios where each model had to find hidden information in simulated networks – but ended up leading to companies being hacked.

Anthropic added: “Claude compromised the impacted organisations’ infrastructure using basic techniques, such as exploiting weak passwords and unauthenticated endpoints.”

This mistake contrasts with the OpenAI incident, in which its AI agent exploited a vulnerability to gain internet access during testing. Anthropic’s incident involved its Opus 4.7, Mythos 5 and an internal research test model, with the earliest incidents dating back to April this year.

The company acknowledged a need for stronger controls in both its own internal and third-party testing environments, particularly as AI models become more capable of carrying out real-world cyber activities.

“Ultimately, many factors contributed to these incidents, but, consistent with a blameless postmortem culture, we’re approaching the fixes as if the responsibility were ours alone,” Anthropic said. “This begins with ensuring every part of our evaluation pipeline is secure, including the manner in which we integrate with external partners.”

It added: “Moving forward, it will include expanding our continuous monitoring of evaluation transcripts for unexpected behaviour, improving our investigation tooling and conducting more rigorous assurance work with the vendors we rely on … These facts give us cautious optimism that with tighter monitoring and controls around evaluation infrastructure, as well as continued investment in alignment, this type of risk can be overcome.”

These instances also come as the US government cracks down on new AI models, having suspended Anthropic’s own Fable 5 and Mythos 5 models to ‘foreign nationals’ earlier this year, citing national security concerns.

Related stories

Anthropic takes Project Glasswing to critical infrastructure operators in 15 countries

Inside the $35bn deal: Apollo and Blackstone’s chip-backed SPV for Anthropic signals a new financing era

Why Amazon is investing $25bn into Anthropic in a major AI deal

Datacloud USA

01 September 2026