The discovery comes days after rival OpenAI disclosed that an autonomous agent powered by its AI models went rogue.
SAN FRANCISCO — Artificial intelligence company Anthropic said on Thursday that its AI model Claude hacked into the systems of three companies during testing after a configuration error gave it internet access.
The revelation follows a similar incident last week in which rival OpenAI disclosed that an autonomous agent powered by its AI models went rogue during a security test and triggered a hack that compromised the infrastructure of AI firm Hugging Face.
According to Anthropic, a misconfiguration allowed Claude models to reach the internet from testing environments that were supposed to be isolated, leading to unauthorised access to three organisations’ systems. The company identified the incidents after reviewing 141,006 test sessions, a process it launched following OpenAI’s disclosure.
How the Breaches Occurred
Anthropic said the incidents involved three separate models: Claude Opus 4.7, Claude Mythos 5, and an internal research model. The earliest cases dated back to April and occurred in evaluation environments that lacked what the company described as standard safeguards.
The breaches took place during so-called “capture-the-flag” exercises, in which models are tasked with finding hidden information in simulated networks. The company said its prompts told the models they had no internet access, but a misunderstanding with its evaluation partner Irregular left the systems connected to the public internet.
“Claude compromised the impacted organisations’ infrastructure using basic techniques, such as exploiting weak passwords and unauthenticated endpoints,” Anthropic said.
Response and Notification
Anthropic said it began reviewing evaluation transcripts on July 23 and suspended all cyber evaluations the same day after finding evidence that Claude may have accessed the internet. It identified all three incidents by July 24 and notified the affected organisations on July 27.
Two of the organisations were unaware of the activity before being contacted, Anthropic said, adding that it was still trying to reach the third.
Broader Implications
The breaches signal that AI’s expanding capabilities are already fuelling the security threat experts long feared, and even top developers can be caught off guard by flaws their models can exploit.
Anthropic said the findings underscore the need for stronger controls in both internal and third-party testing environments as AI models become increasingly capable of carrying out real-world cyber activities.

Leave a Reply