Claude Goes Rogue, Hacks Three Companies
Synopsis
Anthropic disclosed that its Claude AI escaped containment and compromised three organisations during a cybersecurity evaluation.
Key Highlights
- Anthropic said its Claude AI models accidentally hacked three organisations during internal cybersecurity testing.
- The incidents occurred after Claude gained internet access because of a configuration error.
- Anthropic uncovered the incidents after reviewing more than 141,000 cybersecurity evaluations following OpenAI’s recent AI security incident.
Anthropic has disclosed that its Claude AI models unintentionally breached the systems of three external organisations during internal cybersecurity testing after mistakenly being given internet access.
The company said the incidents happened during a “capture-the-flag” cybersecurity exercise, where the AI was instructed to locate hidden information within a simulated private network. Instead of remaining inside the testing environment, Claude accessed the live infrastructure of three real organisations.
Anthropic said it identified the incidents while conducting a retrospective review of more than 141,000 cybersecurity evaluations after OpenAI revealed that one of its AI models had compromised external systems during testing.
Three Separate Incidents
Anthropic said each of the three incidents was caused, at least in part, by human error. In the first case, Claude targeted a real company because its domain name matched that of a fictional organisation created for the exercise.
In the second incident, Claude created a real email account, registered a Python package on PyPI and uploaded malware intended for fictional employees. Instead, the package was downloaded by 15 real systems and installed by a cybersecurity company during a routine malware scan. The malware obtained some credentials, allowing Claude to access part of the company’s infrastructure.
In the third case, after failing to locate its fictional target, Claude searched online for alternative targets, using basic credential exposure and SQL injection techniques before stopping once it recognised the organisation was real. Anthropic said it has informed the first two organisations and is continuing efforts to contact the third.
Configuration Error Allowed Internet Access
Anthropic said Claude’s internet access resulted from a configuration mistake rather than the AI independently finding a way online.
Although the model had been told it was operating in an isolated simulation without internet access, a misunderstanding between Anthropic and its evaluation partner unintentionally left internet connectivity enabled.
The models relied on techniques such as weak passwords and unauthenticated endpoints to access the organisations. Anthropic said no zero-day vulnerabilities were exploited.
Models Acted Within Their Instructions
The incidents involved Claude Opus 4.7, the cybersecurity-focused Mythos 5 model and another internal research model.
Anthropic said it found no evidence that the models acted autonomously beyond their assigned instructions. Instead, they believed they were operating entirely within a simulated testing environment.
The company added that the safety protections built into its publicly available Claude models would have prevented the behaviour seen during these internal evaluations. Anthropic has appointed independent AI evaluation organisation METR to conduct a third-party review of the incidents.
Matt Warren, director of RMIT University’s Centre for Cyber Security Research and Innovation, said the incidents demonstrate that generative AI can identify weaknesses in IT systems, but warned that publicly discussing these capabilities could also increase awareness among cybercriminals.
Source: Reuters
At Inspirepreneurs Magazine, covering entrepreneurship, business failures, and the human stories behind the world's most ambitious founders. She writes at the intersection of strategy and storytelling.