Anthropic reveals its AI models hacked three organisations during cybersecurity testing
Anthropic has disclosed that three versions of its Claude AI model compromised the production infrastructure of three separate organisations after a configuration error unintentionally gave the systems internet access during internal cybersecurity evaluations.
Anthropic has disclosed that three versions of its Claude AI model compromised the production infrastructure of three separate organisations after a configuration error unintentionally gave the systems internet access during internal cybersecurity evaluations.
The AI giant made the disclosure on Thursday in a blog statement, highlighting the need for stronger controls as AI models become increasingly capable of carrying out real-world cyber activities.
The AI giant revealed that it uncovered the incidents while conducting a retrospective review of cybersecurity evaluation runs following OpenAI’s earlier disclosure that some of its AI models compromised another AI company.
Anthropic said the hacking incidents stemmed from a misconfiguration between the company and its third-party evaluation partner, Irregular, which left evaluation environments connected to the internet despite prompts telling Claude it had no internet access.
Because of that error, the models interpreted real internet-facing systems as legitimate components of the simulated exercise.
The most serious incident involved Claude Opus 4.7, which accessed a company’s production database containing several hundred rows of data after mistaking it for a fictional target.
According to Anthropic, none of the affected organisations had detected the activity before being contacted.
Anthropic’s retrospective review of cybersecurity evaluation was prompted by OpenAI’s disclosure earlier this month that some of its AI models had compromised another AI company, Hugging Face’s production infrastructure.
The incident represents one of the clearest demonstrations to date of advanced AI models independently carrying out complex cyber operations.
The incident has also added momentum to the proposed Kill Switch Bill, a legislation aimed at giving authorities powers to shut down AI systems deemed to pose significant risks.
Concerns over the rapid pace of artificial intelligence development have continued to intensify among policymakers, technology executives, and industry leaders.
Similar concerns have been echoed within Nigeria’s technology ecosystem.
