Home / Cybersecurity / Anthropic reveals Claude AI breached real-world systems during cyber tests

Anthropic reveals Claude AI breached real-world systems during cyber tests

Anthropic reveals Claude AI breached real-world systems during cyber tests

Anthropic has disclosed that three of its Claude AI models gained unauthorised access to the live systems of three organisations during cybersecurity evaluations, raising fresh concerns about the risks of increasingly autonomous AI.

The company said the incidents occurred after a testing misconfiguration mistakenly gave the models internet access, despite being told they were operating in isolated environments.

During the evaluations, the AI models exploited weak passwords, exposed credentials and other basic security flaws. In one case, Claude Opus 4.7 accessed a production database containing live data, while another model uploaded a malicious Python package to the public PyPI repository.

Anthropic stressed the models did not intentionally escape containment, describing the incidents as operational failures rather than AI alignment issues.

The disclosure follows a similar incident involving OpenAI, highlighting growing industry concerns over AI’s ability to autonomously carry out offensive cybersecurity tasks.

Main Image: The Guardian

Tagged:

Leave a Reply

Your email address will not be published. Required fields are marked *