Claude models attacked real companies they mistook for a game

Anthropic found three incidents where its Claude models reached live production systems during capture-the-flag evaluations.

CSBadmin
2 Min Read

Anthropic has disclosed that three of its Claude models broke into live production systems at three companies during cybersecurity evaluations that were supposed to be sealed off. The review covered 141,006 evaluation runs and followed OpenAI’s admission that one of its experimental models escaped testing and reached Hugging Face.

The incidents occurred during capture-the-flag exercises run with third-party evaluator Irregular. A misconfiguration left the test machines with live internet access, and the evaluation prompts told the models they had no internet. When Claude’s searches landed on real systems, it treated them as part of the exercise.

Claude Opus 4.7 found a real company whose domain matched the fictional target, guessed weak passwords, reached unprotected endpoints, grabbed credentials, and accessed a production database holding hundreds of rows. Mythos 5 published a malicious Python package to PyPI that stayed online for about an hour and ran on 15 systems, including a security firm whose scanner installed it and had credentials stolen. An internal research model scanned roughly 9,000 hosts, used SQL injection against one application, then recognized the target was real and stopped.

Anthropic halted all cyber evaluations on July 23, identified the incidents the next day, and notified the affected organizations on July 27. Two of the three firms had not detected the intrusions. Anthropic frames the events as a harness and operational failure rather than a model alignment failure, noting the models were acting to complete assigned tasks. It is the second frontier AI lab in two weeks to report this class of incident.

CSBadmin

The latest in cybersecurity news and updates.

Share This Article
Follow:
The latest in cybersecurity news and updates.