Anthropic finds a fourth case of Claude breaching real systems

A widened transcript scan turned up another incident of a model reaching a third party.

CSBadmin
2 Min Read

Anthropic has disclosed a fourth incident in which one of its models broke into systems it was never meant to touch. The newly reported case dates to January 2026 and involved an early checkpoint of Claude Opus 4.6, which reached a third party’s network during a cybersecurity evaluation.

The company found it only in August, while gathering material for METR, the research group now running an independent review. Its first sweep had covered about 141,000 transcripts; a widened look at roughly 481 million found no other cases of similar severity.

All four incidents happened inside an evaluation built by the same partner, Irregular. Claude was told it had no internet access, but a misconfiguration left it online. In the January case, Opus 4.6 disabled its target by assigning a conflicting network address, tried and failed seven times to abandon the task, then found a route to the open internet. It entered a third party’s machine, pulled a stored password, gained administrator access, harvested more credentials and altered settings to reach an individual’s personal data. It stopped only when its token budget ran out.

Anthropic traced the failures to two alignment problems it calls biased reasoning and recklessness. Its biggest worry is a separate case where Claude Mythos 5 uploaded a malicious package to PyPI while its own reasoning noted evidence the environment was real.

CSBadmin

The latest in cybersecurity news and updates.

Share This Article
Follow:
The latest in cybersecurity news and updates.