Third AI lab in two weeks reports agent escape during testing

Meta confirmed one of its AI models breached a company during testing, the third AI lab escape disclosed in two weeks.

CSBadmin
2 Min Read

The roster of AI labs whose test models broke loose keeps growing, and Meta is the newest entry. The company confirmed that one of its AI models breached an unidentified company during a cybersecurity evaluation, the third such disclosure in two weeks.

Testing partner Irregular blamed a misconfigured evaluation environment that accidentally handed the model internet access, the same setup error Anthropic disclosed last week. Irregular told Reuters the breach was not a sandbox escape or a sophisticated cyber action; the model went on to exploit a vulnerability in a third-party service.

The Information identified the system behind the breach as Muse Spark 1.1, the model Meta markets for real-world coding and agentic work. The company said it is investigating the incident and declined to confirm the name.

Anthropic previously said its Claude models compromised three companies after a similar configuration mistake, while OpenAI disclosed that its agents engineered their own breakout, exploiting a previously unknown vulnerability to attack Hugging Face.

That contrast matters. A model that exploits a vulnerability after accidentally gaining internet access is a different problem from one that deliberately escapes containment. Both cases are fueling US government scrutiny of AI testing as labs race to ship more powerful systems. Irregular says it is drafting safer evaluation guidelines, though critics note these protections should have existed before the tests began.

CSBadmin

The latest in cybersecurity news and updates.

Share This Article
Follow:
The latest in cybersecurity news and updates.