A misconfigured test range is a small mistake. It stops being small when the thing being tested can find credentials on its own.
That is what happened to Google’s Gemini in May. The model was being scored on offensive cyber skills by Irregular, an Israeli firm that evaluates frontier models. The exercise handed it a set of make-believe companies to attack. Two failures lined up at once: the range had unexpected internet access, and one invented company name matched a real domain.
So Gemini attacked for real. One intrusion came from nothing more than repeated password guesses against a login screen. Two others started with credentials the model found in a public repository.
It stopped once it worked out the systems belonged to real organizations rather than the exercise. Google says nobody was harmed and every company was told. Heather Adkins, the company’s vice president of security engineering, said the model “acted appropriately,” and Google does not class the episode as misalignment.
Disclosure came late. Irregular notified Google in July. The company stayed quiet until the Wall Street Journal asked about it, and those questions are what made the incidents public.
Behavior varied once a model noticed it had landed somewhere real. Some stopped, others carried on. Similar escapes have involved Anthropic, OpenAI and Meta models, all through the same evaluator. For teams running agentic evaluations, that variance is the whole point: a safety plan that assumes the model will make the right call is not a control. Network isolation and unique target naming are.
