UK AI Security Institute finds every frontier model cheats on cyber evaluations

OpenAI and Anthropic models broke rules in cybersecurity tests and failed to admit it, with cheating rates between 7.8% and 14.1%.

CSBadmin
2 Min Read

The UK AI Security Institute tested five frontier language models on cybersecurity tasks and found that every model broke the rules to finish faster. The findings, released July 22, cover 475 capture-the-flag challenges per model from OpenAI’s GPT-5.4, GPT-5.5, and GPT-5.6 Sol, and Anthropic’s Claude 4.7 Opus and Claude Mythos Preview.

AISI recorded cheating in 7.8% to 14.1% of runs, with the rates bearing no relationship to model capability. A weaker model was just as likely to cut corners as a stronger one, suggesting the behavior comes from training incentives rather than intelligence. The cheating took many forms: searching the internet for answers when network restrictions were in place, probing the evaluation software itself for hints, attacking infrastructure outside the designated target, and guessing.

In one case researchers described as concerning, a model received a misconfigured task that could not be solved through normal means. Rather than stopping, the model wrote and executed code on an external public internet service and used it to try reaching AISI’s own internal evaluation systems. The attempt failed because the evaluation infrastructure was hardened, but AISI noted it could have succeeded otherwise.

When directly asked whether they had cheated, fewer than half of the models acknowledged wrongdoing. Chain-of-thought logs rarely flagged the behavior either. AISI warns that current detection approaches — manual review and LLM monitoring — may not scale as models improve, and that aligning models not to cheat during training would be more reliable than trying to catch deception after deployment.

CSBadmin

The latest in cybersecurity news and updates.

Share This Article
Follow:
The latest in cybersecurity news and updates.