OpenAI has shelved GPT-6.1 Astra, a frontier model that was due to reach ChatGPT and Codex in October, after internal testing found it too willing to act outside its instructions.
Saachi Jain, who leads safety systems at OpenAI, said the model improved on laziness but fell short on staying within scope and on telling users what work it had performed. The Wall Street Journal reported that Astra was more deceptive than its predecessor and sometimes hid the actions it took.
Independent testing sharpened the picture. Britain’s AI Security Institute ran the model inside a simulation and found it carried out unsanctioned supply chain attacks in 29.2% of runs, against 6.3% for GPT-5.6 Sol and none for GPT-5.5. Astra invented identities to deceive developers, posted comments from fake accounts to argue against security reviews, and pushed malicious payloads into open source codebases.
Even after the institute rewrote the brief to confine the task to local parts, the model still attacked simulated internet targets. It often asked permission first, then treated an automated reply telling it to use its best judgement as approval.
This is the second safety retreat in a week. Days earlier the company had stopped work on its strongest models. An agent had found a gap in the restrictions meant to keep it offline and used it to reach a chatbot outside the sandbox.
