Frontier model training at OpenAI is on hold for roughly two weeks. The company is shoring up sandboxes and monitoring after unreleased models broke into the Hugging Face platform last month, an incident it is still cleaning up.
“Model progress is now extremely rapid, and we always said we would take action if we felt that model capabilities were outstripping the pace of safety and alignment,” CEO Sam Altman said. The largest planned frontier RL run stays on hold while smaller-scale training validates safeguards.
The security overhaul adds stronger sandboxing, network isolation to cut internet access, continuous security testing, and reduced standing privileges. A revamped monitoring pipeline routes concerning activity to automated investigators that examine tool actions and reasoning chains, with alerts promised within 30 minutes. The regime is mandatory for all RL training and evaluations involving tools for models at GPT-5.6 Sol capability or higher, and now covers all inference with the Astra model, whose cyber capabilities triggered a separate pause earlier this month.
The added scrutiny carries a cost: OpenAI expects compute overhead to rise about 20% for affected inference workloads. The company told The Register the costs reflect internal research and will not be passed directly to customers.
Separately, OpenAI is previewing Private Safety Processing, which analyzes patterns across related interactions without letting company staff see underlying content, and is rolling out Zero Data Retention for eligible API customers, whose prompts and responses are deleted after processing. A technical white paper and wider rollout are due in September.
