GPT-6 Astra launch pairs perfect ExploitBench run with stricter guardrails

OpenAI ships GPT-6 Astra after the model posts a perfect ExploitBench score and finds two zero-days in testing.

CSBadmin
2 Min Read

Perfect score on an exploit-building exam, two fresh zero-days found during its own testing, and a safety label OpenAI reserves for its riskiest systems. That is the arrival profile for GPT-6 Astra, which the company began rolling out to select organizations on Thursday.

ExploitBench hands a model a known vulnerability and grades how reliably it produces a working exploit. OpenAI reports that Astra, tested without its production safeguards, went a perfect 100 percent on that exam, a jump from the 78.5 percent posted by GPT-5.6 Sol, its previous frontier model. On the broader ExploitGym benchmark, it reached 42.4 percent, up from 30.3 percent.

The company also pointed the model at flaws disclosed in the three months before launch to check whether it could find bugs on its own. Astra discovered two zero-days that were not part of its training material, OpenAI says, and both are now being reported to the affected software vendors.

The version now reaching customers is deliberately limited. It will review code and write patches, and it declines requests to generate proof-of-concept exploits. OpenAI plans to relax those guardrails for vetted defenders in coming weeks through Daybreak, a $1B program that gives utilities, local governments, banks, and open-source maintainers subsidized access, training, and technical help.

In a separate evaluation built after the Hugging Face incident, Astra exceeded its authorized scope in zero percent of runs, where GPT-5.6 Sol did so 48 percent of the time, OpenAI reports. Enterprise access is off by default and must be switched on by administrators, and API pricing starts at $10 per million input tokens.

CSBadmin

The latest in cybersecurity news and updates.

Share This Article
Follow:
The latest in cybersecurity news and updates.