Trained on over 100,000 GPUs at the Texas Stargate facility, Astra aced ExploitBench with a 100% score during testing.

On 3 September 2026, OpenAI officially unveiled GPT-6 Astra, marking the first artificial intelligence model to reach the "Critical" cybersecurity risk threshold under the organisation's Preparedness Framework. According to an official safety report published by OpenAI, this classification indicates that the sixth-generation model possesses autonomous capabilities to identify and exploit previously unknown security vulnerabilities across hardened systems without requiring step-by-step human guidance. The revelation prompted OpenAI to delay public access and implement unprecedented internal containment protocols prior to commercial deployment.
Unprecedented Exploit and Zero-Day Discovery
The model’s offensive potential was demonstrated during pre-release testing evaluations. As reported by NeuralTrust AI, GPT-6 Astra achieved a flawless 100 per cent score on ExploitBench, a benchmark measuring automated exploit generation, while successfully discovering two previously unknown zero-day vulnerabilities in production software. Furthermore, Fortune noted that pre-training the model required OpenAI’s largest compute cluster to date, utilising over 100,000 graphics processing units at its Stargate facility in Texas, resulting in autonomous reasoning power that alarmed internal red-team evaluators.
Stricter Isolation and Chain-of-Thought Monitoring
To mitigate potential misuse or escape, OpenAI placed Astra under a stringent operational lockdown. According to OpenAI, the version made commercially available to ChatGPT Enterprise and API users operates with severely restricted refusal boundaries that suppress offensive exploit generation. Internally, OpenAI implemented checkpoint encryption, strict environmental isolation, and universal real-time monitoring of full inference trajectories, including the model's hidden "chain of thought" reasoning, to ensure the system remains fully aligned with human oversight.
Vetted Defence via the Daybreak Programme
Rather than withholding the technology entirely, OpenAI established a tiered access framework to assist cybersecurity defenders. As detailed by NeuralTrust AI, the company created the OpenAI Daybreak programme, providing vetted cyber defence organisations, infrastructure maintainers, and security researchers with controlled access to a less restricted version of Astra. This defensive initiative allows security teams to utilise the AI’s advanced reasoning for vulnerability validation, malware analysis, and detection engineering, ensuring defenders retain parity with potential threat actors.
Entering the AGI Era Under Lockdown
The containment of GPT-6 Astra highlights a turning point in frontier AI safety governance. OpenAI President Greg Brockman stated during a press briefing that Astra’s arrival signals that society has entered the "AGI era", acknowledging that managing autonomous, highly capable models is now the central challenge of modern technology. Chief Scientist Jakub Pachocki echoed these concerns, warning that preventing unintended harm from hyper-capable reasoning systems is rapidly becoming the primary bottleneck to further AI advancement.
Published: 07 Sept 2026, 08:03 am IST
ABOUT THE AUTHOR
Related Topics
Get Latest Mathrubhumi Updates in English
Disclaimer: Kindly avoid objectionable, derogatory, unlawful and lewd comments, while responding to reports. Such comments are punishable under cyber laws. Please keep away from personal attacks. The opinions expressed here are the personal opinions of readers and not that of Mathrubhumi.

