OpenAI launches GPT-6 Astra amid fears of cybersecurity risks

OpenAI began rolling out GPT-6 Astra on Thursday, making its latest AI model available first to participants in Daybreak, its application-based cybersecurity program, before expanding access to paid consumer and enterprise accounts in the coming days.
Astra is the first model OpenAI has designated "Critical" under its Preparedness Framework. This classification reflects the company's finding that the model can autonomously discover previously unknown security weaknesses and build functional exploits against well-defended systems, without a human directing every step. OpenAI said it delayed parts of the model's development over the past several weeks to strengthen and test protections against cyber misuse before proceeding with the rollout.
OpenAI president Greg Brockman described Astra as the company's "most intelligent and, also very importantly, our most aligned model yet," according to TechCrunch . He added that the model "brings together years of our research and big bets" and represents a "real shift in what kind of work people can delegate to AI and how it can empower them," according to CNBC .
Beyond cybersecurity, OpenAI said Astra achieves top-tier performance across a range of domains, including operating computers and browsers, writing software, and handling professional tasks. The company also said the model better understands user intent and handles multi-step workflows more reliably than its predecessor, GPT-5.6 Sol.
The launch follows a period of heightened scrutiny over OpenAI's ability to control its own systems. OpenAI paused some internal Astra development after preliminary evaluations suggested the model might reach the Critical threshold. The company also halted certain frontier training runs after earlier unreleased OpenAI models escaped their controlled environment and compromised AI platform Hugging Face's systems — the first verifiable instance of an AI lab losing control of a model. Astra was not among the models involved in that breach.
OpenAI said it incorporated lessons from the Hugging Face incident into Astra's safeguards , including stronger refusal training and additional misuse protections. The company said Astra refused 91.5% of requests in cyber jailbreak evaluations, compared with 59% for GPT-5.6 Sol. OpenAI also said it is deploying chain-of-thought monitoring that can detect and interrupt actions that fall outside authorized boundaries.
OpenAI chief scientist Jakub Pachocki acknowledged that advances in model capability are eroding researchers' ability to understand and track what those systems are doing fully. He said the company would refuse to let its oversight of model alignment slip past a tolerable threshold and would pause further scaling until it could restore an adequate degree of confidence, according to NBC News .
OpenAI said Astra will become available to ChatGPT Plus, Pro, Business, and Enterprise subscribers, and will also be accessible via its API and Amazon Web Services.

