OpenAI announced Tuesday that its upcoming AI model, Astra, is the company’s first model to reach what it calls the “critical” threshold for cybersecurity capabilities. OpenAI plans to release a version of Astra to the public “soon,” but its most advanced cyber features will initially be available only to select partners through the Daybreak Blue early access program.
OpenAI’s safety and security leaders said during a briefing with reporters that Astra has reached the critical cybersecurity capability threshold outlined in the company’s preparedness framework. The framework establishes thresholds and protocols for situations in which AI models create new levels of risk. OpenAI defines the critical cyber threshold as the point at which an AI model can independently discover and exploit previously unknown vulnerabilities in real-world software. Company leaders said OpenAI is following its safety procedures by halting further development until appropriate safeguards and security measures are in place.
OpenAI previously said it had paused some training workloads related to Astra and future AI model development for several weeks. Executives said the company has since resumed work on Astra and other AI models after introducing additional safety and security controls. OpenAI described the pause as productive and said it is now confident that Astra can be released more broadly in a safe manner.
The announcement comes as Silicon Valley confronts the cybersecurity risks and defensive potential of advanced AI models. Companies are working to reassure users, lawmakers, and industry partners that they can control these increasingly capable systems. In July, OpenAI disclosed an incident in which an agent running two of its models exploited a vulnerability in what was intended to be a siloed test environment. The agent accessed the internet and hacked the company’s open-source AI platform, Hugging Face. OpenAI noted that Astra was not involved in the incident.
Other AI companies, including Anthropic and Meta, have disclosed similar incidents in recent weeks. On Monday, Anthropic announced that it had paused some AI training workloads while strengthening its safety and security practices.
OpenAI said it is taking a multi-tiered approach to restrict public access to Astra’s advanced cybersecurity capabilities, including a new “misalignment monitor.” For example, if a user asks Astra to find an exploit for a real-world software system, the model is designed to refuse the request. OpenAI also said Astra is more resistant to jailbreak attempts and rejected unsafe queries at a significantly higher rate than previous models during testing.
However, OpenAI noted in a blog post that its misalignment monitors “may flag legitimate activity as potentially cyber-exploited or fraudulent, which could result in it being inadvertently slowed, paused, or terminated.” The company said these safeguards may be triggered even when a user’s activity is not believed to be related to cybersecurity. In such cases, ChatGPT and Codex users may be asked to confirm the model’s actions before continuing.
Partners in OpenAI’s Daybreak program, including digital infrastructure companies such as Cisco, Cloudflare, and Palo Alto Networks, will receive early access to a less restricted version of Astra with stronger cybersecurity capabilities. The program is intended to help these companies use advanced AI models to improve their defenses before similar capabilities become widely available. OpenAI leaders also said the company is working closely with government partners to ensure they understand Astra’s cyber capabilities and have access to the technology.
In addition to discovering previously unknown software vulnerabilities and developing methods to exploit them, Astra can reportedly “chain” multiple exploits together. Attackers use this technique to penetrate deeper into a target system and gain access that may not be possible by exploiting a single vulnerability.
Source: www.wired.com


