OpenAI just built an AI that can find and exploit hacking vulnerabilities completely on its own.

Here's the full story, via WIRED.
What "Critical" Actually Means
"OpenAI announced Tuesday that its forthcoming AI model, Astra, is its first to reach the company's threshold for what it calls 'critical' cyber capabilities." Specifically, "the company says an AI model has reached its critical cyber threshold when it can independently find and exploit previously unknown vulnerabilities in real-world software."
It Doesn't Stop at One Vulnerability
"Astra is not only capable of finding novel software vulnerabilities and developing ways to exploit them for hacking, but is also able to 'chain' multiple exploits together, a technique used to bore deeper and deeper into a target system and gain access that wouldn't be attainable using just one vulnerability."
The Score
"Astra outperforms industry leading AI models such as GPT-5.6 Sol and Anthropic's Mythos on cybersecurity benchmarks such as ExploitBench, which Astra scored 100 percent on."
What OpenAI Did About It
Reaching the threshold triggered a real internal response: OpenAI "previously said that it paused some training workloads related to the development of Astra and a future AI model for several weeks," resuming only after adding "additional safety and security controls." One of those controls is a new "misalignment monitor" that's supposed to make Astra refuse requests to find real-world exploits. OpenAI notes the monitor "may occasionally flag legitimate activity as potential cyber misuse," in which case users may be asked to confirm an action before it proceeds.
Who Gets Access First
"Partners in OpenAI's Daybreak program, which includes digital infrastructure providers like Cisco, Cloudflare, and Palo Alto Networks, will get early access to a less restricted version of Astra with more robust cyber capabilities." The stated goal: let defenders harden their systems before similarly capable models become broadly available.
It's Not Just OpenAI
Anthropic disclosed its own pause on some training workloads the same week "while it hardens its safety and security practices." Both companies have spent recent months publicly disclosing incidents where their own AI agents acted outside intended boundaries, including OpenAI's earlier Hugging Face hacking incident (which it says did not involve Astra).
The Takeaway
The AI industry's own cybersecurity capabilities are now outpacing what most organizations have defended against. As WIRED notes, cybersecurity experts still emphasize that longstanding best practices remain durable, but organizations that haven't fully implemented them are now at "even more urgent risk."
Source: WIRED (Maxwell Zeff, Lily Hay Newman): https://www.wired.com/story/openai-astra-first-ai-model-with-critical-cyber-abilities/