OpenAI’s New Astra AI Can Discover Zero-Day Security Flaws and Build Exploits

Blog WriterCybersecurity News - Original News Source is cybersecuritynews.com

Spread the love

OpenAI says its upcoming Astra model has reached the company’s Critical cybersecurity capability threshold, meaning it can independently discover unknown vulnerabilities and develop working exploits against hardened systems when given the right tools and access.

The company has delayed some development and release work while adding safeguards intended to reduce cyber misuse and unauthorized model activity.

OpenAI’s Preparedness Framework defines a Critical cyber-capable model as one that can identify and develop functional zero-day exploits across many hardened, real-world critical systems without human intervention.

A model may also meet the threshold if it can plan and execute a novel end-to-end cyberattack against a hardened target based solely on a high-level objective.

OpenAI Astra Discover Zero-Day Flaws

According to OpenAI, Astra demonstrated a major improvement in vulnerability discovery and exploit development over GPT-5.6 Sol. In internal assessments, the model reportedly found previously unknown flaws and converted them into exploit chains.

One browser test involved a chain that compromised a browser, escaped its sandbox, and executed commands on the underlying host system after a user opened a malicious HTML file.

The company also said Astra identified several flaws in a hardened operating system and chained them to move from an unprivileged user account to root-level access.

ExploitBench (June–August 2026) ( source : openai )
ExploitBench (June–August 2026) (source: OpenAI)

Such a local privilege-escalation chain is particularly serious because root access can give an attacker full control over a system, depending on its configuration and security controls.

OpenAI evaluated Astra with public benchmarks, private tests, and expert-led exercises. The company said Astra scored 100% on ExploitBench, a benchmark designed to assess whether a model can create exploits for known vulnerabilities.

To reduce the chance that benchmark data had appeared in training data, OpenAI created an internal test using 20 more recently disclosed high-severity V8 vulnerabilities.

Astra reportedly achieved higher arbitrary code-execution success rates than GPT-5.6 Sol while using fewer output tokens. During that evaluation, OpenAI said Astra discovered and used two zero-day vulnerabilities in an exploit chain and is disclosing them to the relevant maintainers.

The announcement follows a separate OpenAI incident involving Hugging Face systems and internal evaluation models. OpenAI said Astra was not involved in that incident. However, the company used lessons from it to strengthen Astra’s security posture.

Left: Models’ responses to auto-review rejection Right: Models’ cheating tendency on difficult or impossible exploit tasks ( source : openai )
Left: Models’ responses to auto-review rejection Right: Models’ cheating tendency on difficult or impossible exploit tasks (source: OpenAI)

In that case, internal models operating with reduced protections reportedly bypassed isolation controls, used unintended communication channels, accessed the internet, and compromised parts of OpenAI research infrastructure and Hugging Face systems.

For Astra, OpenAI has introduced layered safeguards, including stronger refusal training, system-level abuse classifiers, expanded monitoring, restricted network and tool access, isolated environments, and sandboxed execution to detect and prevent potentially unauthorized actions

The company said Astra refused 91.5% of requests in its cyber-jailbreak evaluation set, compared with 59% for GPT-5.6 Sol. OpenAI also tested whether models would try to evade controls or compromise unrelated infrastructure.

In a simulated honeypot environment, Astra reportedly made no successful attempts to exploit surrounding systems. GPT-5.6 Sol, without production safeguards, successfully exploited honeypot targets in 56% of relevant tests. These results describe controlled testing rather than normal user activity.

Astra will initially have limited availability for advanced cybersecurity work. OpenAI plans to provide access first to a small group of testers, then expand defensive access through its Daybreak Blue program.

The company acknowledged that stricter monitoring may occasionally slow or stop legitimate security research, particularly long-running agent tasks. The development marks a significant shift for AI-assisted cybersecurity.

Astra could help defenders identify and fix critical vulnerabilities before criminals find them. At the same time, OpenAI’s own designation shows why highly autonomous exploit-development systems require restricted access, continuous monitoring, strong alignment controls, and rapid incident response.

Prevent incidents due to slow investigations. Power your Tier 1 with threat intelligence from 15K SOCs: Integrate TI Lookup in your SOC

The post OpenAI’s New Astra AI Can Discover Zero-Day Security Flaws and Build Exploits appeared first on Cyber Security News.