
Photo: Lucas Andrade / Pexels
OpenAI's Astra Is the First AI Model Rated "Critical" for Cybersecurity — It Found Zero-Days on Its Own
Chris Harper
2 min read
Sep 2, 2026 · 20:07 UTC
TL;DR: OpenAI's Astra is the first model rated "Critical" under its Preparedness Framework — it autonomously found zero-day vulnerabilities and built working browser-escape and OS privilege-escalation chains during evaluations.
In expert-led assessments, Astra scored 100% on ExploitBench (which measures exploit-chain generation from known CVEs), discovered two previously unknown zero-days during separate evaluations on recently disclosed flaws, built a full browser-compromise chain that escaped a sandbox and executed host OS commands, and chained multiple hardened-OS vulnerabilities into a privilege-escalation path.
OpenAI reports the model declines 91.5% of cyber-related jailbreak attempts — up from 59% for GPT-5.6 Sol. That still means roughly 1 in 11 targeted attempts succeeds, which is a real failure rate for security-sensitive deployments.
The "Critical" designation means Astra will be released, not blocked, but with gated access through OpenAI's Daybreak Blue program for vetted security researchers — the same capability-gated access pattern Anthropic applied last month for Enterprise Frontier Safeguards.
Why it matters: Autonomous exploit-chain discovery — no per-step human guidance — changes the threat model for agentic systems with broad tool access to shell, file system, or network. If you're building agents with those permissions, ExploitBench capability rating is now the reference point for what adversarial misuse looks like.
Sources: Path to Astra: critical capabilities and frontier safeguards — OpenAI · OpenAI's Astra becomes first model to cross "Critical" cybersecurity threshold — SecurityWeek · OpenAI Astra: First AI Model With "Critical" Hacking Abilities — Decrypt