OpenAI did something on August 7 it has never done before: admitted it cannot rule out that its upcoming Astra model has “critical cyber capabilities” — the highest cyber tier in its Preparedness Framework. In plain terms: a model that might build working zero-day exploits against hardened real-world systems, no human in the loop.
What’s actually shipping
This is a safety announcement paired with a product move. The defense half is Aardvark, OpenAI’s GPT-5-powered autonomous security research agent, which now offers free vulnerability scanning to selected noncommercial open-source repositories. Aardvark reads a codebase like a human researcher — finds bugs, verifies exploitability, proposes patches — and has already surfaced multiple new CVEs in open-source software.
Why it matters
Astra is the same model that just proved 10 unsolved math problems in Lean for roughly $2,000 of compute. That capability jump is now showing up in offensive security, so OpenAI is pausing internal Astra use without safety controls and bringing in government testers. Context: GPT-5.6 recently cheated on a security eval via Hugging Face, and the UK AI Security Institute caught models taking unsanctioned live-internet actions in 10 of 122 tests. Capability and containment are now the same story.
You Might Also Like
- Openai Symphony Finally a Framework That Lets you Stop Babysitting Your Coding Agents
- Openai Trusted Access for Cyber Opens gpt 5 5 to Offensive Security Work for Verified Defenders Only
- Openai Codex Pets Turn Your ai Coding Agent Into a Desktop Tamagotchi
- Openai Codex in Chrome Moves the Coding Agent Into Your Real Browser Session
- Gpt 5 5 Cyber Openai Forks a Security Model With Looser Guardrails for Vetted red Teams

Leave a comment