OpenAI just told the world that its newest model, GPT-6 Astra, is capable of finding and weaponizing security flaws well enough to earn the company’s own “Critical” cyber-risk rating — the highest tier in its Preparedness Framework, and the first time any OpenAI model has crossed it. That headline is real. What it means for you, practically, is a lot murkier than the headline suggests, and the gap between the two is worth sitting with.
My position: OpenAI publishing this rating is a genuine step toward transparency, and Astra’s underlying capability jump is real and worth taking seriously. But “Critical” is a label OpenAI assigned to its own model using its own threshold, tested by its own team, with no independent auditor’s sign-off disclosed alongside it. Treating a self-graded exam as the final word on how dangerous this model is — or how safe the mitigations around it are — is the part that deserves more skepticism than it’s getting.
What OpenAI’s “Critical” Cyber Risk Rating Actually Means
The Preparedness Framework isn’t a marketing term. Under OpenAI’s own definition, a model hits the Critical threshold if it can identify and develop functional zero-day exploits — across all severity levels — in hardened, real-world systems without a human walking it through each step. Or, alternatively, if it can take a high-level goal like “compromise this target” and independently devise and execute the entire attack chain itself. OpenAI lays this out in detail in its own safety overview for GPT-6 Astra, the document this whole story traces back to.
Astra reportedly hit both bars. It scored 100% on ExploitBench, a benchmark that measures whether a model can turn a documented, publicly known vulnerability into a working exploit. That’s the easier half of the test — the flaw is already known, the model just has to build the weapon. The harder number is the 39% it scored finding genuinely novel vulnerabilities from the prior three months, flaws nobody had catalogued yet. During pre-release evaluation, Astra didn’t just find two such zero-days — it chained them together into a full browser-compromise sequence that escaped the sandbox and executed commands on the host machine, unassisted.
That last detail is the one that should actually change how you read this story. A model that can pass a benchmark of known exploits is impressive but bounded. A model that strings together previously unknown flaws into a working sandbox escape, on its own, is a different category of result — closer to an autonomous red-teamer than a coding assistant with security knowledge.
The Zero-Days Are the Real Story, Not the Label
It’s easy to get stuck on the word “Critical” and miss the more interesting fact underneath it: this is the clearest public evidence yet that frontier models are starting to do original offensive-security research, not just recite it. Claude has separately been credited with surfacing cryptography flaws in released software — we covered that discovery in detail here — and the pattern across labs is the same: these systems are getting good enough to find bugs humans missed, in both directions. That’s genuinely useful for defenders who get there first. It’s genuinely dangerous if an attacker’s model gets there first instead.

OpenAI’s response to its own finding was to tighten the model’s leash rather than withhold it. GPT-6 Astra ships with stricter isolation around agentic deployments, checkpoint encryption, and what OpenAI describes as universal monitoring of full model trajectories — including chain-of-thought reasoning, not just outputs. On cyber jailbreak evaluations, Astra refuses 91.5% of disallowed requests, up from 59% for the previous flagship, GPT-5.6 Sol. Accounts flagged as higher-risk get an even more conservative behavior boundary applied on top of that.
The Cyber Risk Safeguards Are Real. So Is the Blind Spot.
None of that is nothing. A jump from 59% to 91.5% refusal on adversarial cyber prompts is a meaningful engineering result, not a rounding error. But here’s the skeptical-investigator read: every one of those numbers — the refusal rate, the ExploitBench score, the “Critical” classification itself — comes from OpenAI testing OpenAI’s model against a framework OpenAI wrote. There’s no independent lab confirming the 91.5% figure holds up against attack techniques OpenAI didn’t think to test. There’s no third party verifying that “universal monitoring” actually catches a determined bad actor rather than just the evaluation suite it was tuned against.
That’s not an accusation of dishonesty. It’s a structural gap. Self-assessment against a self-defined threshold is the same shape of problem as a company auditing its own financial statements: the incentives to be rigorous exist, but so does the incentive to publish a number that reassures rather than alarms. A framework this consequential — one that’s effectively deciding how much autonomous hacking capability gets deployed to millions of users — is exactly the kind of thing that benefits from outside eyes, and right now those eyes aren’t part of the published record.
Devil’s Advocate: Maybe This Is Exactly How It Should Work
Steelmanning OpenAI’s approach
There’s a fair counterargument here. OpenAI didn’t have to publish a “Critical” rating on its own model at all — disclosing your own worst-case capability assessment is not the industry norm, and doing it invites exactly the kind of scrutiny this article is applying. The Preparedness Framework, whatever its blind spots, is a published, versioned document that lab-watchers, journalists, and competitors can hold OpenAI to going forward. If Astra later gets caught being misused and the safeguards described here didn’t hold, that’s now a documented promise OpenAI broke, not a vague assurance it never made. Radical transparency about your own risks, even self-graded, is still more accountability than silence.
That’s a real point, and it’s why this isn’t a story about OpenAI hiding something. It’s a story about an industry still figuring out what “trust but verify” looks like when the entity being trusted is also the only one currently doing the verifying. Independent evaluation organizations exist for exactly this reason, and the more consequential these ratings become, the harder it will be to justify keeping them purely in-house.
What Actually Changes for You
If you’re an everyday ChatGPT user, almost nothing changes today. The safeguards described above sit at the infrastructure and account-monitoring level, not in front of the chat box you type into. If you’re a developer building on OpenAI’s agentic tools, this is worth reading closely — higher-risk account flags and tighter agentic isolation could affect what your integration is allowed to do, and it’s a reasonable moment to check what the “Critical” designation says about default access levels for automated systems. And if you’re just trying to make sense of the broader AI-safety conversation, Astra is a useful data point for a debate we’ve covered before on where AI oversight actually sits — between labs, governments, and independent auditors, and who gets the final say when a model crosses a line like this one.
The honest takeaway is that “Critical” is both an accurate description of a real capability jump and a term whose full weight you’re being asked to take on OpenAI’s word. Both things are true at once. The zero-day-chaining result is genuinely new territory for a publicly deployed model, and the safeguards built around it look substantive rather than cosmetic. But a rating system with no outside referee is a rating system that will always tend to grade itself generously, even with the best intentions in the room. Watch what happens the next time a lab’s own framework says “Critical” about something — and watch even closer for who, besides that lab, gets to check the math.
Leave a Reply