A row of illuminated blue server racks in a modern data center, representing the enterprise infrastructure behind AI cybersecurity and governance decisions.
← Back to Blog
AI Governance

GPT-6 Astra Scored 100% on an AI Cybersecurity Benchmark. Enterprise Access Defaults to Off.

Head portrait of Alex Goryachev
Alex Goryachev·September 5, 2026·5 min read

OpenAI's GPT-6 Astra scored 100% on the ExploitBench cybersecurity benchmark, up from 78.5% for the model before it, and OpenAI shipped it with enterprise access turned off until an administrator opts in.

Key Takeways

  • OpenAI's GPT-6 Astra scored 100% on the ExploitBench cybersecurity benchmark, up from 78.5% for GPT-5.6 Sol, and found 2 previously unknown zero-day vulnerabilities during OpenAI's own testing.
  • Astra ships with enterprise access off by default, so an administrator has to turn it on, and the public version declines advanced offensive work such as writing proof-of-concept exploits.
  • Cybersecurity analyst Sanchit Vir Gogia noted that the Critical designation is largely a disclosure requirement under OpenAI's own Preparedness Framework, and that the testing methodology changed between the two models.
  • NIST's draft Special Publication 1353, a quick-start guide for using AI in Cybersecurity Framework analysis and reporting, is open for public comment through roughly October 15, 2026.

OpenAI's newest model scored 100% on ExploitBench, a cybersecurity benchmark where the model before it scored 78.5%. During OpenAI's own testing, GPT-6 Astra found 2 zero-day vulnerabilities that nobody had documented anywhere. OpenAI shipped it on September 4. For enterprise customers it arrives switched off, waiting for an administrator to decide.

The company said the hard part itself, in its own announcement: "Astra is a significant jump in cyber capabilities and meets the Critical threshold in cybersecurity under our Preparedness Framework." That is a sentence a company chose to publish about its own product on launch day.

OpenAI's framework has now acted twice on the same threshold

In August I wrote here about that same Preparedness Framework pausing an earlier version of this same model before release, after it crossed this same cybersecurity line. The open question then was whether a voluntary framework does anything at the moment acting on it becomes expensive.

The answer arrived 4 weeks later. Astra goes first to a limited set of organizations, then to ChatGPT Plus, Pro, Business and Enterprise users. Enterprise administrators have to switch it on by hand. The public version declines advanced offensive work, including writing proof-of-concept exploits. Defenders can still use it for secure code review and patching. OpenAI hardened the model against jailbreaks and added misalignment-monitoring classifiers that can pause or stop unauthorized activity on their own. A staged loosening for vetted defenders, which OpenAI calls Daybreak, rolls out over the coming weeks.

"Its ability to identify and develop zero-day exploits can help defenders find and patch weaknesses, but it also creates a need for stronger safeguards." (OpenAI, September 4, 2026)

Now the part that keeps this honest. Sanchit Vir Gogia, a cybersecurity analyst, pointed out that the Critical label functions largely as a disclosure requirement under OpenAI's own framework. Crossing the line obligates the company to say so. The testing methodology also changed between the two models, which makes the jump from 78.5% to 100% harder to read as pure capability. The precise version of this story is the interesting one anyway. A company crossed a line it wrote for itself, published the crossing, and put staged access controls behind it. A governance mechanism did its job, which is rarer than a benchmark score.

The decision moved to an administrator's console

After a keynote, someone from a security team usually finds me with a version of the same question. This month it has a new shape. My leadership will ask about this by Monday, and I need to tell them something. Default-off means a person with a name now owns that answer. They have to decide whether their own defenders get a model that can find zero-days. Whoever is studying the same benchmark results with worse intentions will never wait for an approval workflow.

That decision changes the work underneath it. Secure code review that took a senior engineer a week compresses into an afternoon. That shortens the half-life of skills the engineer spent a decade building. The company that answers with retraining keeps the judgment it spent those years accumulating. The company that answers with a smaller security team keeps the tool and loses the people who can tell when the tool is wrong.

Rules written for a slower world

Most corporate security policy runs on an annual review cycle with a quarterly exception process. Astra's access rules have moved twice in about 4 weeks, and OpenAI says Daybreak will move them again shortly. Rules written for a slower world is the phrase I keep coming back to on this blog's AI Governance pillar, and this pattern is exactly what it describes. A company's own rule-making cycle runs slower than the products it writes rules about. The policy arrives describing something that has already changed shape. A benchmark score is a capability. A default setting is a decision.

NIST has a draft out for public comment right now, the same body whose work I covered here two days ago on agent identity standards. Special Publication 1353 is a quick-start guide for using AI in Cybersecurity Framework analysis and reporting, open for comment through roughly October 15. That process is doing a different job from a vendor's access switch. OpenAI's control governs who can use one product this quarter. NIST's guide will shape how thousands of organizations, with different budgets and different risk, do this work in a way they can be audited against. Something built to hold up that broadly has to be built in public. The comment period is also the part practitioners can contribute to directly. If your security team has a view on how AI belongs in CSF analysis, the record is open until October 15. A comment from someone doing the work lands differently than another vendor white paper.

The questions worth asking before you switch it on

I am taking 3 questions into the AI governance work I do with the California State University system this month. They transfer to any organization facing the same switch. Which of our defenders gets this first, and what will they do with it in the first 30 days? If our engineers patch faster now, who checks what the model claims to have found? And what did we budget this year for the retraining that keeps those people sharp enough to check it?

Multiply those answers across a few thousand organizations and they settle something no single company votes on. The hospital system and the utility running a regional water plant sit on software carrying the same undiscovered flaws as everyone else's. Neither has OpenAI's testing budget. Staged access decides which defenders get this capability first and how long everyone else waits. That is a distribution question, and it deserves an intentional answer.

The 100% belongs to the model. The decision belongs to a person with a name, a console and a calendar, and this time that person works for you. In August the question was whether a company's own framework would do anything when doing something cost it money. It has now done that twice. The question for your next security review is narrower. When your defenders can use a tool this capable, what have you funded so the people using it can still judge its work? Take that into the meeting and see who has an answer. If you see it differently, I am easy to find.

Is GPT-6 Astra turned on for my company by default?

Access defaults to off for enterprise customers, so an administrator has to opt in before anyone on the team can use the model. OpenAI is also releasing it to a limited set of organizations first, ahead of ChatGPT Plus, Pro, Business and Enterprise users.

What is OpenAI Daybreak?

Daybreak is the program OpenAI says it will use to loosen Astra's restrictions in stages for vetted defenders, rolling out over the coming weeks. The public version of the model declines advanced offensive tasks by default, while permitting defensive work such as secure code review and patching.

What does the Critical threshold mean in OpenAI's Preparedness Framework?

It is the classification OpenAI applies when a model's capability reaches the level its own framework says requires disclosure and stronger safeguards. Cybersecurity analyst Sanchit Vir Gogia noted that the label works mainly as a disclosure trigger, and that OpenAI's testing methodology changed between GPT-5.6 Sol and GPT-6 Astra.

Here is what makes Alex a credible voice on this topic: Alex advises the California State University system on AI governance, where access decisions like the one Astra just handed enterprise administrators are the daily work.

If your security leadership is weighing what to allow and when, book a conversation →

← Back to Blog
Head portrait of Alex Goryachev
Alex Goryachev

WSJ-bestselling author · Former Managing Director of Innovation, Cisco · Advisor, CSU AI Working Group · LinkedIn Top AI Voice

Work with Alex

Bring this thinking to your organization

Alex works with executive teams at global enterprises on AI strategy, governance frameworks, and organizational readiness. Available for keynotes, C-suite workshops, and advisory engagements.