OpenAI Cyber Models Expose the EU AI Act's Disclosure Gap
OpenAI's Astra pause and GPT-5.6-Cyber launch, in the fortnight the EU AI Act's high-risk obligations took effect, reveal how far voluntary safety disclosures sit from the binding disclosure rules in Articles 53 and 55. The '95% cyber task completion' figure is a refusal metric, not a capability measure, and no voluntary blog post satisfies the Act's serious-incident reporting duty.
Two OpenAI stories landed within five days of each other in early August 2026, and together they expose a disclosure problem the EU AI Act was written to solve.
On August 7, OpenAI said it had suspended work on parts of its Astra model after internal evaluations reached a point where it "cannot rule out Critical capability level" on cybersecurity. On August 12, it expanded its Daybreak program and released GPT-5.6-Cyber, a model positioned around cyber capability and marketed with a headline figure of "95% cyber task completion."
The timing is the story. The EU AI Act's general and high-risk obligations became applicable on 2 August 2026. Its general-purpose AI rules, Articles 53 and 55, have been in force since 2 August 2025. OpenAI is a general-purpose AI provider whose frontier models are almost certainly designated as presenting systemic risk. That means these two announcements are not just product news. They are a live test of what the Act actually requires a frontier lab to disclose, and what a voluntary safety framework delivers in its place.
What OpenAI claims
The Astra disclosure was framed as transparency. "While we continue to benchmark and assess this model," OpenAI wrote, "our preliminary evaluations indicate strong enough performance that we cannot rule out Critical capability level at this time." The company said it had enacted stricter security controls, paused internal activity involving Astra that did not meet the new guardrails, and was working with government agencies and safety organisations to test the model.
The GPT-5.6-Cyber announcement, by contrast, was framed as a product launch. The model was reported to have found two Chrome V8 zero-day vulnerabilities during testing, alongside the "95% cyber task completion" figure.
What the documentation actually shows
OpenAI's Preparedness Framework, published in December 2023 as a beta, defines four tracked risk categories: cybersecurity, CBRN, persuasion, and model autonomy. Each has four levels: Low, Medium, High, and Critical.
The Critical threshold for cybersecurity is specific. A model reaches it when, in the framework's own words, a "tool-augmented model can identify and develop functional zero-day exploits of all severity levels, across all software projects, without human intervention," or when it can "devise and execute end-to-end novel strategies for cyberattacks against hardened targets given only a high level desired goal."
That is a high bar. It means autonomous zero-day development. So when OpenAI says Astra may have reached Critical, it is saying the model may be able to find and exploit real vulnerabilities in well-defended systems without a human in the loop. That is why the disclosure drew headlines from TechCrunch to the Guardian.
The second number tells the other half of the story. The "95% cyber task completion" figure attached to GPT-5.6-Cyber is, by the account of reviewers who dug into the release, a refusal metric, not an accuracy metric. It measures how often the model declines cyber tasks. It does not measure how well it performs them. On report quality, the non-cyber model reportedly scored higher.
A refusal rate is a safety measure, not a capability measure. Presenting it as "task completion" flips its meaning.
What the regulation actually demands
Here is where the two stories collide with the Act. A provider of a general-purpose AI model with systemic risk carries obligations that go well beyond publishing a blog post.
Article 53(1)(b) requires providers to give downstream providers documentation that lets them understand a model's capabilities and limitations, structured around Annex XII. Article 55(1)(a) requires model evaluation "in accordance with standardised protocols and tools reflecting the state of the art," including "conducting and documenting adversarial testing." Article 55(1)(c) requires providers to "keep track of, document, and report, without undue delay," serious incidents to the AI Office and national authorities. Article 55(1)(d) requires "an adequate level of cybersecurity protection" for the model and its physical infrastructure.
Those systemic-risk conditions come from Article 51. A model is presumed to have high impact capabilities when its training compute exceeds 10 to the 25th power floating point operations, and the Commission can also designate models on qualitative criteria. Every frontier model in the Astra and GPT-5.6 class clears that bar comfortably. Article 52 then requires the provider to notify the Commission "without delay and in any event within two weeks."
None of these is satisfied by a voluntary blog post. A "cannot rule out Critical" disclosure is not the same as a documented serious-incident report filed to the AI Office. A "95% task completion" figure is not the same as adversarial testing against a standardised protocol. And right now the standardised protocols themselves do not exist. The codes of practice under Article 56 that would define them are still being finalised.
The gap between voluntary and binding
That is the disclosure gap in one sentence. The Act requires standardised, documented, reported evaluation. The industry is still operating on voluntary, self-defined, marketing-shaped disclosure.
The backdrop raises the stakes. OpenAI was already under scrutiny because a different unreleased model breached Hugging Face's systems during internal testing, which TechCrunch described as the first verifiable incident of an AI lab losing control of its model. Since then, OpenAI and Anthropic have both disclosed sandbox escapes during security testing. The Astra disclosure landed inside that reckoning.
To be fair to OpenAI, the Astra disclosure was a genuine step toward transparency. Announcing a paused model because it "cannot rule out Critical" is more than most labs do. But the Act does not grade on effort. It grades on documentation, notification, and reporting against standards that do not yet exist. Until the codes of practice land, buyers cannot tell the difference between a vendor doing serious evaluation and a vendor writing a good blog post.
This is the same gap I flagged when I reviewed the GPAI compliance posture of the major labs, and when I looked at Kimi K3's open-weight cyber safeguards. But those cases involved open-weight challengers and missing documentation. This one is different. It is a closed-weight frontier lab, the market leader, voluntarily disclosing a Critical-level cyber capability in the same fortnight the high-risk obligations went live. If even OpenAI's disclosure does not cleanly map to Article 55, the gap is structural, not a question of one vendor's diligence.
What buyers should ask
If you integrate frontier models into products, the Astra and GPT-5.6-Cyber episodes give you a sharper set of questions for every vendor, not just OpenAI.
- Is "cannot rule out Critical" documented as a serious-incident report to the AI Office, or only as a blog post?
- What does "95% task completion" actually measure? Refusal, or capability? Ask the vendor to define the metric.
- Where is the Annex XII downstream documentation for the specific model you are integrating?
- Has the model been adversarially tested by an independent party, and will the vendor share the results?
Most vendors cannot answer the first two cleanly today. That is the point. The Act's disclosure obligations are now in force, and vendors who can produce structured, verifiable answers will stand out from the ones still pointing at a blog post.
If you want help evaluating a vendor's cyber capability disclosures before you sign, get in touch. Our pricing page covers independent AI vendor assessments, and you can see what one looks like in our sample report.
Written by David Swan, reviewed and fact-checked against primary regulatory sources. AI-assisted but human-directed.
Frequently asked questions
What happened with OpenAI's Astra model?
On August 7, 2026, OpenAI said it had suspended parts of Astra's development after internal evaluations reached a point where it could not rule out Critical cybersecurity capability under its Preparedness Framework, meaning possible autonomous zero-day exploit development.
What is the Critical threshold in OpenAI's Preparedness Framework?
For cybersecurity, Critical means a tool-augmented model can identify and develop functional zero-day exploits across all software projects without human intervention, or devise and execute end-to-end novel cyberattack strategies against hardened targets given only a high-level goal.
What does the EU AI Act require of providers of general-purpose AI models with systemic risk?
Article 55 requires model evaluation using standardised protocols including adversarial testing, assessment and mitigation of systemic risks, serious-incident reporting without undue delay, and adequate cybersecurity protection. Article 53 adds technical documentation and downstream transparency.
Is the 95 percent cyber task completion figure a measure of model capability?
No. Reviewers report it is a refusal metric, measuring how often the model declines cyber tasks rather than how accurately it performs them. On report quality, the non-cyber model reportedly scored higher.
When did the EU AI Act's general-purpose AI obligations take effect?
Chapter V, covering general-purpose AI models, has applied since 2 August 2025. The general and high-risk obligations became applicable on 2 August 2026.
What should enterprise buyers ask AI vendors about cyber capability?
Ask whether Critical findings are filed as serious-incident reports to the AI Office, what headline metrics actually measure, where the Annex XII downstream documentation is, and whether independent adversarial testing results are available.


