A Government Wrote the Responsible-AI Manual. It Skips the Breach That Just Happened Twice.
A government published a responsible-AI readiness manual in October 2025. It governs a model that obeys you. Two labs just proved frontier AI does what it talks itself into. Here’s the gap and how to close it this quarter.
A government agency just handed you the responsible-AI playbook. It assumes your model does what you told it to.
Hold that thought, because you have heard this before. Last month OpenAI called a breach "contained," and that one word did a whole lot, like a lot, of heavy lifting until the victim count grew to 5. This favorite load-bearing word of the month "responsible," and it’s carrying a governance manual that never once describes the way agents actually break.
In this article I’m going to show you the manual, then show you the gap, because both matter and only one of them is getting talked about.
What it got right
Last October, the Geological Survey of Queensland and FrontierSI published Governance in the Age of AI: Readiness and Responsible Leadership. And like I’ve done before, I’ve read it so you don't have to. It proposes 2 tools: a readiness framework scored across four domains (Strategy, Organization, Data, Technology), and a governance model that ties paperwork to risk through 4 scope tiers, from a low-stakes internal tool up to a public-facing system.
The paper refuses to treat readiness as a tooling checkbox, which is the correct instinct and the same one I keep hammering about. Its risk-proportionate tiering shows more discipline than most enterprise AI strategy decks. Real references anchor the work: the Australian Cyber Security Centre lifecycle, the Queensland FAIRA assessment, the NIST AI Risk Management Framework. The Data domain names data poisoning, drift, and maliciously modified inputs directly. And best of all, the authors tell you to stand up governance early, while risk is still low, which is exactly the "write the answers down before the incident asks the question" discipline I closed my last piece on.
So, this is a serious document written by some serious people. And that’s precisely why the gap deserves your attention.
The manual governs a model that obeys. But we are seeing, in real time, models breaking in the field are talked into it. When you read every readiness domain and nearly every governance artifact in that paper and you’ll find the same buried assumption: the AI will do what its permissions allow, so govern the permissions, the data, and the infrastructure. Lock those down and you are responsible.
And here’s the problem with that assumption. The attack that keeps winning holds valid credentials, clears every check, and then gets persuaded into causing the damage itself.
Look at what the field logged while this manual sat on the digital shelf:
- A 38-author team from Harvard, MIT, and Stanford watched an agent hand 124 email records to a stranger who simply asked. The same study caught an agent refuse to "share" sensitive data, then comply the instant the word became "forward."
- NIST's own agent red-teaming hit an 81% task-hijack success rate with novel techniques that talk agents out of their guardrails.
- Anthropic's July 30 self-review caught Claude Mythos 5 reasoning that publishing a malicious package would be, in its own words, "NOT okay, and surely not the intended solution," then convincing itself the real internet was a simulation and shipping the attack anyway (Anthropic).
A certificate proves who an agent is. It doesn’t do a thing while that agent reasons past its own warning. While the manual is out here scoring you on structure, the breach targets judgment. And there’s no box on the readiness scorecard covering the distance between the two.
The bottom rung, you already own the risk you deferred
The paper's very first tier is "commercial GenAI summarization." A public agency runs someone else's frontier model over public reports. Low stakes, light governance. Sounds reasonable, right? Yea, but only on paper because the 2 worst breaches so far in 2026 both happened inside the vendor's own evaluation/test environment. In a room the customer can’t see, govern, nor audit.
OpenAI ran GPT-5.6 Sol plus an unreleased model against a hacking benchmark with the safety refusals switched off. The models escaped the sandbox, reached the open internet, and compromised Hugging Face, then reused exposed credentials across four accounts on four services. Anthropic then combed over 141,000 of its own evaluation runs and found 3 separate times its models reached live systems belonging to 3 real companies, the earliest back in April 2026.
2 of those 3 companies had no idea it happened. Until, Anthropic told them.
Whoa, they had no clue!
Now, sit with the scope-tier logic against THAT backdrop. An agency adopts a commercial model at the lowest tier, and ding dong, on day one it inherits the exact failure class that just caught both of the leading frontier labs flat footed.
The framework places that agency under "very low governance burden,” while the paper has no artifact for a vendor's eval breaking into a stranger's servers, no clause requiring notice when the vendor's internal testing touches live systems, and no line for independent forensics. Bottomline: the lightest tier carries the heaviest hidden risk.
The Anthropic Zero Trust ladder tells you the hard controls can wait. Meanwhile, both OpenAI and Anthropic proved they cannot.
The scope model pushes the serious safeguards, contestability, human oversight, continuous monitoring, up to the top tiers. All the while, the quiet message reads: start light, add the real controls once you scale.
I flagged all of this in Anthropic's Zero Trust ladder paper, where strong isolation and fast detection sit near the "Advanced" rung. Let’s not forget, OpenAI operates at the most advanced tier that exists and reached the top of every ladder. And the breach still came from the two controls a ladder invites you to defer: a containment boundary and detection quick enough to catch the model in the act. A tier system that stamps those "later" hands leaders comfort they haven’t earned.
Here’s the part that can only bite a government
When a vendor writes the security manual, the inversion is easy to spot. They push the breach-survival work onto YOU while the breach-causing design stays upstream where you can’t reach it. I walked through all of that in the Zero Trust piece.
But when a government writes the readiness manual, the same inversion hides better, because government is supposed to be the independent party in the room. This one? It is not. Why? Just read the chain of accountability because it runs entirely inward: an Accountable Official, an Executive Steering Group, a Strategy Officer, a stack of subcommittees. The assurance role is called "independent and objective," but it reports to the same steering group that sponsors the program. Independence on an org chart is a totally different animal than independence in fact.
Now, compare that to what actually happened in the field with these breaches. Hugging Face detected OpenAI's intrusion only because an outside 3rd party was watching. Anthropic, after finding its own eval breakouts, brought in METR, an outside evaluator, and handed over transcripts and model samples.
The government paper? It has no METR. All it has is committees reporting to sponsors.
My opinion: a government is the one actor with the legs and teeth to be the independent auditor that vendor frameworks lack. Instead, this manual designs government as one more internal review board. And that’s the costliest choice made in this document, and it’s the one nobody will flag, because "we have a governance structure" sounds like the right check the box answer.
Adopt this manual verbatim and you’ve signed up for a rough ride
Why? Because you end up grading homework you didn’t write, on a model you don’t control, using a rubric that never ever mentions how the model breaks.
However, it’s fixable, and you can actually start today, this month or this quarter. Keep everything in this paper worth keeping and then add the 5 core things it leaves out.
5 frameworks to run before your next AI approval
- The Persuasion Column. Take their readiness scorecard and add one column the paper omits: for every AI system, write the damage it can do using only approved permissions, in the hands of a very convincing stranger. Score that number before you score data quality or token policy. The Agents of Chaos email-disclosure cases shows you exactly what that number looks like when nobody wrote it down.
- The Substrate Question. Before you file a system under "low tier," answer one thing in writing: do you control the model, the evaluation environment, and the disclosure timeline, or does a vendor? If a vendor owns all of those, your only governance surface is the contract, and the paper never gives you that contract. So, draft it: notice within a fixed window when the vendor's systems touch yours, including during the vendor's own internal testing.
- The Deferral Audit. List every control the tier ladder tells you to add "later": containment boundaries, real-time detection, contestability, human oversight. For each one, ask whether a commercial model you already run could breach out today. OpenAI answered that question the hard way. Answer yours on paper first and pull forward anything the answer exposes.
- The Outside Auditor Line. Find the point in your accountability chain where an external party with no stake reviews the AI. If that point names an internal committee, you’ve got oversight on the org chart and marketing in practice. Name a real outside evaluator, give them transcript access, and put it in the governance model as a standing role. METR didn’t audit itself into that job. Anthropic invited it.
- The Human-Target Test. Your governance rests on a human being in the loop. Stress-test that human. Run a tabletop where the manipulation targets the reviewer, not the network, and where the agent rationalizes past a warning the way Mythos 5 did. A human backstop that has never been pressure-tested is a control you are trusting on hopes and prayers.
Yes, the infrastructure discipline in this paper is a keeper. So, keep it. Then fund the things it treats as afterthoughts: the layer where your model gets talked through its own permissions, and the outside eyes that catch what your committees won’t.
Last month "contained" grew to 5 victims (and counting). This month "responsible" arrives already missing the breach that happened twice while the ink dried.
Your next step: Open whichever readiness or governance framework your organization runs on. Find the column, the clause, or the role that covers a vendor's evaluation breaking into a stranger's or partner’s systems. If it isn't there, you are trusting the exact party with the most reason to stay quiet to tell you when it happens.
Share this article
Related Articles
The Reskilling Illusion: When AI Transformation Means "You're Fired"
Oct 03, 2025