The Regulator Can't Read the Model

A UK parliamentary committee wants an AI regulator with the power to test frontier models. Its own evidence shows the evaluation capability behind that power doesn't exist yet. Yvette Schmitter breaks down four gaps in the report and a 6-question test any board or legislator can run before calling an oversight body independent.

Yvette
Yvette CEO
October 01, 2026 5 min read

Britain has written itself the right to demand a frontier AI system. It hasn't built the instrument that reads one. Neither has anyone else, and the people who can are the labs being read.

On September 14, 2026, a cross-party committee of the House of Lords and House of Commons published the most serious statement any UK institution has made about AI and the people it affects to date.

The Joint Committee on Human Rights spent 10 evidence sessions and 70+ written submissions arriving at a conclusion its chair, Alex Sobel MP, put it simply: “at present the country is unprepared to deal with AI's consequences, however dire they may be.”

The report's recommendations are solid.

  • A dedicated AI Bill.
  • A risk-based regime with heavier duties for higher-risk systems.
  • Mandatory transparency.
  • Prohibitions on the uses that can’t be squared with human rights.
  • Prior approval for the high-risk ones.
  • Obligations that run the length of the supply chain rather than landing only on the person who switches the system on.; and
  • At the center, an independent oversight body on a statutory footing, with the power to test and evaluate AI systems and to order dangerous ones off the market.

And every headline ran the same line: MPs call for an AI Bill and a new regulator. Each solid, accurate, but stops one step short of the finding that actually matters.

When you read the report against itself, the gap larger than the Grand Canyon exposes what the coverage missed. The Committee recommends a body that can read frontier models. Its own evidence, strewn across the same pages, says the country doesn’t have the instrument that does the reading.

The power it recommends and the capability it records

The prize recommendation asks Parliament to give the AI Security Institute (AISI) a statutory duty to review powerful models, and to require developers to submit new models and new versions for evaluation and testing, along with their technical specifications and training data.

Solid. Now hold that against the evidence the same report collected. The AI Security Institute engages with frontier developers voluntarily. It has, in the report's own words, no regulatory authority and can’t demand access to foundation models before or after deployment. The Equality and Human Rights Commission, the regulator charged with equality and non-discrimination, told the Committee it doesn’t have the technical expertise on AI that other regulators have. The Ada Lovelace Institute went further, calling the EHRC "critically under-resourced to address AI." Not hyperbole. In its own words to the Committee, the regulator said its ability to work on AI is limited by its resources.

So, the report is asking for a right of access it doesn’t currently have, to be exercised by an evaluation capability the country hasn’t built yet. Granting the right to demand a model is one thing. Possessing the laboratory that can tell you what the model will do is straight up another. Sure, a perfectly worded statute can create the first with a clause but the second requires people, process, technology, and money, of which the report budgets none.

This is the point I’ve been beating the drum on from the start. Independence has layers. A body can be independent in who appoints it and still hold a borrowed ruler, because the methods it measures with, and in some cases the access it depends on, come from the same labs it’s meant to grade.

Access isn’t a ruler. A right of entry tells you absolutely nothing if you can’t read what you find inside.

The manifestos of word salad converge on who writes the rules, and not so quietly agree on who measures

While the UK is debating its report, the people who build the models published their own answers, and folks, the timing of all this really matters.

Over the course of 5 weeks this summer, the 3 leaders racing hardest to build the next frontier each set out a governance framework.

Anthropic: Dario Amodei reversed his company's long-held transparency-only position and called for binding rules modeled on the Federal Aviation Administration (FAA): mandatory testing of frontier models by a qualified 3rd party, with the government able to block a release that fails. Read the mechanism closely. The 3rd-party evaluation, per Dario, could be done by a government agency or by a set of private organizations authorized and inspected by the government. He calls this a regulatory-markets approach. Then on September 18, Anthropic named its first embedded evaluator: Accenture, the same firm it already pays as a multi-year Claude-deployment partner. Employee-like access to the models, handed to a company already sitting on the other side of a commercial deal.

Google DeepMind: Demis Hassabis proposed a US-led Frontier AI Standards Body modeled on FINRA, the industry-funded organization that polices Wall Street under government oversight. The same industry that missed Bernie Madoff. Labs would submit models for review of up to 30 days before release, and the body would grow an ecosystem of 3rd party auditors.

OpenAI: Sam Altman welcomed a federal framework that sets consistent safety requirements, while making clear the company would not wait for legislation to act. His own record cuts against mandated pre-clearance: in Senate testimony he warned that requiring developers to pre-clear systems before deployment would be disastrous and was skeptical of mandated testing regimes.

NVIDIA: Jensen Huang argued for a single federal standard over a patchwork of state laws, and for building in the open. Do it in the open, he said. Don't do it in a dark room and tell me it's safe.

Yann LeCun, having left Meta to build world models at his own startup, remains the field's strongest voice for open research and warns that locking AI behind proprietary walls cedes ground to China.

5 leaders (and it’s not lost on me that they are all men), 5 emphases. But as soon as you strip the labels, there’s one structural fact survives in every single machination. The technical capacity to evaluate a frontier model sits inside, or right next to, the labs that build them.

Regulatory markets, a FINRA-style body, a federal framework the labs help design: each keeps the ruler in the same building that makes the product. The market has already said this out loud (Anthropic partnering with Accenture on embedded evaluation). Anthropic's first embedded evaluator is a firm it already pays to sell Claude to enterprises.

If the only people with the “expertise” to evaluate frontier models are the labs building them, one widely shared analysis asked, is an industry-funded standards body a necessary compromise or an inevitable capture? And that’s the question the UK has to now answer for itself, because the frameworks on offer globally all resolve the same way, back to the labs with evaluators they pay.

Four things any readiness check has to survive

If we take the report’s recommendations at face value and hold them against what execution would demand, there are 4 gaps that immediately rise to the surface.

  1. Access without capability. The requirement that developers submit models for evaluation presumes an evaluator. The report's own observers say the UK's evaluation capacity is thin, and in the equality regulator's case, self-described as absent. A requirement without a capability is vapor, a paper power.
  2. The reference model moved its own clock. The report leans on the EU AI Act as its risk-tiering template. Between the evidence sessions and publication, the EU deferred its high-risk obligations from August 2026 to December 2027, through the Digital Omnibus, because the harmonized standards and the compliance ecosystem weren’t ready. So, the template the UK is told to copy has already kicked the can down the road 16 months on exactly the provisions the UK would lean on hardest. Sure, they can copy the approach and guess what? The UK imports the unsolved standards problem, not a finished rulebook. Protection dated to a standard nobody has written yet has a shelf life, and the clock belongs to whichever standards body finishes last.
  3. The government's live instrument points the other way. The Committee wants protections built up across the lifecycle, and its live bill, Regulating for Growth, and live program, the AI Growth Lab launched in June 2026, are built to temporarily switch protections off inside supervised sandboxes to speed adoption. AI Minister Kanishka Narayan told Reuters the UK would consider regulating advanced models only if the current system of voluntary agreements with vendors proves inadequate. So, the report recommends a floor while the government operates a dial, and this year the dial is turned toward disapplication.
  4. The authorship hook the report left on the table. In a single sentence, the report notes that the author of the AI Opportunities Action Plan, Matt Clifford, has since joined the AI firm Anthropic. It records the fact and quickly moves on, as if there’s nothing to see. However, it’s the same exact pattern running through every manifesto above: the people proposing the rules build the systems the rules govern. The report makes the statement but doesn’t connect it to the reason its own recommendation needs measurement independence written in from the first draft.

So, is the UK really ready?

No. Ok, that reads harsh, and don’t shoot the messenger because their report in excruciating detail demonstrates it.

The UK is not ready in 3 specific named areas:

  1. Authorship: The plan that shaped the UK's posture was written by someone who now works for a frontier lab, and the global template is being drafted by the labs themselves.
  2. Measurement: The UK has the right to demand access on paper and, by its own evidence, lacks the instrument to evaluate what it would see.
  3. Timing: The EU template slipped to December 2027 for want of standards, and the UK's own live program is designed to disapply rules on a timer.

So yes, I’ve said, as did they, that they are not ready. So what’s required to make them ready? The UK doesn’t have to drink the ocean to know it’s salty. There are very specific and narrowly scoped actions.

  1. A statutory oversight body with its own funded, staffed evaluation capability.
  2. Methods and benchmarks it doesn’t inherit from the labs it grades.
  3. A constitution that keeps the suppliers of those methods, and the people who run them, separate from the parties being graded.
  4. And a clock the reviewer controls.

The report recommends the body but doesn’t secure the four that makes it work. This is absolutely and categorically NOT a criticism of the Committee’s work. The intent is solid and the urgency is warranted and it’s a caution about the distance between a Bill and a capability.

You can legislate a right of access, but you can’t legislate a laboratory into existence with a clause. The government has approximately 2 months to respond, and the test is whether it engages with the gap or restates its preference for regulation at the point of use by existing regulators (the very regulators who told this Committee they lack the expertise).

Again, this is not a bashing of their work. To borrow their own closing summary and drive the point home, the Committee chose this witness statement: “the only growth worth having is one that protects human rights.” That growth is available to the UK. It will need a regulator that can read the model to get it, and right now it is describing one rather than building one.

The Borrowed-Regulator Test

6 questions a legislator or board can put to any proposed AI oversight regime anywhere on the planet. Use these questions BEFORE calling a body independent.

  1. Capability, not just access. Does the body have its own funded, staffed ability to evaluate a frontier model, or only the legal right to ask for one?
  2. Method provenance. Where do its evaluation methods and benchmarks come from? If the answer is the labs, it is holding a borrowed ruler.
  3. Payer and employer. Who funds and employs the assessors? Trace it to the source before you trust the badge.
  4. Clock ownership. Does the reviewer control the timeline, or does protection begin only once a standard nobody has finished is finally written?
  5. Separation of supplier and subject. Are the parties who supply the methods, benchmarks, and personnel walled off from the parties being graded? (Here’s a question: How will Anthropic wall off Accenture and vice versa?)
  6. Direction of travel. Is the government resourcing the floor it recommends, or operating a dial that disapplies it?

An establishment that fails 3+ or more of these isn’t a regulator. It’s a relationship with vendors that has government letterhead on top.

Which is the whole reason this work exists, and the line I will close every piece on.

We can be your auditor. We can be your vendor. We cannot be both.

Related Articles