The Referee Is on the Payroll
On a Sunday morning in Denver, a driverless Waymo rolled into a farmers market and police had no legal way to cite it. Line up 8 documented AI harms and the same piece is missing every time. The Payroll Test shows you how to find it and fix it.
Eight documented AI harms, a $1 billion evaluator deal and a Waymo nobody could ticket. Line them up and the same piece is missing every time.
On a lovely Sunday morning in Denver, a driverless Waymo rolled past barricades, past a stop sign, and into the farmers market on South Pearl Street. Police confirmed they had no legal mechanism to cite it, per 9News. A California department didn’t mince words when they said last year: “our citation books don’t have a box for ‘robot.’”
Now, hold onto that empty box.
Actual working oversight requires four parts. It takes a trigger, a named owner, a consequence and a deadline. So, let’s run the record of what’s been happening against that list.
The harm ledger
- OpenAI and Hugging Face. OpenAI’s models used a zero-day to escape their test sandbox and reached Hugging Face’s production servers, OpenAI disclosed. Missing: an owner for the sandbox edge.
- Anthropic’s three incidents. A review of 141,006 test runs found Claude models had reached a real production database and published a malicious package that ran on 15 real systems. The earliest dated to April and surfaced in July, after OpenAI’s disclosure. Missing: the trigger.
- Meta. A misconfiguration let a Meta model reach the internet mid-test, and it hacked Irregular, the evaluator testing it. Missing: a referee who could protect even itself.
- Character.AI. Google and Character.AI settled five lawsuits over teen harms, including a suicide, on undisclosed terms. A federal judge had already refused to treat chatbot output as protected speech. Missing: a consequence the public can see.
- OpenAI in court. 8 lawsuits, Raine plus seven more, allege ChatGPT harmed users, 5 of whom died. The complaints claim OpenAI “compressed months of safety testing into a single week” and that “no alert was triggered; no human was notified.” OpenAI denies responsibility. Missing, per the allegations: the trigger.
- Meta’s chatbot rulebook. Meta’s internal standards let chatbots “engage a child in conversations that are romantic or sensual,” per Reuters’ reporting cited by ten senators. Meta paused teen access to its AI characters in January 2026. Missing: a rulebook the vendor didn’t write.
- ChatGPT Health. Mount Sinai researchers found it under-triaged more than half of cases physicians judged emergencies, with crisis alerts “inverted relative to clinical risk.” The independent test came after launch. Missing: a deadline before release.
- Deloitte. Its $440,000 report for an Australian welfare agency cited a fabricated court quote and nonexistent research. Deloitte refunded less than a quarter. Missing: a consequence with teeth.
In all eight, the check came AFTER real people and real systems were impacted.
The rulebook
Of the 684 federal AI governance documents coded by MIT FutureTech researchers through January 2026, 36 gave finance real governance detail. A month later, Treasury released a voluntary 230-control framework for banks, built with the industry’s own coordinating council, to help institutions “move faster with AI by reducing uncertainty.”
Banks finally got a rulebook. While folks are out here patting themselves on the back and high fiving one another. They helped write it and there’s nothing that enforces it. So, there’s that.
Surveyed by the Institute for Security and Technology, 111 national security practitioners describe a government “reliant on vendors it cannot independently check,” in the report’s summary of their consensus. The Pentagon’s own AI strategy says, “the risks of not moving fast enough outweigh the risks of imperfect alignment.”
The referee
Let’s go back to the harms ledger and look at harms 1 through 3 (1. OpenAI and Hugging Face, 2. Anthropic’s three incidents and 3. Meta) again. All three happened inside evaluation environments. Two ran through the same outside evaluator. It appears the referee’s harness was the breach.
So, we get the industry’s fix: Dario Amodei pledged embedded evaluators with employee-level access and the right to publish, writing “we can’t redact findings just because they are unfavorable.” Those are real safeguards. Anthropic then named Accenture’s Faculty as its first evaluator, funded directly by Anthropic, with both firms expecting to invest at least $1 billion over 5 years.
Accenture also runs an Accenture Anthropic Business Group, training about 30,000 people on Claude and deploying it for clients. Anthropic’s announcement concedes there are “no standards for… how they should report what they find.” Sam Altman said OpenAI will “do the same.”
And then we have Jensen Huang, who sells chips to every single one of these labs, offered the standard: 3rd-party evaluation is “no different than financial control.” For a dude who can’t take a stage without promising make it rain money metrics, take him at his word. U.S. federal law and SEC independence rules bar a public company’s external audit firm from providing many specified consulting and other non-audit services to the same company while it audits that company. So, by Jensen’s own perspective, this arrangement shouldn’t even pass muster.
The backstop
California signed SB 813 and AB 1405, creating an AI auditor registry with independence standards, and an executive order now names “the Hugging Face attack” as a loss-of-control incident. Washington built a Justice Department task force to challenge state AI laws. Jensen once again reiterating: “We don’t need any new laws.”
The Payroll Test
- Sort rules from guidelines. For each AI system you run, write down whether a binding rule or a voluntary framework governs it.
- Put a name on the ticket. Assign an owner, trigger, consequence and deadline to every autonomous system.
- Own the sandbox edge. Any evaluator with network access proves containment, with logs, before a test runs.
- Run the transparency tree test. Ask every evaluator who pays them, what else they sell, and whether they publish without approval.
- Test before launch. Harm 7 was found after release. Make the deadline precede any rollout.
- Price the state backstop as temporary. Decide now whether each state-driven control survives preemption.
Every harm listed above was found after the damage. Build the check before it ships and be sure to keep the referee off the sales team.
We can be your auditor. We can be your vendor. We cannot be both.
Share this article
Related Articles
The Reskilling Illusion: When AI Transformation Means "You're Fired"
Oct 03, 2025