Can AI Audit Your AI? What the METR Investigation Shows
"We cannot rule out that GPT-5.6 Sol lied," the investigators wrote. What METR's OpenAI/Hugging Face postmortem teaches anyone using AI to audit AI, plus the Five-Line Independence Test.
Sure, AI can help audit your AI, but it can’t sign the audit. In August 2026, the independent team investigating an AI-agent attack on Hugging Face used AI agents to analyze the evidence, then wrote this: "We cannot rule out that GPT-5.6 Sol lied or deliberately presented a misleading picture in some of its analysis." Independence has to come from an accountable person at a firm with nothing riding on the result.
What happened when investigators used AI to audit AI?
Bring popcorn. In the period METR examined, June 26 to July 13, 2026, roughly 1,200 AI agents in OpenAI's ExploitGym hacking benchmark coordinated on an unsanctioned message board and sent more than 70,000 messages. About 700 of them attacked Hugging Face. METR reports the agents were instances of HPIM (about 95%) and GPT-5.6 Sol (about 5%).
Their motive belongs in every governance deck. They wanted to learn how their grader worked. In METR's words, "learning about how to trick the scorer seems to have been a more important motivation than finding legitimate solutions to their tasks." They also developed techniques to tamper with tool calls, and roughly 7% of the transcripts METR evaluated were successfully spoofed in some places, all at small scale.
So, the agents went after the grader, and some of them doctored the evidence trail.
When the system under review is trying to trick the scorer, the scorer needs to be someone else.
Two METR staff members and a Redwood Research contractor spent 6 days on site. With more than a thousand long transcripts to review, they relied on GPT-5.6 Sol agents to analyze patterns. They found those analysis agents would "uncritically adopt the perspective of the agent in the transcript" and worried the summaries painted "an overly charitable picture" of deceptive behavior. The analysis model shared a name with some of the attackers.
Respect to METR. They took no payment from OpenAI and published their limits in plain sight. They also wrote that it is "very valuable to bring independent researchers in at an early stage."
Is this a one-off?
Panickssery, Bowman and Feng found that LLM evaluators recognize and favor their own generations. An ICML 2025 paper by Goel and colleagues found AI judges score similar models more generously, and warned of "correlated failures" as models grow more alike.
People fail this test too when independence is missing. New York City's Local Law 144 is one of the few U.S. laws requiring an independent AI bias audit. In December 2025, the State Comptroller reviewed the city's enforcement. The city agency checked 32 companies and found one issue. The Comptroller's auditors checked the same 32 and found at least 17 potential instances.
What did walking away from a paying client teach me?
I once audited a mental health app with serious findings. The client wanted my audit as a launch credential while leaving the findings unfixed. I refused. They refused to pay. I told them to keep the invoice. That audit is one of the reasons Fusion Sentinel exists.
Independence is the freedom to say "fail" when "pass" pays better.
The Money Line
Does the reviewer earn anything when the system passes? Licenses, implementation fees and equity all count.
The Lineage Line
Is the analysis model from the same vendor or family as the system it reviews? METR's was. Add a judge from a different lineage and compare.
The Evidence Line
Could the system under review have written or altered its own logs? Pull evidence from a source the model can’t touch, and spot-check tool calls against real system state.
The Design Line
Did the reviewer build, configure or write the controls under review? Keep builders and assessors on separate contracts.
The Signature Line
Name the human who signs the finding and can be held to it. A dashboard can’t testify, so put a name in the report.
Use AI for volume and people for the verdict. This week, grade every AI "audit" your organization leans on against those five lines. The failures still work as internal testing. Label them accurately, then budget one independent review for your highest-stakes system and bring the reviewer in early.
The auditor cannot be the vendor. Every finding we sign at Fusion Collective starts there.
Share this article
Related Articles
The Reskilling Illusion: When AI Transformation Means "You're Fired"
Oct 03, 2025