The Auditor Who Doesn't Sell the Models

Yvette
Yvette Managing Partner
September 02, 2026 5 min read

FAR.AI graded the four frontier flagships head-to-head and found a hundredfold safety gap. The finding is real. Your risk lives in the fine print.

For once, the people grading frontier AI safeguards don’t sell frontier AI. And that single fact makes FAR.AI's new AI Security Leaderboard the most useful safety document your board will read this quarter.

Here’s what the independent Berkeley nonprofit did. It tested the four leading models under identical conditions across chemical, biological, radiological, nuclear, explosive, and cyber threats.

  • Two of them, Claude Fable 5 and GPT-5.6 Sol, held against every attack in the suite.
  • Two of them, Grok 4.5 and Gemini 3.1 Pro, came apart. A working universal jailbreak, one that unlocks a whole category of dangerous requests, cost roughly $58 to find on Grok and roughly $278 on Gemini.
  • And on the two that held, the same search never landed, which puts the floor above $14,200.

FAR.AI calls it a spread of more than a hundredfold between the sturdiest and the flimsiest systems on sale today.

And that gap? It buys an attacker one thing: choice. When one model refuses, the request just walks next door. A safeguard that stops a bad actor at one company does little to nothing for anyone if the same prompt succeeds at the next.

When “if it bleeds it leads” determines headlines, this one made it. But it’s the story underneath all of this is why this report exists at all, and what a clean score can and can’t certify for the people writing the checks.

Independent measurement finally showed up

Readers should know where Fusion Collective plants its flag: No party can grade its own homework.

The safety claims that reach your desk almost always start with the seller: the system card the lab wrote, the evaluation the lab ran, the risk label the lab handed itself.

FAR.AI did the thing this market keeps ignoring: it built its own standard, its own attacks, and its own scoreboard, and it published a ranking the vendors do not control. And the end result carries a considerable amount of weight. Why? Well, because the tester has no model to sell. So, I’m giving them the flowers they are due. This is what independent measurement looks like, and almost nothing else in the category qualifies.

However, the ruler is still partly borrowed

In my article, "The Ruler Belongs to the Labs," I drew a line that this report actually puts to the test. Independence of authorship is a separate property from independence of measurement. FAR.AI wrote the report free of the labs. However, the instrument it wrote with still leans on the same family of systems it is judging. A jailbroken model generated the attacks. Open-weight models rated them. A GPT-class model scored every transcript.

But FAR.AI is candid about all of it and calibrated the pipeline against 600 human labels, publishing the error rates in the open:

  • 89% agreement
  • 2.3% false-positive rate; and
  • 23.6% false-negative rate.

So, the authorship is independent and the measurement is a model grading models with people checking the seams. Both facts belong in your read. The second one rarely survives the summary.

What a clean score actually certifies

If you look closely at "held against every attack," there’s three limits that sit inside it, and FAR.AI names all three.

First, the two models that held were tested in a later round and never entered the human-validated set. That set covered their predecessors. Their clean result lines up with those earlier models, which also showed nothing, and it wasn’t checked directly on the versions now carrying the badge.

Second, the evaluator misses. On the models where jailbreaks existed, it failed to flag close to a quarter of the real ones. Here’s what the means for an exposed model; that means the true count likely runs higher than the headline. For a clean model, that same blind spot is the reason "none found" reads as a ceiling rather than a guarantee.

Third, the suite is static by design. It leaves out the adaptive, self-tuning attacks that tend to work best against the hardest targets, which FAR.AI defers to future work. "Held against every attack we ran" is an honest sentence with a boundary drawn around it.

And now, when you put it altogether this is what you get:

  • The exposure of Grok and Gemini is demonstrable, repeatable, and probably undercounted.
  • The safety of Claude Fable 5 and GPT-5.6 Sol is a floor: well-built, well-earned, and bounded by those three limits.

Don’t get me wrong, a floor is an achievement. So, treat it as the starting line for your diligence.

The gap is a choice, and that’s the whole damn accountability story

There’s one finding that deserves a spotlight that it didn’t get and it’s this: every model in the test refuses these requests on its own, with no attack applied, and every model recognizes the goals as harmful.

So, what does that mean? Well, the gap traces to one thing: engineering investment.

FAR.AI names the defenses behind the clean scores, points out they are publicly described, and notes that competitors already ship them. Now, the exposed vendors could build the same protections. But they have chosen, so far, not to.

And for a board? Well, that reframes the whole exercise. Safeguard strength is a decision a vendor makes, which means literally means YOU can price it and write it into a contract.

Why you can act on any of this

Well, because the messenger is independent of the labs where it counts, in its funding.

FAR.AI runs on safety philanthropy, not on lab contracts, which is the one property a vendor's own safety page can never really claim. It’s also the property the law declined to require. California's SB 53 kept the disclosure mandate and let SB 1047's 3rd-party audit requirement fall away.

So, the single independent audit in the room showed up via a nonprofit's initiative and on no one's order. Remember when I said, “match the promise to the messenger's record?” Well, this messenger's record supports the promise. And remember what was it that actually stopped the worst potential outcome of the AISI reported hack? The defense wasn’t a reviewer sitting inside the system, gating the agent's every move. The defense was a human standing out on the receiving end, an open-source maintainer doing Yeomans’ work; who got suspicious and said no. The person who stopped this was oppositional to the attack and external to the lab, standing right where the attack was aimed, doing unpaid work, already buried in AI-generated, vibe-coded slop. And that fragility is worth naming once again because the entire supply of independent measurement in this market rests on one organization choosing to do the yeoman’s work.

The Clean-Score Stress Test

My recommendation is for you to read the leaderboard as the strongest opening question in the category and then ask these six it was never built to answer. Run these before you trust any green light - on FAR.AI's scoreboard or your vendor's slide.

  1. The Provenance Test. For any safety score you are shown, name who built the test and who funded them. When the tester sells the model, mark the score self-attested and weight it accordingly. FAR.AI clears this bar. Most system cards do not.
  2. Version Match. Confirm that the exact model version in your contract is the one that was validated, not an earlier sibling. A clean score inherited from a predecessor is encouraging but it’s 3 quarters short of a dollar for a clean score on the model you are actually deploying.
  3. Miss-Rate Read. Ask for the evaluator's false-negative rate and read "zero findings" through it. A tester that misses a quarter of real attacks is telling you that "none found" is a floor. So, size your comfort level to the floor.
  4. Static Versus Adaptive. Ask whether the testing included adaptive, self-tuning attacks or only a fixed suite. The fixed suite is the bare minimum bar. The adaptive attacks are the ones motivated actors actually run. So, get the difference in writing.
  5. Defense-in-Depth Receipt. Ask your vendor to name its layers: reasoning monitors, activation-based checks, independent input and output filters, instruction-hierarchy training, and testing of combined attacks. Because one guardrail is one point of failure and that’s a reason why planes have powerful engines so if one goes the plane can still fly, your VPC has a failover and you pack a carry on with your prescriptions and clothes you can wear just in case your checked luggage doesn’t land when you do.
  6. Checkbox Guard. When a vendor markets "meets the Minimal Standard," remember what that means: a floor that publicly known defenses already clear. So, if you choose to reward it as a floor cleared just make sure you keep asking for everything stacked above it.

The leaderboard is a gift to anyone who must make a real decision about AI safety, and it’s a gift because an outsider built it. So, take the finding and thank the people who did the work. Then carry the fine print into every purchase and every board meeting. Treat the clean score presented to you as a floor you can build on and the exposure as a signal you can price. And for the love of all things Dolly Parton, remember why you can trust the scoreboard in the first place. Because someone with no model to sell held it up.

That’s the line Fusion Collective was built to hold.

We can be your auditor. We can be your vendor. We cannot be both.

Related Articles