Stanford Told Congress What's Broken. Nobody Told Congress How to Check.
In 2024 Stanford HAI & Black in AI handed Congress a rigorous warning about how AI harms Black Americans,then stopped at the one page that mattered: who checks the machine.That blank page is the case for independent AI audits.Because 97% of FDA-cleared medical AI was never tested on a live patient.The company that built the model wrote its own report card,and regulators cleared it on that evidence.The builder became the auditor.That arrangement fails in every industry that tries it.
Nobody Checked the Machine: The Case for Independent AI Audits
In February 2024, Stanford HAI and Black in AI handed the Congressional Black Caucus a 25-page white paper on how AI affects Black Americans. The scholarship is rigorous and the citations honor the Cite Black Women movement. The authors gave out the flowers and did their job.
Then the paper stops. Right at the edge of the question that matters. The authors say it themselves. The paper is "an educational document, laying out the relevant issues and debates, rather than a set of definitive policy recommendations."
Huh, rewind? The most credentialed voices in the field wrote a warning to Congress and left the enforcement page blank?
I counted the questions the paper asks and never answers.
14.
Which medical decisions should AI assist? Who evaluates equity across demographics? How often should medical AI be re-checked? Who owns oversight? How will these systems be audited?
Asked. Filed. Abandoned.
Here's the thing. Those aren’t research questions; they are audit questions. Researchers describe problems while auditors verify claims.
The number that should end every meeting
Buried on page 14 sits the most important statistic in the document. A 2021 survey found that 97% of FDA-approved medical AI devices were evaluated using only retrospective data.
They were never tested on live patients.
Take a minute to really, really think about that. 97% of the medical AI cleared for use on your mother, your daughter, your community was graded on homework the vendor turned in about its own past performance. Nobody independent touched those devices before they touched patients.
The paper presents this as a sample-size problem and I can tell you, it is not. It’s a structural problem. The entire evaluation regime runs on vendor-supplied evidence. The company that built the model wrote the report card. And you know what? That arrangement has a name. The auditor became the vendor. It fails every time, in every industry, and AI will not be the exception.
Tested on whom?
The 97% stat has an even uglier twin sitting right next to it in the same paragraph, and almost nobody quotes it. The paper reports that many of those approved devices were tested only on limited populations clustered around a few geographic sites. The authors' own conclusion: performance "may not be broadly generalizable."
Translated out of academic-speak; it means a device gets validated on the patient mix at a handful of hospitals. Then it gets approved for everyone. The device/tool was never tested on the population it will be levied upon. If the validation sites skew white, affluent, and suburban, then the first real test on Black patients, rural patients, and Medicaid patients happens in production; on live unsuspecting people.
Uncontrolled.
Unmeasured.
Unconsented.
And we already know how that story ends and the paper goes on to document it in every other chapter. Facial recognition performs worse on darker skin, and Black students got locked out of their own exams. Deepfake detectors work best on white faces.
We also know how it ends across borders, because they have already run the experiment. A Danish company trained an AI on emergency calls from Copenhagen to detect cardiac arrest over the phone, then put it into production in Seattle. Danish-trained ears, listening to American emergencies. When Europe tried to move the same technology to France and Italy, the models had to be rebuilt from the ground up, because a system trained on Danish callers could not hear French panic. Here's the part that should stop every buyer cold: even in Copenhagen, on the exact population it was trained on, the randomized trial found no significant improvement in dispatcher recognition, and the model flagged so many false alarms that fewer than one in five of its alerts was a real cardiac arrest. Retrospective accuracy on the home population did not survive contact with production on the home population. Now imagine the ocean in between.
Now, let’s run the tape the other direction and the lesson sharpens. IBM trained Watson for Oncology on protocols from a single elite New York cancer center. When Denmark's national cancer center tested it against their own oncologists before deployment, it agreed with them 33% of the time. Denmark rejected it. Read that carefully. The Danes did the one thing most American hospitals never do: they tested the tool on their own patients before trusting it with them.
And you don't need an ocean for the training population to miss yours. A sepsis prediction model developed on data from a handful of health systems was deployed at hundreds of American hospitals. When Michigan Medicine independently validated it on 30,000 of their own patients, it caught 33% of actual sepsis cases. It missed two thirds. Hundreds of hospitals had been running it on populations it had never seen, and almost none of them had checked.
A model is a mirror of its training data. Feed it a population that doesn't include your patients, and it will be confidently, precisely wrong about your patients. Accuracy claims without a demographic denominator are not accuracy claims. They are averages hiding the people the tool fails.
So, the question every health system, every school district, every agency should ask before deployment is brutally simple. Tested on whom? Show me the validation population next to my population. If the vendor can’t produce that comparison, the vendor has not validated the product. It has validated a different product for different people and borrowed the certificate.
Race removed. Bias survived.
The paper's own best-case study proves the point. Researchers evaluated a risk tool that health insurers used to decide which patients needed extra care. The developers removed race as an input entirely. Careful design. Every box checked.
However, the algorithm still assigned lower risk scores to Black patients who were sicker than white patients with the same score. Why? Because the data itself carried the discrimination. The tool used cost of care as a proxy for medical need, and the historical record it trained on reflected a healthcare system that spent less money on Black patients. Not because they were healthier. Because of every barrier this country built between Black people and a doctor's office. The algorithm learned that spending pattern and enforced it at scale. Recalibrated correctly, the tool would have flagged more than twice as many Black patients for early intervention.
Understand what that means. You can’t fix biased data by deleting a column. The bias lives in the outcomes the data records, and the model will find it through any proxy available. Cost. Zip code. Utilization history. The pattern is the point of the machine.
Twice as many.
That’s the gap between a checkbox and a verification. And here’s the detail almost everyone misses. The bias was not caught at the design stage. No, it was caught in production, by independent evaluators, examining a system that had already passed every internal review. The paper even notes that these learning systems "change significantly over time" as they ingest new data, with no requirement that anyone re-check them.
A model can pass every test on the 23rd of September and drift into harm by December. Pre-deployment review is a photograph. Behavioral drift is a film. You can’t audit a film with a photograph.
2026 renewed the paper's warranty
If a whitepaper from February 2024 sounds dated, January 2026 made it current again. During the week around the J.P. Morgan Healthcare Conference, the “leading” frontier AI labs walked into medicine. OpenAI launched ChatGPT Health for consumers and ChatGPT for Healthcare for hospitals, with Boston Children's, Cedars-Sinai, HCA, Memorial Sloan Kettering, Stanford Medicine Children's Health, and UCSF already rolling it out. Within days, Anthropic answered with Claude for Healthcare, wired into CMS coverage databases, ICD-10 codes, the provider registry, and PubMed. Microsoft moved Claude into its Foundry platform for healthcare workflows alongside its Copilot line.
Note the venue; these products launched at a finance conference, pitched to investors. Not to regulators.
Now if we apply the whitepaper's own test to the new arrivals. Who validated them? You answered correctly if you said the vendors.
OpenAI cites development input from more than 260 physicians and physician-led testing on internal benchmarks. Anthropic cites its own model evaluations. The claims are large, and every one of them is vendor-supplied.
The 2024 scandal was medical devices approved on retrospective vendor evidence inside an FDA pathway. The 2026 tools skip the pathway entirely and enter the hospital as productivity software and wellness features, not as medical devices.
No clearance.
No validation requirement.
No re-evaluation clock.
ECRI, the independent patient safety organization, named misuse of general-purpose AI chatbots the number one health technology hazard of 2026, stating plainly that these tools are not regulated as medical devices and have not been validated for healthcare purposes.
The first independent check arrived 6 weeks after launch, and it tells a familiar story. A Mount Sinai team published a structured stress test of ChatGPT Health in Nature Medicine: 60 clinician-authored cases across 21 clinical domains, tested under 16 conditions, 960 total responses. Among the cases physicians agreed were emergencies, the tool undertriaged 52%. Diabetic ketoacidosis, told to wait a day or two. Impending respiratory failure, told to wait a day or two. The lead author put it plainly: this is something that can kill someone in a couple of hours. And when a family member in the scenario downplayed the symptoms, the tool tended to downgrade the urgency. The machine believed the minimizer.
The responses to the study complete the pattern. OpenAI said the test didn't reflect real usage. An academic group argued the exam-style format drove the failures. But do you see what nobody disputed? The number. And note what everyone is arguing about: the test. After deployment. On millions of users. That argument is exactly what it looks like when no standard evaluation is required before deployment. The builder graded its own exam, someone else graded it differently, and the fight is now over the grading rubric while the product stays live. That’s not a safety program; it’s a marketing department with a denial desk.
The liability structure completes the pattern. Read the vendors' own usage policies. They require a qualified professional to review outputs before any clinical use. Ok, let’s translate that clause. When the tool is right, the vendor sells the win. When the tool is wrong, the clinician owns the miss. Accountability flows downhill to the person with the least visibility into how the system was built, what it was trained on, and whom it was tested on.
And to be exact about the standard: none of the 3 labs gets a pass.
Not OpenAI.
Not Anthropic.
Not Microsoft.
Remember, some of these companies fund the research institutes that write the whitepapers. OpenAI established a pilot program of fellowships and focused research grants of up to $100,000 and up to $1 million in API credits for work that builds on related policy ideas outlined in their Industrial Policy document and as of May 2026 started convening discussions at their new OpenAI Workshop in Washington, DC.
Some build genuinely useful technology. Neither fact changes the measure, which is the same for all of them. Show independent validation on the populations you serve, on a schedule, in public, or the claim is marketing. An auditor who plays favorites is a vendor with a different hat.
I know what happens when nobody checks
This isn’t abstract for me. During the action hero generation phase, ChatGPT took a photo of me and replaced my face with Jensen Huang’s. My work, my image, reassigned to an Asian man by a system nobody verified. The Stanford paper describes the "strip-mining" of Black creative genius as theory. Well, that makes me the exhibit.
The paper goes on to warn that AI-generated content erodes trust, appropriates Black creativity, and performs worse on darker skin. All true. All documented. And all of it continues, because documentation without verification is a press release for the problem.
What Congress actually needs
The CBC doesn’t need another literature review. It needs an assurance architecture, and these are the 4 moves to get there:
Independence by law. No AI system making consequential decisions about health, credit, education, or liberty gets certified on the developer's own evidence. Separate the builder from the checker. And classification games don't clear the bar. A model that summarizes charts, drafts discharge instructions, or weighs differentials sits in the clinical decision path whether the label says medical device, productivity suite, or wellness feature. The auditor cannot be the vendor.
Verification on the population, not the convenience sample. No certification without evidence that the validation population matches the deployment population. Demographic performance floors, stratified results, and a public answer to the only question that matters: tested on whom?
Verification on a clock. Learning systems drift. Certification must expire. Re-evaluation on a defined cadence, on current data from the population actually being served, published where patients and parents can read them.
Liability with a name on it. Every question the Stanford paper punts to "policymakers must consider." What it really is; is a question about who pays when the system fails. So, assign it. A rule without an accountable party is a suggestion.
And guess what? None of this requires new science. ISO 42001 exists. Audit methodology exists. Independent assurance firms exist. But what's missing is the requirement.
The Stanford paper gave Congress the awareness. 2.5 years later, that awareness has aged into urgency. The next document on a CBC staffer's desk should not describe the problem again. It should verify whether anything changed.
Checkbox theater ends where independent verification begins. The diagnosis is done. Send in the auditors.
Share this article
Related Articles
The Reskilling Illusion: When AI Transformation Means "You're Fired"
Oct 03, 2025