
There's a finding in the Lancet Digital Health's April 2024 study that I keep coming back to. Not the headline number — the 7.1% of AI-drafted patient portal messages that posed severe harm risk — but the companion to it: 90% of the physicians in the study reported trusting the AI's performance. That trust level and the 33.4% error catch rate exist in the same dataset. The physicians who were most confident in the AI were, by the study's structure, operating on a safety net that caught one-third of its own failures.
I've been doing clinical AI safety assessments for a while, and I think about "physician oversight" differently now than I did before reading that study. "Physician oversight" as a regulatory strategy — which is what California's AB 3030 "read and reviewed" exemption codifies — assumes that the oversight mechanism works. The Lancet data establishes what working looks like in practice.
The CMIO Question I Couldn't Answer Cleanly

A conversation with a CMIO at a mid-size health system, roughly six months into building our Healthcare AI Safety for Health Systems practice, crystallized the core problem. She was asking about AB 3030 compliance — specifically whether her existing review workflows made her legally safer or more exposed. I gave her the standard answer: the "read and reviewed" exemption exists, her physicians were technically reviewing AI-generated messages before delivery, her compliance team had documented the workflow.
Then I walked her through the Lancet numbers. The 35–45% of erroneous drafts submitted entirely unedited. The mechanism behind the error rate: when an AI draft is well-written and empathetic, physicians enter a cognitive state where the quality of the prose substitutes for independent clinical verification. Her review workflow existed on paper and in practice. Whether it constituted actual safety oversight was a different question.
Her read of those two numbers — 90% trust, 33.4% catch rate — was that the exemption made her technically compliant and practically exposed at the same time. That's exactly right, and it's not a problem you can document away.
The legal exposure doesn't come from being out of compliance with AB 3030. It comes from being in compliance with a standard the underlying clinical evidence has hollowed out.
The AB 3030 "read and reviewed" exemption is satisfied by a workflow. What the Lancet evidence says is that the workflow catches one-third of its own failures. Compliance and safety are not the same thing.
The technical fix — uncertainty highlighting, active confirmation workflows that interrupt the rubber stamp, citation linking that traces each note claim back to the audio segment it came from (Abridge's Linked Evidence feature does this; it's one of the few honest attempts at making "reviewed" mean something) — is an engineering problem, not a compliance checkbox. You can't document your way out of automation bias.
What I Found When I Started Looking at Vendor Accuracy Claims

My confidence in vendor accuracy claims started dissolving when I worked through the Pieces Technologies Assurance of Voluntary Compliance from the Texas AG settlement in September 2024. Pieces had claimed a less-than-0.001% "critical hallucination rate" for clinical documentation software deployed at Houston Methodist, Children's Health, Texas Health Resources, and Parkland. The AVC imposed a five-year transparency mandate requiring disclosure of the things the claim hadn't included: how "critical" was defined, which use cases were sampled, what denominator was used.
I spent time with that document. What struck me wasn't that Pieces had gamed the metric — it was that there was no agreed-upon way to verify whether they had. The information needed to evaluate the claim was precisely the information the claim didn't include. Working through the AVC's disclosure requirements made it clear this wasn't a Pieces-specific problem. It was the structure of a market where ambient AI vendors define their own accuracy metrics and health systems don't have the infrastructure to independently test them.
The Q1 2025 discharge assistant incident made it concrete: a deployed AI recommended a medication for a patient explicitly listed as allergic to that drug class. The actual clinically actionable misstatement rate was 0.98% — twelve times higher than the vendor's claimed 0.08%. A nurse caught it. Abridge ($316M Series E, April 2026, $5.3 billion valuation, Best in KLAS for Ambient AI two years running) and Ambience Healthcare ($243M Series C, Cleveland Clinic rollout) are doing serious work on clinical traceability. But across the market broadly, the gap between what vendors claim and what's independently verifiable is the default.
I now go into every clinical AI engagement with one question: can you show me the denominator? If the answer is no, the accuracy claim is unverifiable — and any governance framework built on top of it is built on sand.
The Bias Layer I Didn't Expect to Find This Deep

When I started running demographic performance analyses on deployed clinical AI, I expected performance gaps. I didn't expect the mechanisms to run as deep as they do — or to be this embedded in clinical infrastructure that predates AI software entirely.
The Epic Sepsis Model was the first deployment I pulled external validation data on, and the gap was wider than I expected. External validation at Michigan Medicine found AUC of 0.63 against a developer-reported 0.76–0.83. Sensitivity was 33%; positive predictive value was 12% — the model was correct on 1 in 8 patients it flagged. But the bias dimension isn't just the overall accuracy: Black and Hispanic patients carry nearly twice the sepsis incidence of white patients, and the model's calibration for those demographics was not addressed in the original deployment. A February 2026 multicenter prospective validation of the overhauled ESM v2 was published in JAMA Network Open — the performance and demographic calibration questions remain live.
Pulse oximetry is the part that took me longest to absorb because the bias isn't in the AI software — it's in the medical hardware the AI was trained on. Pulse ox devices are medical hardware, not AI software. But clinical AI trained on their output inherits their measurement errors as if the readings were ground truth. NEJM data shows Black patients are approximately three times more likely to experience occult hypoxemia undetected by pulse ox, with SpO₂ overestimated by 0.6 to 1.5 percentage points in darker-skinned patients. FDA draft guidance from January 2025 recommends 150 diverse participants in validation studies — up from 10 — using the Monk Skin Tone scale with 25% minimum per group. That guidance doesn't retrofit the devices in the field or the AI models already trained on their readings.
The California Maternal Data Center data is the one I find hardest to engage with abstractly. AI early warning systems missed 40% of severe morbidity cases in Black patients. Black women face 42.8–50.3 maternal deaths per 100,000 live births against 13.0–14.5 for white women. McKinsey puts the cost of that gap at $24.4 billion in US GDP impact and $385 million in annual savings. The early warning systems weren't engineered to be racially disparate — they emerged that way from historical clinical data that carried its biases in. The AI layer inherits those biases, and in some cases amplifies them, because nobody ran the subgroup performance audit at deployment. Running equalized odds analysis and Population Stability Index tracking at each deployment site costs $50–200 per expert-adjudicated clinical case. That's a resource gap, not a conceptual one.
The Governance Question That Actually Has Sharp Edges

What I've found talking with CMIOs about AI governance is that the structural question — do you have a governance committee? — is rarely where the gap is. Censinet's 2026 data says 84% of health systems have governance committees, and only 59% have formal documented approval processes before AI implementation. CIOs sit on 63% of those committees; CMIOs on only 45%. The governance infrastructure exists structurally, but it's built around approvals, not verification.
The regulatory question I field most from CMIOs mapping their compliance timelines is when their current governance model stops being adequate. My read: it stopped being adequate when Colorado's AI Act (SB 24-205) — extended from February to June 2026, with no further extension expected — made annual reviews and documented risk assessments a legal requirement for systems making consequential healthcare decisions. The EU AI Act's Annex III high-risk classification takes effect August 2, 2026, with penalties up to €15M or 3% of global turnover. Texas's Responsible AI Governance Act is already live, with per-violation penalties reaching $200,000. HITRUST r2 certification — v11.7 required for new assessments by March 31, 2026 — costs $150,000–$300,000 and takes 9–14 months. What these regulations require in common is documentation that vendor accuracy claims were independently verified, that subgroup performance was tested before deployment, that the governance trail predates any adverse event. An approval committee generates none of that.
Malpractice claims involving AI tools increased 14% from 2022 to 2024, concentrated in radiology, cardiology, and oncology. The standard of care is shifting, with some malpractice insurers adding AI-specific exclusions or conditioning coverage on clinician AI training.
The question I ask every CMIO now is not "do you have a governance committee?" It's "can you produce, today, the evidence that your sepsis model's AUC matches the vendor's claim on your patient population?" Not a model card. Not a SOC 2 on file. Tested, documented, timestamped evidence that predates any adverse event. Across every tool in the portfolio, across every patient demographic you serve.
Most can't produce that. The Healthcare AI Safety for Health Systems work is built to close that specific gap — vendor-neutral, independent, and designed to generate the documentation regulators and plaintiffs' attorneys are starting to ask for.
The Lancet finding I keep returning to isn't just about physician review. It's a description of what happens when any safety mechanism operates at the edge of human cognitive bandwidth — when trust in the tool outruns what the tool has earned. Most health systems are somewhere in that gap right now. The ones building independent verification infrastructure aren't doing it out of regulatory fear — the enforcement actions haven't landed for most of them yet. They're doing it because some CMIO read the same data I did, sat with those two numbers, and decided that a safety net catching one-third of its own failures isn't a foundation you build a program on.