
When I read the March 9, 2026 discovery order in Lokken v. UnitedHealth (2026 WL 658883), the part that stopped me wasn't the scope of what plaintiffs could access — training data, development documents, validation reports. It was the retroactivity. The court was asking UnitedHealth to produce documentation of decisions made years before anyone knew those decisions would be discoverable. That's a different problem from building a compliant system going forward. It's an architecture problem for decisions you've already made.
Most MAOs hadn't absorbed this yet. I started asking a specific question in conversations: if a federal court today issued discovery over your UM AI, what would your litigation team actually point to? The honest answer, more often than not, was: the governance dashboard, the aggregate override reports, maybe the model card from the vendor. None of those answer what a discovery order actually demands — the specific model version that was running on the day of a specific prior authorization, the training data it was built on at that deployment, the validation results that were current then, and a structured record of what it weighted in that decision.
Governance dashboards are built for prospective compliance. The Lokken order revealed that the real liability is retroactive.
The First Compliance Conversation That Changed How I Framed This

My first detailed conversation with a health plan compliance officer went a direction I hadn't expected. I'd gone in thinking the gap was technical — explainability tooling, SHAP output, maybe drift monitoring. Before we got there, she pulled up their Evidence of Coverage document and read me the coverage-decision language: "clinical services staff," "medical judgment," "individualized review."
I asked her to walk me through the actual UM workflow. What I was looking at, within the first few minutes, was an AI-generated denial recommendation that a nurse reviewer approved in under 90 seconds — with no documented independent clinical assessment. The EOC promised one thing. The workflow delivered another. That's not a model calibration problem. That's the Lokken breach-of-contract argument, sitting in the current workflow, waiting for someone with a plaintiff's bar background to notice it.
The OIG had already quantified the systemic version of this: 13% of prior authorization denials were for requests that should have been granted under existing coverage rules. That number is across the industry, not a single plan. The methemoglobinemia patient in Lokken — discharged based on her diagnosis group's average recovery timeline while her actual blood oxygen levels were ignored — wasn't an outlier. She was the appeal that happened to get filed, in a universe where only 0.2% of Medicare beneficiaries appeal at all.
Why I Stopped Treating This as a Monitoring Problem

My initial instinct, shared with most of the people I'd been talking to in the governance platform space, was that the solution was better monitoring: more explainability, more drift detection, better fairness tooling. Fiddler AI does this well. Credo AI and IBM Watsonx.governance package the policy-compliance layer on top.
What changed my thinking was a specific scenario I kept running in my head: if a court ordered discovery over a specific denial from 18 months ago, what would Fiddler AI's current dashboards tell you? They'd tell you how the current version of the model is performing against current standards. They wouldn't tell you what model version was running 18 months ago, what training data it was built on at that deployment, or how it weighted this specific patient's clinical indicators on this specific date.
The governance platform market is built for prospective compliance audits — CMS reporting, bias monitoring going forward, EU AI Act readiness for the August 2027 deadline. Retroactive litigation defense is a different architecture. The documentation it requires — version-controlled training data manifests, per-decision structured logs, deployment records tied to specific date ranges — doesn't get produced by monitoring. It gets produced by an engineering decision made before the decision you're now being asked to defend.
This is the lens I now use when thinking about the January 2027 FHIR mandate. The HL7 Prior Authorization APIs — CRD, DTR, PAS — will create an immutable electronic per-decision audit trail for every covered PA transaction. Starting January 1, 2027, every MAO will be building a discoverable record of each decision, one FHIR transaction at a time. My view: the plans that build structured per-decision documentation before that deadline will be managing compliance. The ones that don't will be generating discovery, transaction by transaction, in real time.
The Conversation That Shaped the Architecture

The architectural question I kept hitting was about the distance between what a governance dashboard produces and what an actual litigation defense requires. The clearest version of that conversation I had was with a health plan's general counsel. The context was a review of their current UM AI governance posture. He asked — in a tone that was more statement than question — what they'd point a federal judge to if a coverage denial from 18 months ago was challenged in class action discovery.
The model card from their vendor didn't cover the specific deployment version from 18 months ago. The aggregate quarterly override reports didn't document what happened in a specific PA. The SHAP outputs their data science team produced were aggregate, not per-decision. There was no training data version control that tied back to that deployment period.
That's the gap our work on Medicare Advantage AI Governance & Algorithmic Compliance was built to close: explainability middleware at the claims-system transaction level, version-controlled model documentation with deployment-period tagging, and the CMS-0057-F compliance infrastructure that maps the FHIR requirement to the current Facets or QNXT configuration the plan actually runs. Only 12% of health systems have formal AI governance frameworks (Censinet, 2026) — which tells me most of these gaps are still open.
What I've Learned About the Regulatory Convergence

I used to present the regulatory timeline as a future planning problem. Texas, Pennsylvania, EU AI Act, CMS-0057-F — each with different deadlines and different requirements. The way clients treated it, I was seeing the same response: a compliance calendar mapped to future milestone dates.
What shifted my framing was watching what happened after March 31, 2026 — the first CMS deadline for public reporting of PA metrics at contract level. Denial rates, turnaround times, appeal overturn rates: now visible to regulators, media, and plaintiff attorneys simultaneously. The $19.7B that the AMA and industry data put on annual hospital spending fighting denials had always been the cost falling on providers. The public metrics made the plan-side exposure visible in the same moment, in the same data.
A publicly reported 90% overturn rate at contract level isn't just a reputational problem — it's a signal to the plaintiff bar about where to look next. The Texas Attorney General's civil investigative demand authority, effective January 2026, and the Pennsylvania disclosure legislation moving through the legislature mean that state-level pressure is arriving on a timeline that doesn't wait for the federal class action calendar. Multi-state MAOs are managing a patchwork right now.
CMS reinforced this from a direction most plans hadn't anticipated. In February 2026, CMS began Payment Year 2020 RADV audits using AI-powered anomaly detection to flag unsupported diagnoses and statistical outliers. The same agency requiring plans to govern their AI is now using AI to audit the revenue claims those plans filed. The plans that built governance infrastructure before the audit cycle are in a materially different posture than the ones discovering documentation gaps in the middle of an audit.
The Diagnostic I Now Use First

When I'm in a first conversation with a health plan about UM AI governance, I've stopped leading with architecture and started with a single question: can you produce a structured, version-controlled audit record for a specific coverage decision made 18 months ago?
Not "does your model produce explainable outputs" — most models do. Not "do you have a governance dashboard" — most plans have something. The specific question: the model version that was running at that time, the training data it was built on, the validation results current at deployment, and what it weighted in that decision.
The answer to that question tells me more about the actual governance posture than any vendor list or compliance calendar does. If the answer is no — and it usually is — the architecture problem is the same one the Lokken court is now examining, at UnitedHealth's scale, in a federal discovery proceeding.
The architecture that actually closes this gap is documented in detail at Medicare Advantage AI Governance & Algorithmic Compliance — the per-decision logging layer, the version-controlled deployment records, the FHIR compliance mapping for the January 2027 deadline.
I'm curious what other people building this infrastructure are finding. The documentation architecture that actually satisfies retroactive discovery tends to converge once plans are serious about it — but getting to that architecture requires having the honest conversation about what a discovery order asks for versus what a governance dashboard provides. If you're working through that gap, I'd be interested in what you're seeing.