
The 90% appeal overturn rate on AI-generated denials in Medicare Advantage isn't a model performance problem. It is a contract liability problem — and the Lokken v. UnitedHealth litigation has exposed its architecture in federal court.
The methemoglobinemia patient at the center of the case had blood oxygen levels her diagnosis group's average couldn't capture. nH Predict — UnitedHealth's utilization management AI — weighted population-level recovery timelines heavily and individual clinical indicators minimally. Her family paid $16,768 out-of-pocket to prevent premature discharge. A federal court has since granted discovery into nH Predict's training data, development documents, and validation reports.
That 90% overturn rate carries a context most compliance teams haven't fully absorbed: only 0.2% of Medicare beneficiaries appeal their denials. The appeals that did reach the courts exposed a systematic pattern — a 90% reversal rate across the cases that were brought — and that's what drove the class action. Every denial that wasn't appealed remains unchallenged in the record. The OIG separately found that 13% of prior authorization denials were for requests that should have been granted under existing coverage rules. Every Medicare Advantage plan using AI in utilization management is carrying that same actuarial weight in their denial inventory, whether or not anyone has sued yet.
The question isn't whether your algorithms will face scrutiny. It's whether they will survive it — in court, in CMS audit, and in the press on the day your denial rate is publicly reported at contract level.
The Contract Problem Your EOC Language Creates

Evidence of Coverage documents at most Medicare Advantage organizations promise that coverage decisions are made by "clinical services staff" using "medical judgment." Those are contractual representations — not aspirational language. When an AI model generates the denial and a human reviewer rubber-stamps it within seconds, opposing counsel can argue that the "clinical services staff" language is a misrepresentation. That's how a coverage dispute becomes a bad-faith tort.
The Lokken case made this explicit. When NaviHealth managers narrowed the acceptable variance from nH Predict's projections from 3% to 1%, they converted a decision-support tool into an automated gatekeeper. Clinicians who overrode the algorithm faced disciplinary action. The human-in-the-loop became performative, and the court found that every denial generated under those conditions carried breach-of-contract exposure.
The UM committee override rate spreadsheet most compliance teams maintain — showing aggregate override behavior quarterly — doesn't document what the algorithm weighted in a specific prior authorization and why. It can't close the gap between what the EOC promises and what the algorithm actually did. When opposing counsel asks what "clinical services staff" did before this specific denial was issued, the override rate is not an answer.
What the Discovery Order Actually Changed

The March 9, 2026 discovery order in Lokken (2026 WL 658883) granted plaintiffs access to AI development documents, training data specifications, and model validation reports. This is not a narrow case-specific ruling — it signals that training data provenance records are discoverable in federal class actions involving AI coverage decisions.
Every MAO should now operate on the assumption that their AI documentation will be reviewed by opposing counsel. The problem: most UM AI deployed today has no structured per-decision logs, no version-controlled training data records, and no documented validation results tied to specific coverage decisions. That documentation doesn't just need to exist — it needs to exist in a form that a federal judge can read and that your own lawyers can reconstruct years after the decision was made.
CMS reinforced this pressure from a different direction. In February 2026, CMS began Payment Year 2020 RADV audits using AI-powered anomaly detection to flag unsupported diagnoses and statistical outliers. CMS is using AI to audit your AI. Plans whose algorithms generated unsupported diagnoses at scale are discovering that audit risk on the revenue side is compounding with litigation risk on the coverage side.
Three Deadlines That Have Already Started Running

CMS-0057-F created a prior authorization timeline most plans have nominally addressed operationally. January 1, 2026 required 72-hour expedited and 7-day standard PA turnaround, along with restrictions on reopening previously approved inpatient admissions. March 31, 2026 was the first deadline for public reporting of eight PA metrics at contract level — denial rates, turnaround times, and appeal overturn rates are now visible to regulators, media, and plaintiff attorneys simultaneously.
A publicly reported 90% overturn rate at contract level is not only a reputation problem. It signals the plaintiff bar and state AGs about where to look next. The denial-rate profile you published in March 2026 is now an open invitation.
The January 2027 deadline is the structural one. HL7 FHIR Prior Authorization APIs — the CRD, DTR, and PAS transactions — will require a full electronic per-decision PA trail for every covered transaction. That trail will be discoverable. A plan whose AI generates denials without structured per-decision logging will be building a liability record one FHIR transaction at a time starting January 1, 2027.
State pressure is running in parallel. Texas settled the first healthcare generative AI investigation — against Pieces Technologies in September 2024, requiring accurate accuracy disclosures — and the Texas Responsible AI Governance Act, effective January 2026, gives the Attorney General broad civil investigative demand authority. Pennsylvania has introduced legislation requiring human provider review before any AI-driven denial, mandatory insurer disclosure of AI use, and annual compliance statements. Multi-state MAOs face a patchwork of disclosure and audit requirements, each with different thresholds and each potentially triggering state AG interest if publicly reported PA metrics look anomalous.
The convergence risk is specific: CMS is scaling its own AI-powered audit capability while simultaneously requiring plans to govern their AI. You are being audited by the same technology you're being asked to govern.
Where Existing Tools Leave the Gap

The AI governance platform market has grown quickly. Fiddler AI offers SHAP/LIME explainability, drift detection, and bias monitoring — genuinely useful tooling for aggregate model evaluation. Credo AI and its peers (IBM Watsonx.governance, Holistic AI) package policy compliance packs for EU AI Act, NIST, and ISO frameworks with automated evidence collection. These platforms solve a real problem — monitoring a deployed model against current standards. What they don't do is build the explainability layer into the decision workflow itself.
If your UM AI has the same architectural flaw nH Predict had — weighting population-level throughput more heavily than individual clinical indicators — monitoring that model more carefully doesn't fix it. A SHAP waterfall chart from a governance dashboard shows how features contribute to predictions in aggregate across the validation set. It doesn't produce the per-decision audit record that a discovery order requires: what the algorithm weighted in this specific prior authorization, on this specific date, under the specific model version that was running at the time.
Prior authorization automation vendors address a different problem entirely. Cohere Health claims 47% administrative cost reduction; FinThrive and Availity focus on workflow throughput. They're solving speed, not court-defensibility. A PA workflow that processes in three minutes instead of three days, with the same explainability gap, is a faster liability generator. The $19.7B in annual hospital spending on denial overturn — from AMA and industry data — represents the cost falling on providers, not plans. The governance architecture your plan needs is about the cost that eventually falls on you.
The legacy claims system vendors — Cognizant and TriZetto for Facets and QNXT, the systems most large MAOs actually run — also sell AI add-ons for those platforms. They're not positioned to surface governance gaps in their own systems. Building custom explainability middleware that integrates at the transaction level with your specific claims configuration requires a vendor whose incentive is the governance outcome, not the platform sale.
What Court-Defensible Governance Requires

The documentation infrastructure a discovery order actually demands is more specific than most governance frameworks contemplate. For a coverage decision made 18 months ago, you need: the specific model version that was running at that time; the training data it was built on, with version control tracing back to that deployment; the validation results that were current at deployment; and a structured record of what the algorithm considered and weighted in that specific decision.
Most governance platforms audit the current model against current standards. That's useful for CMS compliance reporting and for bias monitoring going forward. It doesn't satisfy a retroactive litigation defense.
Our work on Medicare Advantage AI Governance & Algorithmic Compliance was built around this retroactive documentation requirement. The architecture covers explainability middleware integrated at the claims-system transaction level, version-controlled model documentation tied to specific deployment periods, and CMS-0057-F compliance infrastructure that maps the January 2027 FHIR API requirements to current UM workflow configurations. Only 12% of U.S. health systems have formal AI governance frameworks in place (Censinet, 2026) — which means plans that build this now establish a demonstrable compliance posture before the RADV audit cycle and the FHIR deadline converge.
The Question Worth Asking Your Vendor Today
Not "does our model produce explainable outputs?" — most models produce some form of feature attribution. The question that matters is whether you can produce a structured, version-controlled, litigation-ready audit record for a specific coverage decision made 18 months ago, covering the specific model version that was running at that time, the training data it was built on, the validation results that were current at deployment, and the specific clinical indicators it weighted in that decision.
If the answer is no, the architecture problem is the same one Lokken exposed.
The plans taking this seriously are finding that the governance infrastructure required to answer that question affirmatively — the per-decision logs, the training data manifests, the version-controlled deployment records — is also the infrastructure required to comply with the January 2027 FHIR mandate and to defend against state AG civil investigative demands. The problems converge.
If your team is navigating the gap between what your governance dashboard produces and what a federal discovery order might actually require, we've documented the architecture requirements in detail at Medicare Advantage AI Governance & Algorithmic Compliance. We'd be interested in what your compliance team is finding — the plans building this rigorously tend to discover the same structural gaps in the same places.