
An algorithm denied medical coverage to thousands of elderly patients. When those decisions were appealed, nine out of ten were reversed — the AI had been wrong almost every time. But only 0.2% of patients ever appealed, because most were too sick, too old, or too overwhelmed to fight back. That gap between a 90% error rate and a 0.2% appeal rate became one of the most profitable failures in healthcare AI history.
The algorithm was called nH Predict, and it belonged to UnitedHealth Group. In February 2025, a federal judge allowed a class action lawsuit to proceed against the company, marking what may be the most significant legal ruling on AI accountability to date. We've spent months analyzing this case and its implications — not because it's an isolated disaster, but because it's a blueprint for everything that can go wrong when organizations deploy AI without rigorous governance. Our interactive analysis of the case and its implications traces the full arc from technical failure to legal reckoning.
What happened at UnitedHealth isn't a cautionary tale about AI gone rogue. It's a cautionary tale about AI gone unchecked.
A $1 Billion Algorithm That Couldn't See Its Patients
UnitedHealth's Optum division acquired nH Predict for over $1 billion in 2020. The tool was designed to predict how long Medicare Advantage patients would need post-acute care — time in skilled nursing facilities or rehabilitation centers after a hospitalization. It drew on a database of six million patient records to generate a "target" discharge date.
The problem wasn't that the algorithm made predictions. The problem was what it ignored.
nH Predict worked by cross-referencing historical outcomes — essentially pattern-matching against past patients with similar diagnoses. But it had no mechanism to account for the specific realities of the person in front of it: whether they had a caregiver at home, whether they could afford medication, whether their particular complications required extended care. It could tell you what usually happened. It couldn't tell you what should happen.
A model that predicts average outcomes will, by definition, fail every patient who isn't average.
This is the core distinction between correlation and causation in AI. A correlative model says, "Patients with diagnosis X typically stay 14 days." A causal model asks, "What factors cause this patient to need more or less time, and what happens if we cut their coverage short?" nH Predict never asked the second question.
The Numbers Behind the Denials

The operational impact was staggering. Before nH Predict was widely deployed, UnitedHealth's post-acute care denial rate hovered between 8.7% and 10.9%. By 2022, it had surged to 22.7% — more than doubling. Skilled nursing facility denials specifically rose to nine times their baseline level.
Meanwhile, the algorithm enabled UnitedHealth to process coverage reviews six to ten minutes faster per case. Speed went up. Accuracy went down. And the company's revenue climbed from $240 billion in 2019 to a projected $340 billion in 2025.
The economics were perverse but effective. Even with a 90% error rate on appeals, the system was profitable because almost nobody appealed. The elderly and disabled patients most affected by these denials were precisely the people least equipped to navigate a complex, multi-step appeals process. The algorithm didn't need to be right. It just needed to be hard to challenge.
When only 0.2% of patients appeal and 90% of those appeals succeed, the system isn't making good decisions — it's making profitable ones.
When "Decision Support" Becomes a Mandate

Perhaps the most troubling finding from the Senate investigation and subsequent litigation was how nH Predict was actually used in practice. UnitedHealth maintained publicly that the algorithm was a "guide" — a decision-support tool that informed human judgment. The evidence told a different story.
Internal documents and whistleblower testimony revealed that NaviHealth managers instructed clinical staff to keep patients' actual lengths of stay within a narrow variance of the algorithm's prediction. That target started at 3% and was later tightened to 1%. Care coordinators were directed to time their progress reviews to coincide precisely with the algorithm's predicted discharge date — not with the patient's actual medical trajectory.
Clinicians who deviated from nH Predict's projections to accommodate a patient's real medical needs faced disciplinary action or termination. Experienced doctors and nurses — people whose professional and ethical obligation is to the patient — were reduced to rubber-stamping outputs from a model they knew was flawed.
This is what we call algorithmic coercion: the point where a tool designed to assist human judgment instead replaces it entirely, while the organization maintains the fiction of human oversight.
Consider Carol Clemens. After a life-threatening episode of methemoglobinemia — a blood disorder that can be fatal — she required intensive skilled nursing care. Despite clear clinical evidence of ongoing need, nH Predict's projections were used to terminate her coverage. Her family paid over $16,768 out of pocket to prevent her premature discharge. The litigation alleges UnitedHealth counted on patients like Clemens being too impaired or under-resourced to fight back.
The Legal Ruling That Changes Everything
On February 13, 2025, U.S. District Judge John R. Tunheim allowed the class action — Estate of Gene B. Lokken v. UnitedHealth Group — to proceed. The ruling matters far beyond this single case.
The court found that UnitedHealth's own policy documents promised coverage decisions would be made by "clinical services staff" and "physicians." By substituting those humans with an algorithm that effectively dictated outcomes, the company may have breached its contract with policyholders. The judge allowed claims for breach of contract and breach of the implied covenant of good faith to move forward.
Equally significant: the court waived the usual requirement that patients exhaust their administrative appeals before suing. The reasoning was blunt — given the "irreparable injury" patients faced and the "futility" of appealing to a system with a 90% error rate, forcing patients through that process would amount to requiring them to participate in a broken system.
When a court waives exhaustion of remedies because the system itself is fundamentally flawed, that's not just a legal ruling — it's a verdict on the technology.
For any organization deploying AI in high-stakes decisions, this ruling establishes a clear principle: if your AI system promises human oversight but delivers algorithmic automation, you face contract liability. The legal shield many companies assumed they had is thinner than they thought.
Why "Wrapper AI" Is a Liability in Regulated Industries
The nH Predict failure exemplifies what happens when organizations treat AI as a thin layer of automation rather than a deeply governed system. In the industry, these are often called "wrapper" solutions — applications that take an existing AI engine, wrap it in a custom interface, and deploy it without building proprietary logic, causal reasoning, or audit infrastructure underneath.
In low-stakes contexts, wrappers can be useful. In regulated industries — healthcare, insurance, financial services — they're a ticking liability. They inherit whatever biases exist in their foundational models. They can't explain their reasoning in ways that satisfy regulators or courts. And they create a dangerous illusion of sophistication over what is, fundamentally, pattern-matching.
The alternative — what we describe as Deep AI — requires a fundamentally different approach. Instead of asking "what does the data predict will happen?", deep AI systems ask "what causes this outcome, and what happens if we intervene?" This shift from prediction to causation is what separates a tool that can be audited, explained, and trusted from one that can't.
Our detailed technical research maps this distinction across five dimensions: defensibility, vendor independence, compliance readiness, bias mitigation, and accountability. In every category, the gap between wrapper approaches and deeply governed systems is widening — and regulators are noticing.
The Regulatory Walls Are Going Up
The regulatory landscape has shifted dramatically. In January 2025, the FDA issued draft guidance establishing a seven-step credibility assessment framework for AI models used in medical and regulatory decision-making. The framework requires organizations to clearly define what question their AI is answering, specify its exact role in the workflow, assess the consequences if it's wrong, and validate it rigorously — including stress-testing for edge cases.
nH Predict would have failed every step.
Simultaneously, the EU AI Act began phased enforcement in 2025, classifying healthcare AI systems as "high-risk" and requiring mandatory conformity assessments, transparency disclosures, and human oversight mechanisms. Non-compliance penalties run up to 7% of global turnover. The World Health Organization has separately flagged "automation bias" — the tendency for clinicians to defer to an algorithm even when it contradicts their own clinical judgment — as a specific threat to patient safety.
And the disclosure pressure is mounting from investors, too. By 2025, 72% of S&P 500 companies had disclosed material AI risks in their SEC filings, with reputational damage as the top-cited concern.
This isn't a future regulatory environment. It's the current one.
What Should Explainability Actually Look Like?

Regulatory compliance isn't just about checking boxes. It requires AI systems that can show their reasoning — what the field calls Explainable AI, or XAI. Think of it as the difference between a doctor who says "take this medication" and one who says "here's your diagnosis, here's why I'm recommending this treatment, and here are the alternatives we considered."
Two technical approaches matter here. SHAP (SHapley Additive exPlanations) provides a global view of what factors drive an AI's decisions across many cases — useful for auditors who want to know whether an insurance model is systematically weighting zip code (often a proxy for race) too heavily. LIME (Local Interpretable Model-Agnostic Explanations) explains individual decisions — why this patient was denied, based on which specific factors.
For Carol Clemens, a LIME-equipped system would have flagged that the model was ignoring her dangerously low blood oxygen levels in favor of an average recovery timeline for her diagnosis category. An auditor reviewing SHAP outputs might have caught the systematic pattern months earlier.
Equally critical is confidence scoring — the system's own assessment of how reliable its prediction is. When a patient presents with a rare condition underrepresented in the training data, the AI should explicitly flag its uncertainty and route the case to a human reviewer. nH Predict had no such mechanism. It delivered its predictions with the same authority whether it was confident or guessing.
But Doesn't Human Oversight Solve This?
This is the objection we hear most often: "We have humans reviewing AI outputs, so we're covered." The UnitedHealth case demolishes that argument. They had humans in the loop too — clinicians, case managers, physicians. But when those humans were punished for disagreeing with the algorithm, the "loop" became a rubber stamp.
Human oversight only works when the human has the authority, the information, and the institutional support to override the machine. If your organization penalizes clinicians for deviating from algorithmic recommendations, you don't have human oversight. You have human theater.
What About Industries Outside Healthcare?
The principles exposed by this case apply wherever AI makes high-stakes decisions about people. Insurance underwriting. Loan approvals. Hiring algorithms. Parole recommendations. Any system where an algorithm's output materially affects a person's life, livelihood, or liberty faces the same governance questions: Can you explain the decision? Can a human meaningfully override it? Have you tested for bias? What happens when the model is wrong?
The FDA's credibility framework, the EU AI Act's risk classifications, and the NIST AI Risk Management Framework all converge on the same core requirements: define the AI's role clearly, assess the consequences of failure honestly, validate rigorously, and document everything.
AI governance isn't a compliance exercise. It's the difference between a tool that enhances human judgment and one that quietly replaces it.
What Needs to Change — Starting at the Top
Algorithmic governance can no longer live in the IT department. It's a board-level responsibility. Organizations deploying AI in consequential decisions need cross-functional AI governance committees — not as advisory bodies, but with real authority to approve use cases, maintain a registry of every model in production, and enforce rollback mechanisms when performance degrades.
Three things should happen immediately in any organization using AI for decisions that affect people:
Audit your override rates. If humans almost never disagree with the AI, you don't have oversight — you have automation bias. Investigate why.
Test your appeals process. If the people affected by AI decisions can't realistically challenge them, your error-correction mechanism is broken, no matter what your error rate actually is.
Require confidence scoring. Any AI system making high-stakes recommendations should be able to say "I'm not sure about this one" — and that flag should trigger automatic human review.
The UnitedHealth case will be studied for years. Not because an algorithm made mistakes — all algorithms do — but because an organization built a system where mistakes were profitable and corrections were nearly impossible. That's not a technology failure. It's a governance failure.
The question every leader should be asking right now isn't "Is our AI accurate?" It's "What happens when it's wrong — and who has the power to fix it?"