
When nine out of ten AI-driven decisions are wrong — and the company using it knows this — what do you call that? UnitedHealth Group's nH Predict algorithm denied post-acute care coverage to elderly Medicare patients with a 90% error rate on appeal. But only 0.2% of patients ever appealed. The algorithm wasn't broken. It was profitable. This is the most important AI governance crisis of 2025, and it should terrify every executive deploying AI in high-stakes decisions — not just in healthcare, but in insurance, lending, hiring, and beyond.
I've spent the last several months studying the UnitedHealth case in detail — the court filings, the Senate investigation, the whistleblower testimonies. What I found wasn't a story about a bad algorithm. It was a story about what happens when organizations treat AI as a cost-optimization machine and strip away every safeguard designed to protect people. We published our full analysis here, but I want to share what hit me hardest — and what I think it means for anyone building or buying AI today.
How Did a $1 Billion Algorithm Miss Dying Patients?
UnitedHealth acquired nH Predict through its Optum division for over $1 billion. The tool was built on 6 million patient records and designed to predict how long Medicare patients would need post-acute care — time in skilled nursing facilities, rehabilitation centers, recovery programs.
On paper, that sounds reasonable. In practice, the algorithm functioned like a spreadsheet with amnesia. It cross-referenced historical outcomes to generate a "target" discharge date, but it couldn't account for whether a patient had a caregiver at home, whether they could afford medications, or whether they had complications that didn't fit neatly into a statistical average.
Think of it like a GPS that calculates your route based on average traffic patterns but can't see the bridge that's washed out ahead. It gives you a confident answer. The answer is wrong. And if you follow it, you drive off a cliff.
The algorithm didn't fail because it was inaccurate. It failed because accuracy was never the point — throughput was.
What Numbers Should Keep Healthcare Executives Awake?
The operational data from the Senate investigation tells a story of systematic escalation. Post-acute care denial rates jumped from roughly 9-11% to 22.7% — more than doubling. Skilled nursing facility denials surged to nine times baseline levels. UnitedHealth's revenue, meanwhile, climbed from $240 billion in 2019 to a projected $340 billion in 2025.
I remember the moment my team and I mapped these numbers side by side. Revenue going up. Denial rates going up. Error rates on appeal at 90%. And the appeal rate itself at 0.2%. We sat with that for a while.
The economic logic is brutal in its simplicity: if your algorithm wrongly denies a thousand claims, and only two people fight back, you've saved money on 998 of them. The 90% error rate isn't a bug — it's irrelevant to the business model, because the system banks on "administrative friction." Most elderly or disabled patients simply don't have the cognitive, physical, or financial resources to navigate a multi-stage appeals process while simultaneously needing medical care.
When Does the "Human in the Loop" Become a Rubber Stamp?
Every responsible AI framework talks about keeping humans in the loop — making sure a person reviews the machine's output before it affects someone's life. UnitedHealth had humans in the loop. It just punished them for doing their jobs.
Whistleblower testimonies revealed that NaviHealth managers set rigid targets: case managers had to keep patient stays within 1% variance of the algorithm's prediction. Not 10%. Not 5%. One percent. Clinicians who deviated — who looked at a patient and said "this person isn't ready to go home" — faced disciplinary action or termination.
I've talked to people who build AI governance frameworks, and this case broke something in the conversation. We've been debating technical safeguards — confidence scores, explainability tools, bias audits — and those matter enormously. But none of it matters if the organizational culture treats the algorithm as the authority and the human as the compliance risk.
A human-in-the-loop who gets fired for overriding the machine isn't a safeguard. They're a fig leaf.
Consider Carol Clemens. After a life-threatening episode of methemoglobinemia — a blood disorder that left her critically ill — she needed intensive skilled nursing care. Clinical evidence supported this. The algorithm didn't care. Her coverage was terminated based on nH Predict's projections, and her family paid over $16,768 out of pocket to prevent her premature discharge. The litigation alleges UnitedHealth counted on patients like her being too sick to fight back.
February 2025: The Court Said Enough
On February 13, 2025, U.S. District Judge John Tunheim ruled that the class action — Estate of Gene B. Lokken v. UnitedHealth Group — could proceed. This ruling matters far beyond one lawsuit.
The court allowed breach of contract claims to move forward on a specific and devastating basis: UnitedHealth's own policy documents promised that coverage decisions would be made by "clinical services staff" and "physicians." By letting an algorithm effectively dictate outcomes — and disciplining humans who disagreed — the company may have violated its own contractual commitments to policyholders.
Even more significant, the court waived the exhaustion requirement. Normally, Medicare beneficiaries must grind through every level of internal appeal before they can sue. Judge Tunheim found that requiring patients to exhaust a process with a 90% error rate would constitute "irreparable injury" — legal language for: we're not going to make dying people participate in a system designed to outlast them.
This is the moment AI governance stopped being a compliance exercise and became a litigation reality. If your AI system is demonstrably broken and you keep deploying it, courts won't hide behind procedural technicalities.
Why "Wrapper AI" Is a Liability in Disguise

The nH Predict crisis crystallizes something I've been arguing for years: there is a fundamental difference between AI that automates a surface-level process and AI that actually understands what it's doing.
Most enterprise AI today falls into what I call the "wrapper" category — a thin layer of automation over someone else's model. You take an existing AI engine, put a custom interface on it, and ship it. It's fast to build, easy to sell, and nearly impossible to audit. When something goes wrong, nobody can explain why the system made the decision it made, because nobody built the reasoning — they just built the interface.
nH Predict was a more sophisticated version of this problem. It had its own model, but that model was purely correlation-driven. It could tell you that patients with diagnosis X historically stayed 14 days. It couldn't tell you why a specific patient might need 21 days. It had no concept of causation — no ability to reason about what happens when you remove care from someone who isn't ready.
Correlation tells you what usually happens. Causation tells you what will happen to this person if you make this decision. In healthcare, that gap kills people.
The alternative — what we've been building toward at Veriprajna — is what I call Deep AI: systems that incorporate causal reasoning, that can explain their logic in terms a clinician or auditor can evaluate, and that flag their own uncertainty instead of projecting false confidence. For the full technical methodology behind this distinction, we've published a detailed breakdown.
The Regulatory Walls Are Going Up Fast
If the court ruling didn't get your attention, the regulatory landscape should.
In January 2025, the FDA issued a 7-step credibility framework for AI models used in medical and regulatory decisions. It requires organizations to clearly define what question the AI is answering, specify its exact role in the workflow, assess what happens if the AI is wrong, and validate the model against those consequences — not against cost savings. nH Predict would have failed every single step.
The EU AI Act, now in phased enforcement, classifies healthcare AI as "High-Risk" — requiring conformity assessments, transparency disclosures, and mandatory human oversight. Penalties run up to 7% of global turnover. For a company UnitedHealth's size, that's potentially tens of billions of dollars.
The WHO has specifically warned about "Automation Bias" — the documented tendency for trained professionals to defer to an algorithm even when their own expertise tells them it's wrong. When you combine automation bias with a corporate culture that punishes override decisions, you get exactly what happened at NaviHealth.
And here's a number that should reframe every board conversation about AI risk: 72% of S&P 500 companies now disclose material AI risks in their SEC filings. This is no longer an engineering concern. It's a fiduciary one.
What About Companies That Aren't in Healthcare?
I hear this objection constantly: "We're not denying medical care. Our AI just handles [invoices / hiring / customer service / risk scoring]. This doesn't apply to us."
It applies to you. The legal principle established in the Lokken ruling — that substituting AI for promised human judgment can constitute breach of contract — has no industry boundary. If your terms of service, employment contracts, or lending agreements promise human review, and your AI is making the actual decision while a human rubber-stamps it, you have the same structural vulnerability.
The FDA framework, the EU AI Act, the NIST AI Risk Management Framework — these aren't healthcare-specific ideas wrapped in healthcare-specific regulation. They're governance principles. Define what the AI is deciding. Assess what happens when it's wrong. Make sure someone with authority can override it without getting fired. Document everything.
If you can't explain why your AI made a specific decision to a regulator, a judge, or a customer, you don't have an AI governance problem. You have a ticking liability.
What I Think This Means
The UnitedHealth crisis isn't an anomaly. It's a preview. Every industry deploying AI in high-stakes decisions — insurance, lending, hiring, criminal justice, child welfare — faces the same structural temptation: an algorithm that's wrong but profitable, paired with a friction-heavy appeals process that suppresses accountability.
Three things need to change:
AI must explain itself — not to other engineers, but to the people affected by its decisions and the regulators overseeing them. Tools like SHAP and LIME (tools that break down which specific factors drove each AI decision) exist to make this possible. They can show which factors drove a specific decision and whether those factors include proxies for race, age, or income. If your AI can't do this, it shouldn't be making consequential decisions.
The second shift is harder: humans must be empowered, not performative. A human-in-the-loop framework means nothing if override decisions trigger career consequences. Governance committees need clinical, legal, and patient-safety representation — with actual authority to pause or kill a model.
Finally, boards must own this. Not IT. Not a buried compliance team. The boardroom. That means maintaining a registry of every AI model in production, tracking performance against fairness and accuracy metrics, and having a kill switch for any model that drifts from its validated parameters.
The question isn't whether your AI will make a mistake. It's whether your organization is designed to catch it — or designed to profit from it.
I don't think we're going back to a world without AI in healthcare or insurance or any other high-stakes domain. Nor should we. But the nH Predict disaster proved that speed and scale without governance isn't innovation — it's negligence with a technology budget.
The companies that will earn trust in the next decade aren't the ones deploying AI fastest. They're the ones who can look a patient, a customer, or a judge in the eye and explain exactly why the machine made the decision it made — and what happens when the machine is wrong.
If you're grappling with how to build that kind of accountability into your own AI systems, I'd genuinely like to hear what's working and what isn't. This is a problem none of us have fully solved yet.