Klarna replaced 700 support agents with AI. Costs dropped 40%. Then Q1 2025 closed with a $99 million net loss.
Within months, they rehired humans.
The part most people miss: the AI wasn't broken. It handled password resets and routine chats well, dropping cost per transaction from $0.32 to $0.19.
What nobody validated was the other 20% — the multi-currency refund on a cancelled flight, the disputed merchant charge. The interactions that decide whether a customer stays or a regulator calls. That's where it collapsed.
This is the gap behind the stat that 70-85% of enterprise AI projects never reach production.
Generic guardrails catch toxicity and leaked PII. They do not catch an AI that miscalculates an insurance reserve, cites a repealed statute, or approves a loan that breaks fair-lending rules. On legal due-diligence tasks, AI error rates run 69-88%, and courts logged 729 AI-hallucination incidents in legal filings by the end of 2025 — none a toxicity filter would catch.
A green governance dashboard tells you the policy boxes are checked. It does not tell you the AI is giving correct answers for your domain. Only one of those shows up in a quarterly loss.
With the EU AI Act's high-risk rules live in August 2026 and penalties reaching EUR 35M per violation, "it passed QA" is no longer a defense.
If you've put an AI system into production: what broke first — the routine path everyone tested, or the edge case nobody did?
#AIGovernance #EnterpriseAI #AIValidation
Published on Facebook · June 21, 2026
On social media
See this post on its original platform
In our archive