
The Platform Got the Deduction Wrong. The IRS Penalty Wouldn't Have Cared.
I was in a client's tax technology review when I first saw it live — not in a test environment, in production. The OBBBA car loan interest deduction, classified as above-the-line by the AI. Section 63(b)(7) is explicit: it's below-the-line. The consensus error wasn't one platform's bad training data. It was three platforms, same misclassification, same confident output. The IRS accuracy-related penalty for getting that wrong is 20% of underpayment. The fraud penalty is 75%.
The person who filed the return didn't know. The platform didn't flag it. The IRS would not have cared about the model's training data.
That observation reshaped what we built with Tax Compliance AI Verification. Not because AI tax tools are bad — Thomson Reuters CoCounsel and CCH Axcess Expert AI make teams faster in ways that are genuinely valuable. But faster at filing the wrong answer is not progress.
The Audit Rate Changed the Risk Math

My frame on AI tax tools changed the day the IRS published its intention to raise the large corporate audit rate from 8.8% to 22.6%. That number sits in your head differently when you've seen the accuracy penalty worksheet — 20% of the underpayment, automatically. The fraud penalty is 75%.
This is the IRS investing in its capacity to find what the AI missed. The tax compliance AI market is investing in preparing returns faster. Those are not the same direction.
The vendors building preparation infrastructure are doing real work — I watched those tools ship over the past year and the pattern was consistent. ONESOURCE's claimed 65% reduction in routine reporting time, EY.ai for tax running on IBM's watsonx and targeting 80% automation of foreign tax compliance, KPMG's Tax AI Accelerator on Azure OpenAI — each of them is a legitimate response to a real problem. Business tax compliance costs $126 billion annually. Federal compliance forms consume 11.6 billion labor hours. The staffing crisis — accounting enrollment down, senior practitioners retiring — makes that math worse each year.
None of these tools were designed to verify their own output. That's not a criticism; it's an architectural observation. Thomson Reuters verifies Thomson Reuters. Wolters Kluwer verifies Wolters Kluwer. The cross-platform neutrality question — what checks the output when a client's workflow spans four ERPs and three AI tools — isn't answered by any of them.
What the Heppner Ruling Did to the Privilege Calculation

A partner at a client firm called me the day Judge Rakoff's Heppner ruling came down, February 10, 2026. The question was direct: their team had been using a public Claude session to verify uncertain tax positions before finalising. Was that now subpoenaable?
The answer from their outside counsel was yes. Using a public AI tool for tax communications waives attorney-client privilege. The session logs were discoverable.
This was not a theoretical future risk. It was a live engagement, a working team, a current-year filing cycle. The way the compliance AI market had been operating — prepare with enterprise software, verify informally with public models — turned out to be a privilege exposure that had been accumulating silently. Approximately 800 AI citation error cases have been logged across 25 countries through late 2025. Fifty percent of UK accountants are now aware of businesses suffering direct financial losses from AI-generated errors. Those cases were caught after they mattered, not before.
IRM 10.24.1, the IRS's own AI governance policy formalised in February 2026, requires enhanced human oversight for AI outputs that serve as the basis for decisions with legal or material effect. The IRS built governance infrastructure for how it uses AI. The taxpayers it audits largely haven't.
What Verification Actually Requires

Blue J is the closest thing to a verification product in the market — a disagree rate below 1 in 700 against primary source rulings, $122 million in Series D funding, 220+ jurisdictions via the IBFD partnership. It's a probabilistic research engine. When an ASC 740 FIN 48 uncertain-tax-position is on the workpaper and the penalty for getting it wrong starts at 20%, probabilistic is the wrong register.
What we built runs deterministic checks against primary source statutory text. The OBBBA deduction is below-the-line because Section 63(b)(7) says it is — and if any preparation tool in the client's composite workflow says otherwise, the verification layer catches it before signature. The architecture runs entirely within the enterprise environment. No public endpoints, no data leaving the firewall. Post-Heppner, that's a compliance requirement, not a preference.
The clients who need this most are the ones running AI fastest across the most fragmented data environments. The 78% of companies running four to seven ERP systems aren't on a single-vendor stack — they're assembling compliance outputs from multiple tools and signing off on the composite. The verification gap is widest at the seam between platforms.
The OBBBA error wasn't in any one platform's test suite. It was in all of them simultaneously. Consensus error is the failure mode that no single-vendor quality check can catch.
What I took from the production misclassification I saw, and from the Heppner call, and from watching IRS audit rates climb — is that the question isn't whether AI should be in the tax workflow. It's whether there's something deterministic sitting between the AI's output and the signature. If your platforms include any of the ones running the OBBBA misclassification, that question is already live. The Tax Compliance AI Verification page explains how we answer it.