
The Sixth Circuit's $30,000 sanctions in March 2026 didn't come from a case that didn't exist. They came from real cases cited for propositions they didn't support. That's contextual hallucination — and it's the failure mode that Mata v. Avianca didn't solve, and that the legal AI platforms now deployed across more than half of Am Law 100 firms weren't designed to solve either.
Mata is largely patched. Harvey and Lexis+ with Protege ground their output in real case databases. Shepard's and KeyCite catch the docket number that resolves to nothing. The fabricated-citation failure mode mobilized the industry, and citation fabrication is now meaningfully reduced.
What didn't get patched is the 33% hallucination rate that Stanford and JELS researchers found for Westlaw Precision in a 2025 peer-reviewed study — not 33% fabricated cases, but 33% of complex queries returning citations to real cases that don't say what the AI claims they say.
What the 33% Actually Measures

The Stanford/JELS study's methodology is the part most reporting glosses over. Westlaw Precision's hallucination rate wasn't measured by checking whether it returned real docket numbers. It was measured by verifying whether the holdings the AI attributed to those real cases were accurate.
That distinction is the whole problem. KeyCite will tell you whether a case exists and whether it's been directly overruled. It won't tell you whether the AI accurately characterized what the case held — or whether the holding has since been narrowed by a line of subsequent decisions that distinguished it without formally reversing it.
A concrete illustration from Delaware corporate law: a litigation associate researching director oversight liability builds an analysis on Stone v. Ritter (2006). The citation is real. KeyCite shows no red flag. The holding summary is accurate for 2006. What the AI missed is that Marchand v. Barnhill (2019) substantially expanded the Caremark oversight duty, and subsequent Chancery opinions have developed a "mission critical" compliance standard that materially changes what Stone means for a 2026 filing. The case hasn't been overruled. The analysis built on Stone alone is wrong.
Lexis+ with Protege — which LexisNexis launched in February 2026 to replace Lexis+ AI after walking back its "100% hallucination-free" marketing language — carries a 17% contextual hallucination rate on the same methodology. Lower than Westlaw's. Wrong one time in six.
The citator says "good law." That means the case wasn't overruled. It says nothing about whether the AI's characterization of the holding is accurate, or whether subsequent decisions have narrowed the proposition the AI is citing it for.
This is the gap. Shepard's and KeyCite were built to track direct negative treatment: reversal, overruling, significant criticism. A line of decisions that progressively narrows a holding's practical scope — without formally overruling it — is a different category of problem requiring a different kind of verification.
What 25,000 Harvey Agents Actually Mean for QA

Harvey's March 2026 funding round valued the company at $11 billion, with $190 million in ARR and more than half of Am Law 100 as clients. The platform now hosts 25,000 custom agents — multi-step autonomous workflows that run research, synthesis, drafting, and issue-flagging in sequence.
Thomson Reuters CoCounsel launched agentic workflows in early 2026, adding autonomous document review and "Deep Research" to what was already a dominant legal research position.
The efficiency case for agentic legal AI is legitimate. The verification challenge that comes with it is underappreciated. An agentic workflow executing a multi-step research task might run a dozen queries, synthesize results, draft a preliminary analysis, and surface a completed memo — and every intermediate citation in the chain needs to hold. If any one step relied on a contextually hallucinated holding, the final output can be internally coherent and wrong.
The New Orleans case from February 2026 made this compounding failure pattern visible. The attorney used both ChatGPT and Westlaw Precision AI — two distinct research layers — and still submitted 11 fabricated or mischaracterized citations. The second layer was supposed to catch what the first missed. It didn't. Agentic workflows create many more of those intermediate layers. A citator check at the end of the chain doesn't retroactively verify what happened in the middle.
The Supervisory Partner Is the One Holding the Exposure

ABA Formal Opinion 512, issued in July 2024, is the first comprehensive ABA guidance on generative AI. Six ethical obligations flow from it: competence, confidentiality, communication, candor toward the tribunal, supervisory responsibilities, and fee constraints around time spent learning tools the firm uses regularly.
The supervisory-responsibility clause is the one with immediate sanctions exposure.
The supervising partner who signs a filing carries personal sanctions exposure for AI output generated by any attorney on the matter — whether the partner reviewed the session logs, whether they knew which tool was used, or whether they reviewed the output at all.
The Sixth Circuit's March 2026 $30,000 ruling distributed sanctions across two attorneys. At least one was not the person who drafted the problematic citation work. Signing the filing is enough.
Malpractice carriers are adjusting. Published analysis from Wiley's legal team has flagged AI-usage exclusion language appearing in new E&O policy language. Renewal questionnaires now include AI usage disclosure sections. Firms running Harvey's agents or CoCounsel's agentic workflows without documented verification protocols may find coverage gaps at precisely the moment they need coverage most.
The 68% of legal professionals who have used unapproved AI tools at least once — a North Carolina Bar Association figure — aren't just creating compliance risk. They're creating a coverage risk that sits on the supervisory partner's shoulders, often without the supervisory partner's knowledge.
What a Verification Pipeline Is Actually Doing

The solution is not a better AI platform. Every platform's marketing implies the next model will close the gap. The Stanford study's peer-reviewed numbers — from independent academic researchers rather than the vendors themselves — suggest the improvement trajectory isn't fast enough for firms operating at scale in high-stakes practice areas.
The verification layer sits between AI output and human review. What we've built at Veriprajna's legal AI verification and governance practice starts from the Stone v. Ritter failure mode: a standard retrieval system surfaces the case but not its downstream citation context.
GraphRAG — a knowledge graph of citing references — changes what's checkable. The 14% retrieval relevance improvement over standard vector retrieval in legal-context testing isn't about speed; it's about being able to interrogate whether the proposition the AI attributed to a case has been narrowed, distinguished, or functionally overridden by subsequent decisions that a citator would never flag. The knowledge graph retrieves the case and its downstream legal evolution. Standard RAG retrieves only the case.
Alongside that, 300+ judges operating under different AI requirements make standing order compliance a data problem as much as a legal one. Some courts require only disclosure. W.D. North Carolina effectively bars generative AI for drafting. Florida added a state-level mandate in February 2026. Manually tracking what each court requires — and generating compliant certification language per judge, per matter — is a workflow that should be automated rather than delegated to the associates already under filing deadline pressure. The pipeline maintains a current map of active standing orders.
The audit trail that runs through all of this is also what ABA 512's supervisory-responsibility obligation requires in practice. Every AI-assisted research session should produce a record of which model ran which queries, which citations were surfaced, and which verification checks passed or flagged. That log is what your general counsel will need to point to — and what a malpractice carrier will request when a claim is filed.
The Regulatory Window That's Already Closing

The EU AI Act begins enforcement in August 2026. Colorado's AI Act takes effect in June 2026. Multiple state legislatures have bills moving that, per Wiley's analysis, treat AI systems as "products" subject to strict liability — meaning firms deploying legal AI without adequate verification infrastructure may face product-liability exposure, not just malpractice risk.
Fewer than one in five law firms have a formal AI policy today. Nearly 70% of legal professionals now regularly use generative AI tools. That gap between adoption and governance has produced 1,222 documented court incidents in the Charlotin database — and the number is accelerating.
Firms building verification infrastructure now — before the Colorado and EU deadlines, before the next Sixth Circuit ruling — are doing it as competitive positioning. The firms that wait are doing it under deadline pressure, with whatever budget remains after a sanctions ruling.
The verification architectures that hold up across practice areas and jurisdictions are still being built. If your technology committee is working through how to close the gap between what your AI platforms deliver and what ABA 512, the Sixth Circuit, and the incoming regulatory layer actually require — particularly around agentic workflow QA and multi-jurisdiction standing order mapping — the verification and governance frameworks we've developed may be a useful reference point. We'd be interested in what your team is finding. The jurisdictional patchwork is where most governance frameworks have the least coverage, and what firms collectively document here will shape what "adequate" looks like when August 2026 arrives.