
I was reading the Sixth Circuit's March 2026 sanctions order — $30,000, two attorneys — when I realized we'd spent most of the past year building a verification product for a problem that was already largely solved.
The order wasn't about fabricated citations. The cases existed. The docket numbers were real. KeyCite returned green flags. The problem was that the AI had cited real cases for propositions they didn't actually support. That's contextual hallucination — and it's categorically different from what Mata v. Avianca introduced to the legal profession in 2023. Harvey and Lexis+ with Protege ground their output in real case databases. Shepard's and KeyCite catch the case that doesn't exist. But they weren't designed to verify whether the AI accurately characterized what a real case held.
We had been building a verification layer that caught citation fabrication well and contextual hallucination inadequately. The Sixth Circuit order made that concrete in a way that a 33% hallucination statistic, sitting in a peer-reviewed journal, hadn't.
What Testing Westlaw Precision Taught Me About the 33%

I started testing legal AI platforms systematically after seeing the Stanford/JELS 2025 study — not because the 33% number for Westlaw Precision was a surprise, but because I wanted to understand what the 33% was actually made of.
The fabricated-citation portion was the minority. The dominant pattern was subtler: the AI retrieved a real case, summarized a portion of its holding accurately, and cited it for a proposition that the case either didn't support, supported only in a narrower context than the AI implied, or that subsequent decisions had substantially modified without overruling the original case.
Stone v. Ritter (2006) is the example that clarified it for me. A litigation associate researching director oversight liability under Delaware law builds on Stone — real case, accurate KeyCite, broadly correct on 2006 Delaware law. What neither the AI nor the associate flags is that Marchand v. Barnhill (2019) expanded the Caremark duty, and subsequent Chancery opinions have built a "mission critical" compliance standard that changes the practical application of Stone for any 2026 Delaware filing. Stone is still good law. The analysis built on Stone alone is wrong.
A citator doesn't catch this. Shepard's and KeyCite track direct negative treatment: reversal, overruling, significant criticism. The gradual narrowing of a holding's scope across a line of decisions that distinguish rather than overrule is invisible to them. And Lexis+ with Protege, which replaced Lexis+ AI in February 2026 after LexisNexis walked back its "100% hallucination-free" marketing language, still carries a 17% contextual hallucination rate on the same study methodology. Better than Westlaw's 33%. Wrong one time in six.
What I was building for was the AI inventing a case. What the Sixth Circuit's $30,000 ruling was actually about was the AI misreading a real one.
The Knowledge Graph We Had to Build When RAG Wasn't Enough

We tried to build the citation-context layer on standard vector retrieval first. The failure case was Stone v. Ritter: vector search surfaced the case, surfaced Marchand v. Barnhill, and had no way to establish whether and how Marchand changed what Stone practically means. The relationship between the cases — what subsequent opinions had done with the original holding — lives in the citing reference network, not in the text of the cases themselves.
That's what sent us to GraphRAG and the knowledge-graph architecture at Veriprajna's legal AI verification practice. The 14% retrieval relevance improvement over standard vector RAG isn't the headline benefit. The headline is that the graph can be traversed — you can ask not just "what is this case" but "what has happened to this case's core proposition in subsequent citations," and surface the specific line of decisions that distinguished or narrowed the original holding.
It's the same data Shepard's and KeyCite use to track reversal and overruling — but applied to the narrower and harder question of propositional drift. The question the verification layer answers isn't "is this case good law?" It's "is this holding still understood the way the AI described it?"
The GC Question I Didn't Have a Clean Answer To

My first substantive conversation with a litigation firm GC about legal AI verification went where I expected — Mata, the Sixth Circuit order, the published hallucination rates. The question that caught me flat-footed came about thirty minutes in: what happens to supervisory partners when the AI is wrong and they've signed the filing. The GC wasn't asking abstractly.
I knew the broad answer — the partner is exposed. What I didn't know precisely was the scope. So I went back to ABA Formal Opinion 512 more carefully than I had before.
Opinion 512 was issued in July 2024 as the ABA's first comprehensive guidance on generative AI. Six ethical obligations. The one that changed how I thought about the product we were building is the supervisory-responsibility clause: the supervising partner who signs a filing carries personal sanctions exposure for AI output generated by any attorney on the matter — whether the partner reviewed the session logs, whether they knew which tool was used, whether they reviewed the output at all. Signing the filing is the act that confers exposure.
The Sixth Circuit's $30,000 ruling distributed sanctions across two attorneys. At least one was not the person who generated the problematic citation work. That's what Opinion 512's supervisory clause looks like in practice.
What changed in my thinking after that conversation: the GC's question wasn't really about the AI. It was about governance — specifically whether the firm had any documented evidence that a verification step had occurred between AI output and the filing. Published analysis from Wiley's legal team has since flagged AI-usage exclusion language appearing in E&O policy renewals. Firms running Harvey's agents or CoCounsel's agentic workflows without documented verification protocols may find their malpractice coverage excludes the incident they're trying to claim on.
Why Harvey's 25,000 Agents Changed My Architecture Assumptions

The thing that changed my view on the verification problem's scope wasn't a new hallucination statistic — it was watching Harvey's agent demos in late 2025.
Harvey's March 2026 funding round valued it at $11 billion, with $190 million in ARR. More than half of Am Law 100 runs it. The platform hosts 25,000 custom agents — multi-step autonomous workflows that chain research, synthesis, drafting, and issue-flagging together. Thomson Reuters CoCounsel launched agentic "Deep Research" workflows in early 2026. LexisNexis's Protege, which replaced Lexis+ AI in February 2026, has four specialized agents and 300+ pre-built workflows.
The verification problem I'd been thinking about was a query-level problem: AI runs a search, retrieves citations, human verifies before filing. Agentic workflows make it a chain problem. An agent running a multi-step research task might execute a dozen queries in sequence, synthesize the results, draft preliminary analysis, and surface a completed memo. Every intermediate citation in that chain needs to hold. If any one step in the middle relied on a contextually hallucinated holding, the final output can be coherent and wrong — and there's no obvious seam where the human review should have happened.
The New Orleans case from February 2026 made this pattern explicit. The attorney used both ChatGPT and Westlaw Precision AI — two separate research layers — and still submitted 11 fabricated or mischaracterized citations. A citator at the end of an agentic chain doesn't retroactively verify what happened at each intermediate step. That's the architecture problem I didn't fully see until the agents started shipping.
The August Deadline That's Now on My Whiteboard

I have two dates written on my whiteboard that aren't tied to a client deliverable: August 2026 (EU AI Act enforcement begins) and June 2026 (Colorado AI Act takes effect).
These aren't background pressure. Wiley's analysis has flagged bills moving in multiple state legislatures that treat AI systems explicitly as "products" subject to strict liability — meaning firms deploying legal AI tools without adequate verification infrastructure may face product-liability exposure on top of malpractice risk. Colorado and the EU are the first deadlines; they won't be the last.
What changes in August 2026 isn't that legal AI becomes illegal. What changes is that the audit trail, the governance documentation, and the verification record that most firms currently don't produce systematically becomes a regulatory expectation. Fewer than one in five law firms have a formal AI policy today. Nearly 70% of legal professionals regularly use generative AI. The 68% who have used unapproved tools at least once — a North Carolina Bar Association figure — are creating a coverage risk on their supervising partners' shoulders right now, whether or not those partners know it.
The 1,222 court incidents documented in the Charlotin database through early 2026 are a trailing indicator. The leading one is that Westlaw Precision's 33% hallucination rate and the agentic workflows being deployed on top of it are structural, not accidental. The architectures being built now will determine how much of that risk gets caught before it becomes a filing.
I keep thinking about the GC's underlying question as the one the industry is circling but not quite landing on. The citation tools answer the fabrication version reasonably well now. The contextual version — wrong holding for a real case, in an agentic chain, on a matter the supervising partner signed but didn't personally research — is where the exposure sits, and it's where verification architecture has to be built deliberately, not patched after the next sanctions order.
If your firm is working through how to close that gap, the legal AI verification frameworks we've built are one reference point. What I'm genuinely more curious about is what the firms that have already tried to build this internally found when they hit the citation-graph problem. The architectures that hold up are still being mapped, and the field learns faster when those experiments are shared.