
A lawyer in New Orleans used both ChatGPT and Westlaw's AI on his brief. He still got sanctioned for 11 bad citations.
Here's the part most firms miss: the scary failure isn't the invented case. Citators like KeyCite and Shepard's catch those — the citation resolves to nothing in the database.
The dangerous one is the case that's real.
The AI cites Stone v. Ritter for a director-oversight standard. The case exists. KeyCite shows a green flag. Everything looks right. But Marchand v. Barnhill (2019) quietly expanded the Caremark oversight duty Stone established, and the analysis built on it is wrong for a 2026 filing.
That's contextual hallucination — a real citation used for a proposition it doesn't actually support. It's what the Stanford study's 33% Westlaw Precision hallucination rate is really measuring. Not fake cases. Wrong analysis of real ones.
A junior associate reviewing under deadline won't catch it, because the citation looks perfect.
So we don't think the answer is a "better" AI tool. Whether a firm runs Harvey, Lexis Protege, or open models, the gap is the same: nothing sits between the AI's output and the partner's signature to check whether the cited case actually says what the brief claims.
That layer — checking subsequent treatment, not just citator status — is what we build.
If you're in practice: how is your firm verifying AI citations today — a manual read of every one, or trust in the tool's green flag?
#LegalAI #AIGovernance