The one question I couldn't answer for an ordinary AI SDR
I kept picturing the same scene while I built this: a compliance partner sitting next to a VP of Sales, pointing at a single line in an outbound email, asking a question that sounds trivial and is not. Prove which source backed that line, and how current was it. I could not make an ordinary AI SDR answer it. The thing drafted, it sent, and it kept no record of why any particular claim was allowed to leave the building. I had already learned the hard way, earlier in the build, that you cannot trust a model to grade its own homework, so this was a different worry: even when the email is fine, the tool cannot prove it was fine.
That gap is not cosmetic. Enterprise AI-SDR churn runs somewhere around 50 to 70 percent a year (UserGems, 2026), and only 7 percent of enterprises have governance built specifically for agentic AI (Deloitte, 2026). Those two numbers belong together. Teams are shipping autonomous outbound at volume while keeping almost no provenance on what got said. When Gmail started rejecting mail at the SMTP level in November 2025 and a spam rate over 0.3 percent buys you a 6 to 12 week domain recovery, "we cannot show our work" stops being a governance footnote and becomes an operational one.
What I decided a receipt had to contain
I started the schema by writing down what a plain log line gives you, then crossing out everything on that list that wasn't actually accountability, and what was left is what earned the word receipt. A log tells you an email went out. I wanted a record that could tell a stranger, without me in the room, why each claim in it was allowed to leave. So the receipt names the model that drafted the email, down to provider and version, because "an AI wrote it" is not a defense anyone can check. It carries the prospect and the risk tier. It embeds the fact sheet the writer was constrained to, and then, for every factual claim in the draft, the verdict, the exact source span that did or did not back it, the source dates, and the recency gap in days. It ends with the veracity score, the specific policy rule that fired, the human approver where one was required, and the CRM write-back, which is stubbed in the demo but keeps its field anyway, because I did not want the record to look complete when a step was only simulated.
The verdict vocabulary is deliberately small and blunt: supported, stale, entity_mismatch, contradicted, unsourced. Only supported survives. Everything else is stripped and, more to the point, is stripped on the record. The receipt is downloadable JSON, one click, one file per email. There is a real sample sitting on disk from a run on claude-opus-4-8. That was the design bet: a claim should trace back to its source in seconds, by anyone, without my help.
One email's downloadable receipt, open on screen. Every claim carries its verdict, the cited source and its published date, recency_days, and each check with the exact expression it computed. One click on Download JSON, one file.
The claim I traced back in seconds
I only believed the receipt worked when I clicked Download JSON on the Northwind email and followed a killed claim back to its source in about the time it takes to read this paragraph. Northwind Logistics is a synthetic mid-market logistics company I built so I could show the pattern without putting words in a real firm's mouth. Its draft claimed the company had "recently expanded into APAC." The grounding overlap between claim and source was a clean 100 percent, the entity matched, and the claim still failed, because the backing source is dated 2019-03-14, which against the demo's working date is 2,652 days old, more than seven years. Recency language over a seven-year-old source is stale, and the receipt does not just say stale. It hands you the source span, the recency_days, and the rule that fired. The word "recently" was the lie, and the receipt names the lie with a number.
The evidence panel for the "recently expanded into APAC" line. Grounding overlap is 100 percent and the entity matches, yet Temporal validity fails: source age 2652 days over the 365-day limit against a 2019-03-14 source. First failure wins, so the claim is stripped.
The same machinery caught a real one. I loaded two paragraphs from Werner Enterprises' actual FY2023 Form 10-K (SEC EDGAR, CIK 0000793074, filed 2024-02-26) into the fact sheet by hand. Werner is a genuine public company, one of the few real records in the demo. The draft line "recently growing your One-Way Truckload fleet to 2,735 trucks" is factually correct and more than two years stale, so it is stripped, and the receipt carries the filing date so you can see exactly why. This is where two numbers stop being interchangeable. The veracity score is supported claims over total claims in the draft, and it is often well short of 100, which is the honest part. Sent integrity is 100 percent by construction, because whatever ships contains only source-backed claims. Provenance is what lets me say that out loud instead of hoping.
The two numbers side by side on the Northwind email. The veracity score sits at 60 percent of the draft while sent integrity holds at 100 percent, because the two stale and wrong-entity claims are struck through and stripped and only source-backed lines survive.
The clean email I watched get held anyway
I watched an email I could not fault get stopped anyway, and that scene reorganized how I think about all of this. Atlas Capital Markets is a synthetic FINRA-regulated broker-dealer in the demo, the contact is a Chief Revenue Officer, and the deal is 220,000 dollars. Its draft came out spotless. Every claim supported, nothing stripped, a perfect veracity score. And the policy gate routed it straight to human review, because the rule is that regulated, or C-suite, or deal size at or above 100,000 dollars sends the email to a person regardless of how clean the draft is.
A draft I could not fault. Every claim verified, veracity score at 100 percent, nothing stripped, and the gate still routes it to human review because Atlas is regulated, the contact is C-suite, and the deal is 220,000 dollars.
My first instinct was that the gate had a bug. It did not. I had quietly assumed that a correct email is a safe email, and Atlas is the case that proves those are separate properties. A claim can be fully sourced and still carry regulatory weight that no automated check should clear on its own. That is why the receipt reserves a field for the human approver: on the high-risk path, the record is not complete until a named person signed it. Governance is about calibrated risk, not just about whether the draft happened to be right.
Why I don't think a stronger model closes this
I used to assume a better base model would eventually make all of this unnecessary, and building it is what talked me out of that. A smarter writer does draft fewer bad lines, and I am glad of it. But provenance is not a capability you can train into weights. A perfect model, one that never invents a fact, still cannot prove to a FINRA examiner or a GDPR auditor which current source backed which claim in an email that already went out, and prove it the same way twice. That proof lives in the system around the model: the checks, the gate, the signed record. The 11x collapse is the cautionary version of ignoring that, 74 million dollars raised and churn reported at 70 to 80 percent when it came apart in 2025 (TechCrunch), and Gartner projects that 40 percent or more of agentic-AI projects will be abandoned by 2027. Volume without a receipt is a short story every time.
I hold myself to a narrow set of claims in return, because for a company named for true wisdom the honesty is the product. The tool does not promise zero hallucination, the model still drafts, and anyone selling you a zero is not being straight. On the 25-case labeled golden set I built to stress the verifier, it returns the right verdict 25 out of 25 times, deterministically, which is a statement about that fixed benchmark and not a promise about your inbox. The connectors are simulated and the fact sheet is pre-built, which is why I say I loaded the Werner filing in by hand rather than that the demo pulled it live. If you want to open a receipt yourself, the AI Sales Intelligence demo lets you download one and read every field.
So the question I would leave a RevOps or compliance leader with is not whether your AI SDR writes a good email. It is this. If someone pointed at one line it sent last Tuesday and asked you to justify it, could your tool hand them a receipt that names the source, the date, and the rule, or would it just shrug?