I put a correct HbA1c of 6.8% beside a “continue metformin” instruction in a synthetic patient-portal draft; the latest eGFR was 28. I built that contradiction into ChartSieve because a factually accurate sentence can make the next, unsafe sentence easier to trust.
I wanted to see whether a clinical AI safety check would stop at “the lab value matches” or follow the draft all the way to the action it recommends. In this case, the deterministic policy gate holds the message for clinician review. No message is sent to a patient, and the demo does not show a clinician completing that review.
I put the true number next to the risky instruction
I started with a patient-portal draft that reads like reassuring follow-up. The HbA1c value matches the synthetic patient record. The medication is also on the medication list. If I inspected only those two facts, the draft would look well grounded.

The case view keeps the plausible draft and the held disposition on the same screen.
I then read the recommendation as a recommendation, not merely as text to be matched. The latest eGFR is 28, and the curated renal rule flags the metformin instruction for review. The record also shows creatinine rising from 1.4 to 1.9 to 2.3 mg/dL. That trend supports the renal context visible to a reviewer; it does not independently determine the disposition. The rule fires on the latest eGFR.
That distinction changed how I framed this demo. “Grounded” is a useful description for a claim that matches a source. It is a dangerous shorthand for “safe to send” when a message moves from reporting a value to advising a patient what to do. The case needs two checks that can disagree: did the draft state the record accurately, and does its proposed action survive a patient-specific safety rule?
I followed the claim rows into the rule finding
I stopped at two green, grounded rows: 6.8% matched the record, and metformin was on the medication list. For a moment that looks like a clean result. Then my eye moves to the renal finding: latest eGFR 28, metformin advice flagged. I cannot put one “grounded” verdict over this whole draft without hiding that conflict. I kept the claim rows grounded and let the separate CONTRAINDICATION_RENAL rule set HOLD_FOR_REVIEW. The screen now holds an awkward but necessary combination: true evidence and withheld advice.

The correct lab and listed medication remain visible while the renal rule supplies the reason for the hold.
I made the policy gate deterministic because the final disposition should not depend on whether a verifier agent phrases its advice forcefully or misses a detail. The demo's verifier agents add grounding, equity, and red-team advisories. They can enrich the review, but they cannot override the rule-driven outcome. That separation is especially important in a case where a fluent draft contains both truth and risk.
This is a curated seed ruleset, not a maintained clinical knowledge base. The record is synthetic, the vendor output is a stub, and the hold is a simulated workflow state. I would not use this result to claim that ChartSieve is clinically validated or ready to govern actual patient communications. The point the build demonstrates is narrower and more useful: a safety gate can preserve correct evidence while refusing the action that evidence does not justify.
I opened the receipt to test the explanation
I wanted the hold to be inspectable without asking anyone to take the firewall's word for it. In the Safety Receipt for this same case, I can see the HOLD_FOR_REVIEW verdict, the grounded claims and their available source text, the CONTRAINDICATION_RENAL finding, and the verifier advisories. The demo records a timestamp and a SHA-256 digest truncated to 16 hexadecimal characters. That is a compact demo receipt, not a cryptographic signature or an immutable production audit trail.

The Safety Receipt exposes the evidence and rule finding behind this seeded hold; its digest is not a production audit guarantee.
I keep coming back to the apparently reassuring HbA1c line in the draft. A reader can verify the 6.8% and still miss the recommendation attached to it. The fixed regression set makes the same limitation visible at a different scale: a baseline that checks stated lab values catches 4 of 22 unsafe artifacts, while ChartSieve's seeded rule checks hold or block 22 of 22 on that 34-artifact synthetic set. Those are bounded test results, not clinical performance estimates.
The full ChartSieve breakdown shows the case and its evidence trail. I built the metformin example to make one review habit harder to skip: after confirming a clinical AI draft's facts, inspect the action those facts are being used to justify.