A per-message safety filter scores each message on its own. A mental-health crisis doesn't arrive in one message. It builds across the whole conversation.
We built a safety layer that wraps an existing behavioral-health chatbot, then ran it beside a stateless per-message moderator on the same conversation. Six turns, each an ordinary wellness question, slowly drifting toward an eating-disorder crisis.
Watch the split screen. The stateless moderator scores every message in isolation and stays green, because no single line is alarming enough to block. On the guarded side the risk meter climbs turn by turn, because it remembers the pattern. At turn 3 it crosses into CONCERN, blocks the reply, and substitutes a clinician-written grounding script. That is two turns before the identical stateless moderator reacts.
The honest part: both stacks run the same classifier and the same escalation gate. The only difference is memory. So the earlier catch comes from statefulness alone, not a smarter model. Across a 40-conversation labeled golden set, that held at a median of two turns earlier, with zero false escalations on 29 benign turns.
And the call is not a prompt. A deterministic 5-level policy gate the clinical team owns makes the decision, and every turn is written to a hash-chained, tamper-evident audit the platform can actually file. Agents advise, code decides.
Safety is an architecture problem, not a prompting problem. A better base model still can't see across turns, and can't hand you a receipt.
Every conversation here is synthetic and this is a demo, not a medical device. If your team runs behavioral-health AI, we'd genuinely like to hear how you catch the crisis that builds across turns. 🩺
#BehavioralHealth #DigitalHealth #AIGovernance #AISafety #ClinicalAI
Published on Instagram · July 16, 2026
On social media