Every message looked fine on its own. The crisis was in the pattern between them. A per-message check never sees it.
Behavioral-health chatbots get moderated one message at a time, each reply scored on its own. But a crisis rarely arrives in one line. It builds turn by turn, and no single message is alarming enough to stop. So a stateless moderator waves each through and reacts only once something is openly dangerous.
We built a demo of a safety layer that keeps a running, cross-turn read on where a conversation is heading. The part a safety team cares about: we tested it against a moderator running the exact same classifier and the same escalation policy. The only thing we added was memory across turns. On a 40-conversation labeled golden set, that one change caught the escalating conversation a median of two turns earlier, with zero false escalations across 29 benign turns. In two, the stateless moderator never flagged anything at all.
The conversations are synthetic and the numbers scoped to that golden set. It's a demo of an architecture pattern, not a medical device. But a stronger base model won't close the gap: the blind spot isn't in any one message, it's in the space between them.
If you're building or running a mental-health chatbot, does anything watch the conversation as it develops, or is safety still scored one message at a time? Curious how teams handle the slow-burn case.
#AISafety #TrustAndSafety #DigitalHealth #AIGovernance #MentalHealthTech
Published on Facebook · July 16, 2026
On social media
See this post on its original platform
In our archive