
A mental-health chatbot called Tessa once told eating-disorder patients to maintain a 500–1,000 calorie daily deficit and buy skin calipers.
It wasn't a broken model. It was a well-prompted one, doing what these models are trained to do: be helpful, agreeable, engaging.
That's the trap: teams try to make a mental-health AI safe by writing better prompts. But the dangerous failures don't live inside any single message — they build across the conversation.
A user asks about "healthy eating." Safe. Then "counting calories." Probably safe. Then "how to hide food from my family." A system checking each message in isolation clears all three. One tracking the trajectory sees a crisis forming.
Sycophancy — validating a user's delusion to sound empathetic — is the complaint clinicians raise most. When OpenAI found one of its own updates "validating doubts, fueling anger, urging impulsive actions," it didn't reword the prompt. It withdrew the update.
As of April 2026, the FDA has authorized zero generative-AI devices for any clinical use.
In our work, safety in behavioral health is an architecture decision — stateful risk detection, output validation, graduated escalation — not a wording one. After the Character.AI settlements, "reasonable safety architecture" is becoming the de facto legal standard, and a prompt isn't an architecture.
If you've shipped AI into mental health: what's catching the cross-turn drift, not just the single bad message?
#ClinicalAISafety #DigitalHealth