
A car dealership's chatbot agreed to sell a $76,000 Chevy Tahoe for $1. The buyer just typed the right words.
He'd told the bot to agree with anything the customer says and call every reply "a legally binding offer." Then he set his budget at $1. The bot said yes.
Two months later, Air Canada's chatbot invented a bereavement refund that didn't exist. A tribunal held the airline liable and rejected its argument that the bot was a "separate legal entity."
Both bots had a system prompt telling them how to behave. Neither had a logic layer deciding what they were allowed to do.
A system prompt lives in the same text stream as the attack, so the most persuasive wording wins. A joint OpenAI, Anthropic and Google DeepMind study found every published prompt-injection defense bypassed more than 90% of the time.
The fix isn't a smarter prompt. It's moving the decision into code. A rule like "if offer < MSRP × 0.9: reject" compares two numbers. No clever phrasing changes an if-statement.
This is the layer the market keeps skipping. The $740M identity-and-access deal that closed this year stops an unauthorized agent. It does nothing about an authorized one that hallucinates a price or a refund window.
Our research found 88% of enterprises hit a confirmed or suspected AI agent incident last year. The gap isn't ambition. It's architecture.
If you've shipped a customer-facing AI, what's actually stopping it from promising something your policy forbids?
#AIGovernance #EnterpriseAI