A retention system can produce a polite message for an offer that should never have been made. Before judging its writing, I want to know why it believes intervention will help this subscriber.
My design standard is that withholding an offer needs a reason precise enough to guide the next decision. Estimated harm, insufficient evidence and unacceptable wording can all leave the subscriber without an offer. They call for different responses from the people operating the system.
When better wording cannot fix the decision
In our Ethical Subscription Retention AI demo, a synthetic subscriber has estimated retention of 87.1% without intervention and 28.2% with it. The difference is negative: the model estimates that contacting this subscriber would reduce retention. These are model estimates from synthetic data, not observed individual counterfactuals.
The configured route records no retention contact and skips the drafting agent. Cancellation remains available in the local preview. The important design choice is the order: the system decides whether an offer belongs here before asking an agent to make it persuasive.
Synthetic subscriber evidence: estimated retention falls from 87.1% to 28.2% with intervention, so the configured route is Do Not Contact. This is an estimate, not an observed individual outcome.
If the intervention itself is expected to harm retention, rewriting the message does not address the reason it was withheld. A product team needs to examine the treatment estimate and the routing rule. Sending the case back to a copy generator would solve the wrong problem.
That distinction also changes how I read a policy comparison. In the seeded cohort, predictive churn targeting and causal uplift both contact 2,239 subscribers. Uplift produces $43,794 more annual incremental value in the synthetic evaluation. The difference comes from who is selected at the same contact budget, rather than fewer contacts. It is evidence for testing selection quality separately from contact volume.
The comparison uses hidden synthetic outcomes and evaluates aggregate targeting policies, not the complete subscriber routing system or realized revenue. Its imperfection matters too: uplift still contacts five of the 673 subscribers whose synthetic retention is harmed by intervention. I would not accept an evaluation that reported the value improvement while hiding those harmful selections.
Missing evidence needs a different response
A second synthetic account fails the demo's configured data-sufficiency thresholds. Its interface withholds causal predictions from decision evidence, skips the agent and offers a clean exit.
This is not another negative treatment estimate. It is a decision to abstain because the evidence does not meet the configured requirements. Combining both cases into a single “no offer” count would conceal whether the system had a reason to expect harm or lacked enough evidence to support the recommendation.
For an operational review, I want those reasons counted separately. Estimated harm calls for scrutiny of intervention effects and routing. Abstention calls for scrutiny of the evidence requirement and the available data. Relaxing that requirement just to increase offer coverage would change the standard of evidence, even if the resulting messages looked acceptable.
An eligible offer still has to clear the copy gate
For a different synthetic subscriber, the demo supports an optional offer after accounting for the discount cost. A deterministic lexical gate outside the drafting model then screens the actual words before they reach the subscriber preview.
A cached model draft cleared the lexical gate for this synthetic subscriber. Accept offer and Confirm cancellation remain separate choices; both buttons affect only the local preview and do not change billing.
All 16 saved model-cache responses cleared the current gate, including drafts under the adversarial maximize-saves objective. The demo's blocked and held examples come from explicitly authored benchmark copy cases, not those model responses. A held message awaits review; the demo does not complete human review. Clearing the lexical screen does not certify every fact, meaning or legal requirement.
This gives a third reason to withhold an offer: the proposed message fails a configured screen even though the subscriber is eligible. That case belongs in a copy-control review. It should not be counted as proof that the intervention estimate was wrong, or that human review has already happened.
For builders evaluating retention AI, my acceptance criterion is a decision record that makes these differences recoverable: the evidence supporting eligibility or abstention, the route chosen, and the actual draft with its screening result when drafting occurs. The demo saves a local run-specific receipt for that purpose. It demonstrates inspectable decisions, with no production billing integration or tamper-evident evidence service.
The full breakdown of Ethical Subscription Retention AI explains the mechanism and its limits.
A single offer-rate metric cannot tell a team whether to improve targeting, gather more evidence or repair its copy controls. Those decisions become possible when the system preserves the reason an offer was withheld.