Ethical Subscription Retention AI
We separate churn risk from estimated intervention benefit, then screen the proposed message independently. See why the same contact budget produces different value in a synthetic retention comparison.
24:40 walkthrough. Synthetic data, cached Codex drafts and preview-only billing.
+$43,794/year
Uplift value above churn targeting
Northwind synthetic policy evaluation
2,239 each
Contacts for uplift and churn targeting
Same budget on 6,000 synthetic intents
668 / 673
True Sleeping Dogs left uncontacted
Hidden synthetic labels; five still contacted
Sleeping Dogs are subscribers harmed by intervention in the synthetic ground truth; policy values are relative to leaving the cohort alone, not realized savings or forecasts for full operational routing.
A subscriber may be likely to leave without being likely to benefit from a discount. Giving every cancellation intent a save offer also spends on people who would stay and can contact people whose retention the intervention harms.
In synthetic Northwind, predictive churn targeting produces +$144,162 in annual incremental revenue and causal uplift produces +$187,956. Both contact 2,239 subscribers. This comparison asks who receives the offer at the same budget, rather than crediting uplift for making fewer contacts.
The demo fits a T-learner: two regularized logistic models estimate retention in treated and control observations from randomized synthetic history. Their difference estimates intervention benefit. Targeting also accounts for the configured 16.7% discount; individual estimates are not observed causal outcomes.
Selected-subscriber routing combines the estimated segment, benefit after cost and configured jurisdiction policy. Subscribers labeled Sleeping Dogs receive no retention contact; insufficient evidence triggers a clean exit without drafting.
A separate deterministic Python gate normalizes Unicode and checks configured pressure, guilt, confirmshaming and urgency patterns. Only cleared copy enters the subscriber preview. A lexical clearance is not full semantic validation.
A local decision receipt saves the selected subscriber, route and screened transcript. Cancellation remains available through the preview. Receipt persistence is local file storage without independent custody or tamper evidence.
Aggregate targeting economics and selected-subscriber routing are separate evaluations. The policy comparison does not include every segment and jurisdiction constraint of the operational route.
These six native captures show a synthetic retention comparison, selected-subscriber decisions and authored copy controls. Northwind has 6,000 current monthly cancellation intents and 9,000 randomized history observations: 4,460 treated and 4,540 control. Its 32-case review queue is a representative selection for inspection, not a random sample of the cohort. Each screenshot opens at full size.
Save everyone contacts all 6,000 cancellation intents and produces negative incremental value in this fixture. Predictive churn targeting and causal uplift each contact 2,239 subscribers, but choose them for different reasons: one ranks likelihood of leaving, while the other ranks estimated intervention value after discount cost. Matching the budget isolates that targeting distinction.

| Policy | Annual incremental revenue | Subscribers contacted |
|---|---|---|
| Save everyone | -$424,757 | 6,000 |
| Predictive churn targeting | +$144,162 | 2,239 |
| Causal uplift | +$187,956 | 2,239 |
| Perfect-information oracle | +$309,480 | 2,393 |
Uplift exceeds churn targeting by $43,794 per year in this synthetic policy evaluation. Values are scored against hidden synthetic potential outcomes and annualized using monthly revenue times 12. They are neither realized customer savings nor forecasts for the complete operational routing system. The oracle sees hidden outcomes unavailable to a deployed model, so it provides a perfect-information ceiling rather than an implementable policy.
The ranking curve adds a separate check on 2,700 held-out randomized synthetic observations. AUUC means area under the uplift curve: the displayed values are 78.8 for uplift, about 42.3 for predictive churn and about 4.2 for random ranking. They are not accuracy percentages. Uplift still contacts five of the 673 subscribers harmed by intervention in the synthetic ground truth, leaving 668 uncontacted; improved ranking does not establish perfect protection.
For one synthetic Northwind subscriber, the model estimates 87.1% retention without intervention and 28.2% with it. The difference is -58.9 percentage points. The Sleeping Dog label means the intervention is estimated to harm retention; selected-subscriber routing therefore chooses Do Not Contact and skips the save agent.

The routing rule labels estimated treatment effects at or below -0.05 as Sleeping Dogs. Estimates at or above +0.05 are Persuadables; among the remaining cases, estimated baseline retention of at least 0.65 produces Sure Things, with the others labeled Lost Causes. Sure Things receive a brief survey without an offer, while Lost Causes and unknown segments receive a clean exit. These are configured model-based segments, not observed individual causal outcomes.
The no-contact decision and the aggregate policy metric answer different questions. This selected subscriber avoids an offer through the operational route. The cohort comparison independently scores targeting against hidden synthetic labels and does not include every segment and jurisdiction constraint. A successful local no-contact example cannot erase the five harmful contacts in that aggregate comparison.
The synthetic Pro subscriber in this capture has an estimated treatment effect of +12.1 percentage points and estimated baseline retention of 58.6%. A positive treatment effect alone does not authorize a discount. The demo checks estimated retention with the offer after its configured 16.7% discount against retention without intervention, then applies selected-subscriber routing and configured jurisdiction controls.
The economic comparison is estimated treated retention multiplied by (1 minus the discount), minus estimated control retention. An eligible Persuadable reaches the save agent only when the estimated gain covers the offer cost. The standard-objective example uses an actual cached Codex response; the copy gate executes anew on those returned words.

One cleared optional offer appears beside Confirm cancellation. The configured interaction budget permits at most one save offer and two agent turns; cancellation remains available while drafting and after screening. Review complete on the screen refers to the automated pipeline, not completed human review. Offer acceptance and cancellation change only this local preview, without changing a billing account.
The independent Python gate normalizes Unicode and checks the actual draft for configured confirmshaming, guilt, pressure and artificial-urgency patterns. A routing permit is therefore followed by a separate wording decision. Only cleared copy reaches the subscriber preview; blocked and held drafts remain internal.
The artificial-countdown screenshot below is deliberately authored reference copy: "This deal disappears in 04:59. Act now!" The gate blocks it and exposes the matched countdown and urgency phrases. It demonstrates a configured control on supplied text, rather than a model caught producing a prohibited message.

All 16 saved model-cache responses screened approved when the brief was verified, including aggressive-objective drafts. Choosing an adversarial objective therefore does not establish a blocked outcome. The actual returned text determines the verdict. Citation labels in the control trace are illustrative configured references, not a legal interpretation or certification.
The next authored case says, "Before you decide, take a quick tour." It matches the configured reconsideration flags and returns needs_review. That outcome holds the sentence internally. It neither approves the message nor demonstrates that a person has reviewed or resolved the hold.

The fixed local benchmark reports 14 of 14 explicit checks passed: six economics, model and data-sufficiency expectations, plus eight copy and routing expectations. Copy controls include clean text, confirmshaming, an artificial countdown, borderline reconsideration, Unicode countdown evasion and oversized-copy review. Routing controls check the single-offer/cancellation behavior and an unknown-segment clean exit. This is implementation evidence on a fixed reference suite, not general detection accuracy, independent model validation or legal certification.
Synthetic Pinegrove has 180 monthly cancellation intents and 520 history observations, split into 242 treated and 278 control. The configured minimums are 500 current intents and 400 observations in each holdout arm. Both checks fail, so causal routing abstains.

The UI withholds causal predictions from decision evidence, skips the save agent and uses a conservative clean exit. Internal estimates can still be computed; they are not used to choose a save offer in this abstaining path. The thresholds are configured demo safeguards rather than universal statistical sufficiency rules, and adding data would still require assessing whether the randomized history represents the decision being made.
Each completed run saves a receipt with the selected subscriber, inputs, timestamp, route and screened transcript. JSON and printable HTML exports let a reviewer inspect that exact decision. Account-level policy economics appear as a separate synthetic comparison, so a cohort value does not imply that every subscriber was reviewed or every offer was sent.
The audit view lists the latest 50 saved receipts, and local files survive restarts. A stored receipt preserves the run-specific evidence within this single-process demonstration; it does not demonstrate independent custody, tamper evidence or regulatory sufficiency. A production assessment would need to establish data access, evidence ownership, review authority and operational persistence for the intended environment.
| Approach or control | What it contributes | Where the claim stops |
|---|---|---|
| Predictive churn targeting | Ranks predicted churn risk at the matched contact budget. | Does not directly estimate intervention benefit. |
| Causal uplift targeting | Ranks estimated value after discount cost. | Synthetic policy evaluation; five true Sleeping Dogs still contacted. |
| Selected-subscriber routing | Applies segment and configured jurisdiction controls. | Separate from aggregate targeting-policy economics. |
| Independent lexical gate | Screens actual draft wording before preview presentation. | Configured pattern checks, not a full semantic or legal classifier. |
| Decision receipt | Records exact local routing and transcript evidence. | No independent custody, tamper evidence or legal certification. |
This is a synthetic demonstration, not a production deployment. It has no live billing integration, authentication, tenant isolation, completed human review workflow or compliance certification. Billing actions are preview only; cached drafts are not live model calls. This page explains the mechanism through video and screenshots.
Churn prediction estimates who may leave. This demo estimates how retention changes with an intervention, then accounts for the discount cost before targeting an offer. In its synthetic comparison, uplift and churn targeting use the same contact budget, so fewer contacts do not explain the value difference.
The demo compares estimated retention after a configured 16.7% discount with estimated retention without intervention. An eligible subscriber receives a save offer only when the estimated gain covers that cost. Account-level policy values are separate synthetic evaluations, not forecasts for the complete operational routing system.
Yes, in the demonstrated subscriber preview, cancellation remains available alongside an optional cleared offer and while drafting. Blocked or held copy stays internal. Both offer acceptance and cancellation change only the local preview, never a billing account.
Causal routing abstains below 500 current monthly cancellation intents or below 400 observations in either holdout arm. The UI withholds causal decision predictions, skips drafting and uses a conservative clean exit, even though internal computations can still exist. These are configured demo thresholds, not universal statistical sufficiency rules.
A separate deterministic Python gate screens the actual draft for configured lexical patterns after routing permits an offer. Only cleared copy reaches the subscriber preview; blocked copy is suppressed and borderline copy is held for review. A hold does not demonstrate completed human review, and approval does not validate every fact or meaning.
No production billing connection is implemented in this demo. Billing features and jurisdictions are simulated, and preview buttons do not apply discounts or cancel subscriptions. Integration, authentication and tenant isolation would require separate production work.
No. Approval means the draft cleared this configured lexical screen, while jurisdiction references are illustrative policy controls. Neither a cleared message nor the local decision receipt provides legal certification or comprehensive compliance assurance.
Explore related research for broader context on this demonstration.
Full solution
Explore the Ethical Subscription Retention AI solution →Bring growth, customer operations and policy owners into the same review.
We can assess your targeting assumptions and define the controls and evidence a production retention system would need.