Retention email tests and learnings
Retention tests pay most on relevance and restraint. Cutting sends to the people who are not reading, tying messages to a real event rather than the calendar, and surfacing value someone already owns beat adding volume. Frequency increases are the most commonly run and most commonly disappointing test in this stage.
- What the retention stage is for and the metric you own
- Test ideas grouped by lever, from cadence to relevance
- The retention tests most teams skip, and how to run one properly
- Guard metrics, and graded evidence from real programs
What the retention stage is actually for
Retention messaging exists to keep a relationship worth having. Its job is to make the next order, the next session or the next renewal feel natural, and to do that without spending the attention you will need later.
The metric you own is repeat rate or active rate over a period you can defend, paired with revenue per contact. Open and click rates are especially misleading here, because the healthiest retention programs often send less.
This is the stage where compounding happens. A cadence change that looks small in a single month decides list health a year out, which is why the read window for retention tests has to be longer than for any other stage.
In CacheMagpie, every test in this cell records its own primary metric, and the metric you own here is repeat rate and revenue per contact. Check that a test's metric lines up with yours before you copy the learning.
Where retention sits in the lifecycle
Retention sits after activation and before winback. Every contact you fail to keep here becomes a winback problem later, and winback is the most expensive stage in the lifecycle.
What you can test at the retention stage
Each group below is one variable. A clean test changes one of them and holds the rest still, which is the difference between a learning and a story.
Cadence and volume
- Test sending less to disengaged contacts against sending the same to everyone, and read revenue per contact rather than total sends.
- Test an added send against a holdout, with unsubscribes and complaints read at the same time.
- Test a frequency preference centre against a fixed cadence.
Relevance and targeting
- Test messages triggered by a real account event against calendar campaigns.
- Test recommendations based on behaviour against category best sellers.
- Test suppressing campaigns for contacts who bought or acted in the last few days.
Value and non-price levers
- Test surfacing an unused benefit or feature the customer already pays for.
- Test a useful summary of the customer's own activity against a promotional message.
- Test guidance or education against an offer, read on repeat rate rather than immediate revenue.
Format and length
- Test one idea per message against a digest with several.
- Test a shorter message with a single action against the full template.
- Test plain text styling for the messages that carry information rather than merchandising.
Lifecycle triggers
- Test a replenishment or renewal reminder timed on the customer's own cycle rather than an average.
- Test a proactive message when usage drops against waiting for the lapse window.
- Test a post purchase sequence that sets up the next order rather than only confirming the last one.
What the graded tests say about aggressive defaults
Across the graded tests in this stage, targeting and segmentation, timing and delay, shorter copy carry more of the winning reads.
The highest-value retention tests most teams skip
Retention roadmaps fill with campaigns, so these structural tests keep getting postponed.
Sending less
Reducing cadence for disengaged contacts is rarely tested because it feels like giving up revenue. Measured on revenue per contact over a full quarter it is one of the most reliable wins in the stage.
The permanent holdout
A small group that receives no campaigns is the only way to know what the program is actually worth. Very few teams run one, and the ones that do stop arguing about attribution.
Event triggered replacements
Replacing a recurring campaign with an event triggered message is a bigger change than any subject line test and usually produces a larger, longer lived effect.
How to run a retention test
Randomise contacts rather than sends, and read over a period long enough to capture at least one purchase or renewal cycle. Retention tests read on a single campaign are measuring a campaign, not retention.
Set the sample size and the read date before the test ships, not after you have seen the first day of data. If the audience cannot reach the sample you need inside a sensible window, change the test rather than the standard: pick a lever with a bigger expected effect, or widen the trigger.
A test card template for Retention
- Hypothesis
- Because contacts in [segment] disengage after [signal], sending [variant] instead of [control] will increase revenue per contact over [period].
- Success criteria
- Revenue or active rate per contact over a full cycle, measured on the assigned population rather than on openers.
- Guard metric
- Unsubscribe rate, complaint rate and the size of the engaged sendable list.
Free to copy and use in your own program. Fill the brackets from a test you have read in the library.
Guard metrics: what a win must not cost
Retention wins are easy to fake by sending more. These reads stop that.
- Unsubscribe and complaint rate per contact, not per send.
- Engaged list size at the end of the period, which is the asset the program runs on.
- Revenue per contact rather than revenue per campaign.
- Inbox placement, since volume increases show up here first.
Evidence from real Retention programs
The library holds 41 graded tests run by Retention teams: 39 report a win for the variant, 1 a loss, and 1 no clear difference.
Winning tests in this cell read in line with the rest of the library, across 17 comparable tests.
- Push vs self-scheduled reminders for daily practice habit
Algorithmically-timed push beat self-scheduling for daily-habit adherence (an exact amount vsan exact amount), but the authors warn the crutch can undercut intrinsic habit and over-frequency fatigues users.
Grade AVariant won2026 - Personalized push (owner + pet name) lifts app engagement
Personalization compounds (owner plus pet name beat owner name alone), but layering in behavioral history on top of names added little, so identity cues do most of the work.
Grade AVariant won2026 - Fewer coupon emails cut churn but dented short-term sales
Cutting send frequency is a churn-prevention lever with a real short-term revenue tax, so it only pays off when you optimize for lifetime value, not this month.
Grade AVariant won2026 - Fewer coupon emails
Send frequency is a churn lever with a real revenue tax, so it only pays when optimized for lifetime value rather than this month's sales.
Grade AVariant won2026 - Behavioral text + email nudges for aid renewal (FAFSA)
Loss-aversion and plan-making SMS/email nudges lifted deadline renewals, and the paper shows targeting the right students matters as much as the nudge itself.
Grade AVariant wonCUNY / ideas422023 - Targeting high-risk churners can backfire
Who you contact matters more than the offer; sending a renewal push to the highest-risk members can actively push them out, so target by responsiveness, not risk.
Grade AVariant won2018 - Time-varying push effect on same-day app engagement (mHealth MRT)
A push reliably lifts same-day engagement, but the lift shrinks the more habituated or already-active the user is, so blanket sending wastes the effect on people who did not need it.
Grade AVariant won2018 - Personalized app-notification promotions by engagement stage
Push works better when timed to where the user is in their engagement lifecycle rather than sent on a broadcast schedule, and the stage can be modeled.
Grade AVariant won2016 - Location-based mobile promotions drive 6-12x purchases
Location-triggered push drove large same-day and 12-day-delayed purchase lifts, and measuring only the immediate response undercounts the true effect by half.
Grade AVariant won - Subject-line testing for newsletters
The act of subject-line testing itself lifted average opens an exact amount, a floor-level return on running the test at all.
Grade BVariant wonThe Remote Company2026 - Duolingo notification optimization cuts daily churn
Notification optimization delivered a large churn reduction, but the discipline was testing carefully to avoid burning the push channel, which is the real constraint.
Grade BVariant wonDuolingo2026 - Notification optimization at Duolingo cuts churn
Sustained notification testing delivered a large engagement/churn win, with the discipline being to protect the push channel rather than maximize sends.
Grade BVariant wonDuolingo2026 - Question-format subject line vs statement (MailerLite)
Phrasing the subject as a direct question to the reader beat the equivalent how-to statement on opens for this newsletter audience.
Grade BVariant wonMailerLite2026 - Sensitivity-based retention targeting cuts wasted spend
Some customer segments churn more when messaged, so a good retention program must detect and exclude them, echoing the Ascarza finding with a newer method.
Grade BVariant won - Push notifications for retention
Lifecycle push can nearly triple 90-day retention, but the same channel over-sent makesan exact amount of users kill notifications, so the guardrail is opt-out, not open rate.
Grade CVariant wonAggregate2026 - Email marketing optimization
Moving from manual sends to data-driven, timely segmentation doubled conversions and tripled email revenue, a whole-program transformation rather than an isolated test.
Grade CVariant wonHerman Miller2026 - Loyalty program presence
A loyalty program lifts repeat rate 15an exact amount by adding an extrinsic reward, strongest when the first reward is reachable within a purchase or two.
Grade CVariant wonAggregate2026 - Replenishment reminders vs generic promo conversion
Timing a reminder to the moment the customer actually runs low converts several times better than a promo blast, because relevance is the whole mechanism.
Grade CVariant wonAggregate2026 - Multi-pack / subscribe-and-save captures 2nd purchase upfront
For consumables where 82an exact amount of repeat purchases are the identical product, converting the second purchase to a multi-pack upfront removes the repurchase decision entirely.
Grade CVariant wonAggregate2026 - Post-purchase flow extended to the 90-day repurchase window
Since an exact amount of repeat purchases happen within 90 days, a flow that stops at day 14 covers a fraction of the window, and extending it lifts second orders 20an exact amount.
Grade CVariant wonAggregate2026 - Non-price retention levers vs discounts by segment
Matching the lever to the segment (credits for high-value, discounts only for price-responsive) recovers churn without training the whole base to expect discounts.
Grade CVariant wonAggregate2026 - Renewal reminder sequence reduces chargebacks and churn
Pre-renewal reminders convert surprise into expectation, cutting chargebacks 40an exact amount and lifting renewals 5an exact amount, with an annual-upgrade bonus.
Grade CVariant wonAggregate2026 - Value/usage-summary in renewal reminder vs plain reminder
Showing the outcomes a customer already achieved makes the renewal harder to walk away from than a bare 'renews soon' notice.
Grade CVariant wonAggregate2026 - Escalating dunning sequence for involuntary churn
Much subscription churn is involuntary (expired cards), so a timed multi-channel card-update sequence recovers revenue that has nothing to do with satisfaction.
Grade CVariant wonAggregate2026 - Sharing scripts vs strategy content in newsletter
Testing content type, not just subject lines, surfaced that copy-paste-ready scripts massively out-engaged conceptual strategy content.
Grade CVariant won2026 - Images/GIF position at top of email vs lower
Visual placement is a testable engagement lever; leading with an image or GIF was run specifically to move click rate.
Grade CVariant wonMailerLite2026 - Sunset/suppression raises aggregate engagement and deliverability
Counterintuitively, mailing fewer engaged people beats mailing everyone, because providers reward aggregate engagement and punish dead weight across your entire domain.
Grade CVariant wonAggregate2026 - Over-reminding disengaged subscribers backfires
For engagement re-activation, more reminders and generic incentives backfired; already-disengaged readers need something genuinely special or should be let go, not nagged.
Grade CVariant lost2026 - Gmail Promotions annotations
Adding pre-open annotations in Gmail's Promotions tab lifted opensan exact amount across thousands of campaigns, a cheap deliverability-surface lever independent of subject line.
Grade CVariant wonAggregate2026 - Key takeaways at top layout
Restructuring so readers can skim the key points at the top revived a low-engagement newsletter, layout as an engagement lever independent of content.
Grade CVariant won2025
Direction, grade and source are free. Exact figures open up once you sign in.
Retention tests by channel
The same stage behaves differently depending on the channel carrying it. These counts are published tests in this cell.
What a winning retention test looks like
A winning retention test moves revenue or active rate per contact over a full cycle while keeping the engaged list at least as large as it started.
The most durable winners reduce or retarget volume rather than adding it, because relevance keeps working and pressure does not.
Frequently asked questions
Does retention messages still work in 2026?
The library holds 41 graded tests here, and 39 of them report a win for the variant against 1 losses and 1 with no clear difference. That is enough to say the direction still holds, and not enough to promise a number for your own program.
What is the most credible retention test in the library?
Push vs self-scheduled reminders for daily practice habit. Algorithmically-timed push beat self-scheduling for daily-habit adherence (an exact amount vsan exact amount), but the authors warn the crutch can undercut intrinsic habit and over-frequency fatigues users.
What tends to fail in retention messages?
For engagement re-activation, more reminders and generic incentives backfired; already-disengaged readers need something genuinely special or should be let go, not nagged. Losing tests are kept in the library on purpose, because knowing what did not move is as useful as knowing what did.
How fresh are these retention learnings?
The newest record in this cell is from 2026 and the oldest from 2016. 0 records have been reviewed and vouched for by the CacheMagpie editor.
Which metric should I judge retention tests on?
The metric you own for this stage is repeat rate and revenue per contact. Every test in this cell records its own primary metric, so check that it matches yours before you copy the learning.
