CacheMagpie logoCacheMagpie· Community Library

Retention email tests and learnings

Retention tests pay most on relevance and restraint. Cutting sends to the people who are not reading, tying messages to a real event rather than the calendar, and surfacing value someone already owns beat adding volume. Frequency increases are the most commonly run and most commonly disappointing test in this stage.

  • What the retention stage is for and the metric you own
  • Test ideas grouped by lever, from cadence to relevance
  • The retention tests most teams skip, and how to run one properly
  • Guard metrics, and graded evidence from real programs

What the retention stage is actually for

Retention messaging exists to keep a relationship worth having. Its job is to make the next order, the next session or the next renewal feel natural, and to do that without spending the attention you will need later.

The metric you own is repeat rate or active rate over a period you can defend, paired with revenue per contact. Open and click rates are especially misleading here, because the healthiest retention programs often send less.

This is the stage where compounding happens. A cadence change that looks small in a single month decides list health a year out, which is why the read window for retention tests has to be longer than for any other stage.

In CacheMagpie, every test in this cell records its own primary metric, and the metric you own here is repeat rate and revenue per contact. Check that a test's metric lines up with yours before you copy the learning.

Where retention sits in the lifecycle

Retention sits after activation and before winback. Every contact you fail to keep here becomes a winback problem later, and winback is the most expensive stage in the lifecycle.

What you can test at the retention stage

Each group below is one variable. A clean test changes one of them and holds the rest still, which is the difference between a learning and a story.

Cadence and volume

  • Test sending less to disengaged contacts against sending the same to everyone, and read revenue per contact rather than total sends.
  • Test an added send against a holdout, with unsubscribes and complaints read at the same time.
  • Test a frequency preference centre against a fixed cadence.

Relevance and targeting

  • Test messages triggered by a real account event against calendar campaigns.
  • Test recommendations based on behaviour against category best sellers.
  • Test suppressing campaigns for contacts who bought or acted in the last few days.

Value and non-price levers

  • Test surfacing an unused benefit or feature the customer already pays for.
  • Test a useful summary of the customer's own activity against a promotional message.
  • Test guidance or education against an offer, read on repeat rate rather than immediate revenue.

Format and length

  • Test one idea per message against a digest with several.
  • Test a shorter message with a single action against the full template.
  • Test plain text styling for the messages that carry information rather than merchandising.

Lifecycle triggers

  • Test a replenishment or renewal reminder timed on the customer's own cycle rather than an average.
  • Test a proactive message when usage drops against waiting for the lapse window.
  • Test a post purchase sequence that sets up the next order rather than only confirming the last one.

What the graded tests say about aggressive defaults

Across the graded tests in this stage, targeting and segmentation, timing and delay, shorter copy carry more of the winning reads.

The highest-value retention tests most teams skip

Retention roadmaps fill with campaigns, so these structural tests keep getting postponed.

Sending less

Reducing cadence for disengaged contacts is rarely tested because it feels like giving up revenue. Measured on revenue per contact over a full quarter it is one of the most reliable wins in the stage.

The permanent holdout

A small group that receives no campaigns is the only way to know what the program is actually worth. Very few teams run one, and the ones that do stop arguing about attribution.

Event triggered replacements

Replacing a recurring campaign with an event triggered message is a bigger change than any subject line test and usually produces a larger, longer lived effect.

How to run a retention test

Randomise contacts rather than sends, and read over a period long enough to capture at least one purchase or renewal cycle. Retention tests read on a single campaign are measuring a campaign, not retention.

Set the sample size and the read date before the test ships, not after you have seen the first day of data. If the audience cannot reach the sample you need inside a sensible window, change the test rather than the standard: pick a lever with a bigger expected effect, or widen the trigger.

A test card template for Retention

Hypothesis
Because contacts in [segment] disengage after [signal], sending [variant] instead of [control] will increase revenue per contact over [period].
Success criteria
Revenue or active rate per contact over a full cycle, measured on the assigned population rather than on openers.
Guard metric
Unsubscribe rate, complaint rate and the size of the engaged sendable list.

Free to copy and use in your own program. Fill the brackets from a test you have read in the library.

Guard metrics: what a win must not cost

Retention wins are easy to fake by sending more. These reads stop that.

  • Unsubscribe and complaint rate per contact, not per send.
  • Engaged list size at the end of the period, which is the asset the program runs on.
  • Revenue per contact rather than revenue per campaign.
  • Inbox placement, since volume increases show up here first.

Evidence from real Retention programs

The library holds 41 graded tests run by Retention teams: 39 report a win for the variant, 1 a loss, and 1 no clear difference.

Winning tests in this cell read in line with the rest of the library, across 17 comparable tests.

Direction, grade and source are free. Exact figures open up once you sign in.

Retention tests by channel

The same stage behaves differently depending on the channel carrying it. These counts are published tests in this cell.

What a winning retention test looks like

A winning retention test moves revenue or active rate per contact over a full cycle while keeping the engaged list at least as large as it started.

The most durable winners reduce or retarget volume rather than adding it, because relevance keeps working and pressure does not.

Frequently asked questions

Does retention messages still work in 2026?

The library holds 41 graded tests here, and 39 of them report a win for the variant against 1 losses and 1 with no clear difference. That is enough to say the direction still holds, and not enough to promise a number for your own program.

What is the most credible retention test in the library?

Push vs self-scheduled reminders for daily practice habit. Algorithmically-timed push beat self-scheduling for daily-habit adherence (an exact amount vsan exact amount), but the authors warn the crutch can undercut intrinsic habit and over-frequency fatigues users.

What tends to fail in retention messages?

For engagement re-activation, more reminders and generic incentives backfired; already-disengaged readers need something genuinely special or should be let go, not nagged. Losing tests are kept in the library on purpose, because knowing what did not move is as useful as knowing what did.

How fresh are these retention learnings?

The newest record in this cell is from 2026 and the oldest from 2016. 0 records have been reviewed and vouched for by the CacheMagpie editor.

Which metric should I judge retention tests on?

The metric you own for this stage is repeat rate and revenue per contact. Every test in this cell records its own primary metric, so check that it matches yours before you copy the learning.