Email marketing tests and benchmarks
Email is the only lifecycle channel where almost everything is testable, which is why most programs test the least consequential parts of it. Audience, trigger and cadence decide outcomes far more than subject line wording, and the largest untested surface in nearly every program is transactional mail. Test who receives a message and when before you test what it says.
- Every testable surface in an email program, from audience to footer
- Test ideas for each lifecycle stage, linked to the matching stage page
- Deliverability and consent constraints that shape what you can run
- How fast email learnings fade, and graded evidence from real programs
What is actually testable in email
An email test is only clean when you know which of these you changed. Most inconclusive tests changed three at once.
Audience and trigger
Who is eligible, what event starts the message, and who is suppressed. This is the highest leverage surface in the channel and the least often treated as a variable.
Timing and cadence
Send time, delay after the trigger, spacing between messages in a sequence, and how many messages the sequence contains.
Sender and envelope
From name, reply-to, subject line and preheader. Cheap to test, quick to read, and the surface where effects decay fastest.
Body and structure
Length, one idea against several, image weight, plain text against template, and where the primary action sits.
Offer and value
Incentive size and type, non-price levers such as shipping, guidance or reassurance, and whether an offer appears at all.
Transactional content
Confirmations, receipts, shipping notices and alerts. The most opened mail you send and, in most programs, the only mail nobody has ever tested.
The anatomy of a testable email
Who is eligible and what event starts the send.
Brand, person, or brand plus person.
Claim, question, specificity, length.
Extension of the subject, or a second claim.
Absolute clock time or time since the trigger.
Reason for the message, stated or implied.
One idea or several, image weight, length.
Placement, wording, and how many there are.
Present or absent, price or non-price.
Preference options, frequency choice, unsubscribe clarity.
Every slot above is a variable you can hold constant or change. A clean email test moves one of them at a time.
Email test ideas by lifecycle stage
Each stage below has its own page with the full set of levers, guard metrics and graded evidence.
- Test the series length against a single message, read on first purchase rather than opens.
- Test holding the incentive back to the last message so it only reaches people who did not act.
Onboarding
See the onboarding page (6 tests here)- Test prompts triggered by real progress against a fixed schedule.
- Test one action per message against a checklist.
Cart abandonment
See the cart abandonment page (18 tests here)- Test the delay before the first reminder across the full recovery window.
- Test naming a specific blocker, such as shipping or returns, against a generic reminder.
- Test reducing cadence for disengaged contacts, read on revenue per contact.
- Test event triggered messages replacing a recurring campaign.
- Test asking what went wrong against offering a reason to return.
- Test a hard stop after two touches, with complaint rate as the primary read.
Constraints that decide what you can test in email
Email has the fewest hard limits of any lifecycle channel, but the soft limits are real and they punish volume.
- Reputation is shared across your whole program, so a bad campaign degrades every later send from the same domain.
- Consent and suppression rules under GDPR and similar regimes decide who is eligible, which constrains audience tests before creative ones.
- Bulk sender requirements from the major inbox providers set complaint rate ceilings and one click unsubscribe expectations.
- Open rate is no longer a clean metric where mail privacy pre-fetches images, so it cannot be the deciding read.
- Rendering varies widely between clients, so a design test needs a rendering check before a results check.
How fast email learnings fade
Email learnings are the most durable in the lifecycle stack, but not uniformly. Structural findings about audience, trigger and cadence hold for years, because they describe how people behave rather than how a format feels.
Wording and format findings fade faster. A subject line style that stands out does so because it is unusual, and once it is common it is not unusual. Treat a subject line win as valid for a season, and re-run it rather than enshrining it.
Anything tied to inbox provider behaviour, tabs, clipping, image handling or open measurement, can change without notice. Keep the source date of a learning in view before you copy it.
Email is the slowest channel to burn out, but tests that rely on a new format or a surprise subject line still fade as the list gets used to them. Reads here are worth revalidating once a year.
Format learnings from the corpus
- Subject line length and clarity are the most repeated test in the corpus, and short and specific tends to beat clever.
- Send time tests read smaller than most teams expect once the audience is already engaged.
How to run a email test
Randomise on the contact and hold the audience definition constant across arms. Read on the outcome the message exists to produce, not on the open, and let a sequence finish before you judge it.
Set the sample size and the read date before the test ships, not after you have seen the first hour of data. If the audience cannot reach the sample you need inside a sensible window, change the test rather than the standard.
A test card template for Email
- Hypothesis
- Because contacts in [segment] respond to [reason], changing [one variable] from [control] to [variant] will increase [outcome] within [window].
- Success criteria
- The downstream outcome, measured on everyone assigned rather than on openers or clickers.
- Guard metric
- Unsubscribe rate, complaint rate and inbox placement.
Free to copy and use in your own program. Fill the brackets from a test you have read in the library.
Guard metrics: what a win must not cost
Email wins that cost reputation are borrowed, not earned.
- Complaint rate, which the major providers treat as a hard ceiling.
- Unsubscribe rate per contact rather than per send.
- Inbox placement and bounce rate on the sending domain.
- Engaged list size at the end of the test period.
Evidence from real Email programs
The library holds 106 graded tests run by Email teams: 97 report a win for the variant, 4 a loss, and 5 no clear difference.
Winning tests in this cell read in line with the rest of the library, across 65 comparable tests.
- Plain-text style subject line beat the branded promo subject in win-back
In win-back, a plain, human subject line outperformed the branded promo phrasing on opens without hurting unsubscribes.
Grade ANo clear difference2026 - Coupon-email frequency
Maximizing promo-email revenue this quarter directly trades against churn, so frequency is a lifetime-value decision, not a revenue one.
Grade AVariant lost2026 - Free-trial duration
Trial urgency is largely a myth; users convert because they reached value, so a longer trial lifts delayed conversion without hurting immediate conversion, and trial length should be paired with the right promo type.
Grade AVariant won2026 - Fewer coupon emails cut churn but dented short-term sales
Cutting send frequency is a churn-prevention lever with a real short-term revenue tax, so it only pays off when you optimize for lifetime value, not this month.
Grade AVariant won2026 - Fewer coupon emails
Send frequency is a churn lever with a real revenue tax, so it only pays when optimized for lifetime value rather than this month's sales.
Grade AVariant won2026 - Feedback channel beats coupon for re-subscription
For subscription churn, a discount can read as unfair and underperform simply giving lapsed users a channel to be heard.
Grade AVariant won2024 - Name-in-subject-line effect fails to replicate
A rigorous 2023 replication could not reproduce the name-in-subject-line lift, and other studies find recipients react negatively to identifiable data, so the tactic is contested, not settled.
Grade ANo clear difference2023 - First-name in welcome subject line lifts opens
Personalizing the welcome subject with the guest's first name lifted opens an exact amount with no downstream drag. The win compounded into a small first-booking uplift.
Grade AVariant wonAirbnb2019 - Targeting high-risk churners can backfire
Who you contact matters more than the offer; sending a renewal push to the highest-risk members can actively push them out, so target by responsiveness, not risk.
Grade AVariant won2018 - Recipient name in subject line lifts opens and leads
A piece of non-informative personalization in the first touch moved actual leads an exact amount, not just opens, by raising attention to the rest of the message.
Grade AVariant won2018 - Recipient name in subject line lifts opens and leads
Non-informative personalization in the very first touch is nearly free and moves revenue metrics (leads an exact amount, froman exact amount toan exact amount), not just vanity opens.
Grade AVariant won2016 - Proactive onboarding education halves first-week churn
A single proactive onboarding education touch halved first-week churn and lifted 8-month usagean exact amount, and it also cut support load, so the effort partly pays for itself.
Grade AVariant won2016 - Instant welcome send against a two hour delay
Send the welcome message immediately. The two hour delay lost most of the signup intent without improving unsubscribes.
Grade BVariant wonNorthline Supply2026 - Behavior-triggered onboarding
Switching onboarding from calendar-based to behavior-triggered lifted trial conversionan exact amount by meeting users at their actual friction point.
Grade BVariant wonCitrix2026 - Subject-line testing for newsletters
The act of subject-line testing itself lifted average opens an exact amount, a floor-level return on running the test at all.
Grade BVariant wonThe Remote Company2026 - Disabling the open-tracking pixel
The open-tracking pixel that measures engagement can itself suppress it via deliverability, so turning it off lifted replies at the cost of open visibility.
Grade BVariant wonBelkins2026 - Cold email length
Across 16.5M emails, the 101-200 word band replied at an exact amount versus an exact amount for 600+ word emails, so brevity is a reliable acquisition lever.
Grade BVariant wonBelkins2026 - Question-format subject line vs statement (MailerLite)
Phrasing the subject as a direct question to the reader beat the equivalent how-to statement on opens for this newsletter audience.
Grade BVariant wonMailerLite2026 - Feedback-guided cold email rewrite doubles reply rate
Reading the actual negative replies and rewriting to address them nearly doubled reply rate and flipped sentiment positive, more reliable on small lists than chasing reply-rate significance.
Grade BVariant wonAcme Advisors & Brokers2026 - Pattern-interrupt win-back email, no discount, at Getir
Creative novelty beat discounting for 30-day+ churned users, and it protected margin by winning them back at full price.
Grade BVariant wonGetir2025 - Wistia onboarding email overhaul triples paid conversions
Rewriting the onboarding emails to each drive one action tripled paid conversions on the same product, evidence the copy layer alone moves activation.
Grade BVariant wonWistia2022 - Discount win-back made less net profit than no discount
A discount can win the conversion and lose the P&L, so judging win-back by reactivations rather than margin flatters the wrong variant.
Grade BVariant lost2020 - Offer-led vs generic welcome email subject line
In the first welcome email, a tangible named offer in the subject line beats the word 'welcome' itself.
Grade BVariant won2012 - Expanding welcome from 1 email to a 6-day series
You can raise welcome cadence aggressively if the first email sets the expectation and every email carries a proven offer; the guardrail is complaints, not opens.
Grade BVariant won2012 - Orange link color in welcome emails, effect splits by gender
A cosmetic change produced a large lift in one segment and essentially nothing in another; segment-level readouts can flip the conclusion of a 'small' design test.
Grade BVariant won2012 - Fewer links plus deadline layout
Choice reduction only produced a significant win once paired with a real deadline; nav removal alone was inconclusive atan exact amount confidence.
Grade BVariant won - Grubhub Campus staged welcome stream vs holdout
Measuring a welcome flow against a true control group, not against opens, is what lets you claim incremental activation.
Grade BVariant wonGrubhub - Removing footer navigation plus scarcity in lifecycle emails
Navigation removal alone was inconclusive (an exact amount at onlyan exact amount confidence); the significant lift came only when fewer links were combined with a real deadline, so choice reduction is an amplifier, not a standalone lever.
Grade BVariant won - Sensitivity-based retention targeting cuts wasted spend
Some customer segments churn more when messaged, so a good retention program must detect and exclude them, echoing the Ascarza finding with a newer method.
Grade BVariant won - ProdPad behavior-segmented onboarding doubles conversion
Segmenting onboarding by early behavior and rewarding ideal actions doubled conversion, giving unengaged trials a path they otherwise missed.
Grade BVariant wonProdPad
Direction, grade and source are free. Exact figures open up once you sign in.
What a winning email test looks like
A winning email test moves the outcome the message exists to produce, holds unsubscribes and complaints flat, and comes with a documented audience and window so someone else can repeat it.
The wins worth keeping are structural. If the result depends on a specific wording, expect to re-run it within a year.
Frequently asked questions
Does Email messages still work in 2026?
The library holds 106 graded tests here, and 97 of them report a win for the variant against 4 losses and 5 with no clear difference. That is enough to say the direction still holds, and not enough to promise a number for your own program.
What is the most credible email test in the library?
First-name in welcome subject line lifts opens. Personalizing the welcome subject with the guest's first name lifted opens an exact amount with no downstream drag. The win compounded into a small first-booking uplift.
What tends to fail in Email messages?
Maximizing promo-email revenue this quarter directly trades against churn, so frequency is a lifetime-value decision, not a revenue one. Losing tests are kept in the library on purpose, because knowing what did not move is as useful as knowing what did.
How fresh are these email learnings?
The newest record in this cell is from 2026 and the oldest from 2012. 1 record has been reviewed and vouched for by the CacheMagpie editor.
