Welcome series A/B tests and learnings
The welcome series is the one moment where attention is guaranteed, so the tests that pay are structural rather than cosmetic. Sequence length, what the first message asks for, and how fast the first message lands move first purchase and long term engagement more than subject line wording. Test the shape of the series before you test the words inside it.
- What the welcome stage is for and the metric you own
- Test ideas grouped by lever, from timing to sequence length
- The welcome tests most teams skip, and how to run one properly
- Guard metrics, and graded evidence from real programs
What the welcome stage is actually for
A welcome series has one job: turn a fresh permission into a first real action. That action is a purchase in retail, a completed profile in a marketplace, a first funded account in fintech, or a first meaningful session in a product. Everything else, including opens and clicks, is a proxy that will happily go up while the real number stays flat.
The metric you own here is the rate at which new contacts complete that first action inside a fixed window, usually 14 or 30 days. Fix the window before you test anything, because a longer window flatters slow sequences and a shorter one flatters aggressive ones.
The stage is also where you set the expectation for every message that follows. Cadence, tone and the amount you ask for in the first week teach a new subscriber what your program is, and that lesson is hard to unteach later.
In CacheMagpie, every test in this cell records its own primary metric, and the metric you own here is welcome to first order rate. Check that a test's metric lines up with yours before you copy the learning.
Where welcome sits in the lifecycle
Welcome sits between acquisition and onboarding. What you promise at sign-up sets what the welcome series has to deliver, and what the welcome series teaches carries into every later stage.
What you can test at the welcome stage
Each group below is one variable. A clean test changes one of them and holds the rest still, which is the difference between a learning and a story.
Timing and delay
- Send the first message immediately against a short delay, and read completed first actions rather than opens.
- Test the gap between message one and two, since a tight gap works for high intent sign-ups and a loose one for content sign-ups.
- Test triggering from the sign-up event against triggering from the first session, when the two do not coincide.
Sequence length and cadence
- Test a three message series against a single message, and read unsubscribes alongside conversion.
- Test cutting the series short as soon as someone converts, rather than running it to the end.
- Test adding a final message that only goes to people who did nothing, instead of extending the series for everyone.
The ask in message one
- Test a single clear next action against a menu of options.
- Test asking for the purchase against asking for a smaller commitment such as a preference or a category choice.
- Test leading with proof, such as reviews or numbers of customers, against leading with the product itself.
Offer and non-price levers
- Test the welcome incentive against no incentive, and read repeat rate and margin rather than first order rate alone.
- Test holding the incentive back to a later message, so it reaches only the people who did not act.
- Test a non-price lever, such as free returns, guidance or a saved setup, against the discount.
Targeting and personalisation
- Test branching the series on the sign-up source, since a checkout sign-up and a pop-up sign-up are not the same person.
- Test using declared preference data collected in the first message against using nothing.
- Test suppressing the series entirely for people who already bought before it started.
What the graded tests say about aggressive defaults
Across the graded tests in this stage, targeting and segmentation, timing and delay, shorter copy carry more of the winning reads.
The highest-value welcome tests most teams skip
These tests rarely reach the roadmap because they need more coordination than a subject line, and they are the ones with the most room left in them.
Series length itself
Most teams inherit the number of messages from whoever built the program and never question it. Running the current series against a shorter one is a single test that tells you whether the extra sends are earning their unsubscribes.
The incentive
The welcome discount is often the largest recurring cost in the program and the least tested part of it. A holdout that receives no incentive, read on repeat purchase rather than first purchase, usually changes the conversation.
The exit condition
What happens when someone converts halfway through the series is rarely designed and almost never tested, yet it decides whether a new customer's first week feels attentive or mechanical.
How to run a welcome test
Split at the point of entry so both variants see the same mix of sign-up sources, and let the whole series run before you read anything. Reading a multi-message sequence on day two measures the first send, not the series.
Set the sample size and the read date before the test ships, not after you have seen the first day of data. If the audience cannot reach the sample you need inside a sensible window, change the test rather than the standard: pick a lever with a bigger expected effect, or widen the trigger.
A test card template for Welcome
- Hypothesis
- Because new subscribers from [source] act fastest in the first [window], sending [variant] instead of [control] will increase the share who complete [first action].
- Success criteria
- First action rate inside a fixed [14 or 30] day window, measured on everyone who entered the series, not on openers.
- Guard metric
- Unsubscribe and complaint rate across the full series.
Free to copy and use in your own program. Fill the brackets from a test you have read in the library.
Guard metrics: what a win must not cost
A welcome test that wins on first action and quietly burns the list is a loss. Read these alongside the primary metric.
- Unsubscribe and complaint rate measured across the whole series, not per message.
- Repeat purchase or second action rate, so a discount driven win is visible as a one off.
- Margin on first order, when the test touches an incentive.
- Deliverability signals on the sending domain, since a welcome series reaches every new address you collect.
Evidence from real Welcome programs
The library holds 11 graded tests run by Welcome teams: 11 report a win for the variant, 0 a loss, and 0 no clear difference.
Winning tests in this cell read above the rest of the library, across 9 comparable tests.
- First-name in welcome subject line lifts opens
Personalizing the welcome subject with the guest's first name lifted opens an exact amount with no downstream drag. The win compounded into a small first-booking uplift.
Grade AVariant wonAirbnb2019 - Instant welcome send against a two hour delay
Send the welcome message immediately. The two hour delay lost most of the signup intent without improving unsubscribes.
Grade BVariant wonNorthline Supply2026 - Offer-led vs generic welcome email subject line
In the first welcome email, a tangible named offer in the subject line beats the word 'welcome' itself.
Grade BVariant won2012 - Expanding welcome from 1 email to a 6-day series
You can raise welcome cadence aggressively if the first email sets the expectation and every email carries a proven offer; the guardrail is complaints, not opens.
Grade BVariant won2012 - Orange link color in welcome emails, effect splits by gender
A cosmetic change produced a large lift in one segment and essentially nothing in another; segment-level readouts can flip the conclusion of a 'small' design test.
Grade BVariant won2012 - Grubhub Campus staged welcome stream vs holdout
Measuring a welcome flow against a true control group, not against opens, is what lets you claim incremental activation.
Grade BVariant wonGrubhub - Lifecycle email rebuild
Rebuilding the lifecycle program lifted average opens from underan exact amount to overan exact amount, a whole-program transformation with a real before/after delta.
Grade CVariant wonRockin Royalty2026 - Welcome offer type: fixed vs percentage discount vs mystery vs gift
Judge the opt-in offer on welcome flow placed orders, not popup submits; offer type alone moved conversion a multiple, and dollar framing beat the mathematically similar percentage.
Grade CVariant won2026 - Gentle nurture vs aggressive offer welcome flow, supplements
Welcome flow aggressiveness should match business model economics; for consumables the profit lives in purchase two, so speed to first purchase beat brand nurture.
Grade CVariant won2026 - Vitrazza welcome flow rebuilt around testimonials
For a high-consideration product (luxury glass chair mats), social proof and founder voice inside the welcome flow did the persuasion work a discount usually does.
Grade CVariant wonVitrazza2024 - Three-email story-led welcome series at Island Olive Oil
The comparison is welcome flow vs promotional sends rather than a controlled test, but the magnitude (an exact amount conversion,an exact amount higher click rate,an exact amount higher revenue per email) shows where the automation priority sits.
Grade CVariant wonIsland Olive Oil
Direction, grade and source are free. Exact figures open up once you sign in.
Welcome tests by channel
The same stage behaves differently depending on the channel carrying it. These counts are published tests in this cell.
What a winning welcome test looks like
A winning welcome test shows a lift in completed first actions inside a fixed window, with unsubscribes flat or better, and it holds when you look at the second action rather than only the first.
It also tends to be structural. Changes to sequence length, the ask, or who receives the series survive far longer than a wording change, which is why they are worth the coordination cost.
Frequently asked questions
Does welcome messages still work in 2026?
The library holds 11 graded tests here, and 11 of them report a win for the variant against 0 losses and 0 with no clear difference. That is enough to say the direction still holds, and not enough to promise a number for your own program.
What is the most credible welcome test in the library?
First-name in welcome subject line lifts opens. Personalizing the welcome subject with the guest's first name lifted opens an exact amount with no downstream drag. The win compounded into a small first-booking uplift.
What tends to fail in welcome messages?
No losing test has been recorded in this cell yet, which is a gap rather than a signal. Treat the wins here as directional only.
How fresh are these welcome learnings?
The newest record in this cell is from 2026 and the oldest from 2012. 0 records have been reviewed and vouched for by the CacheMagpie editor.
Which metric should I judge welcome tests on?
The metric you own for this stage is welcome to first order rate. Every test in this cell records its own primary metric, so check that it matches yours before you copy the learning.
