ON THIS PAGE
By Efe Berke Colaker, Founder at GetleadReviewed by the Getlead editorial team for accuracy. Last updated August 2026.
Cold email testing usually proves nothing, because the underlying rates are low and the samples are small. Two hundred sends per variant cannot separate a genuine improvement from a quiet week.
Getting value from testing means accepting fewer, larger tests on the variables that actually move outcomes.
The arithmetic that makes small tests useless
Reply rates sit near 3% at median, which means a variant sent to 200 people produces about six replies. A difference of two replies is noise, and no confidence interval will rescue it.
As a working rule, plan for at least a thousand sends per variant before treating a difference as real, and expect to need more when the true effect is small.
For example, a 3% versus 4% comparison is a meaningful commercial difference and a statistically weak one at a few hundred sends. That combination is why most reported subject line wins do not repeat.
What to test, in order of effect size
- Segment. The largest lever by a wide margin, and the cheapest to change.
- Offer or ask. A smaller ask often outperforms a better argument.
- First line relevance. Situation specific versus category generic.
- Sequence length and spacing, since roughly 42% of replies arrive from follow-ups.
- Subject line. Real but small, and the most tested by a distance.
The ordering is inverted in practice. Teams test subject lines weekly and segments never, which is testing the variable with the smallest effect and the noisiest measurement.
Designing a test that answers something
- One variable. If segment and copy both change, the result attributes to neither.
- Same week. Provider behaviour and prospect attention shift over time.
- Same mailboxes. Sending infrastructure differences swamp copy differences.
- Positive replies as the metric, since total replies include people asking to be left alone.
- Pre committed sample size, decided before the send rather than when the numbers look good.
The last rule prevents the most common self deception: watching a test until it says something pleasant, then stopping. Decide the volume first and read the result once.
Sources and method
First-party data (Getlead, 2026): the verification split of 43.4% confirmed valid, 23.9% invalid, 16.7% catch-all and 16.0% unknown comes from 383,368 addresses analyzed through live SMTP verification, and the 0.51% bounce rate and 35.8% open rate come from 34,973 tracked sends, aggregated and anonymized at campaign level. Full method in our cold email benchmark study.
External sources: median B2B reply rates near 3%, the 5 to 8% range for strong campaigns and the finding that roughly 42% of replies arrive from follow-ups come from 2026 cold email benchmark compilations; US commercial email obligations come from the FTC CAN-SPAM compliance guide.
Third party benchmarks vary in methodology and should be read as directional. Checked in August 2026.
Frequently asked questions
How many sends does a cold email A/B test need?
At least a thousand per variant as a working minimum. At a median 3% reply rate, 200 sends produce about six replies, so a two reply difference between variants is noise rather than a result.
What should I test first in cold email?
The segment. Targeting produces far larger swings than copy, and it is cheaper to change. Offer, first line relevance and sequence spacing follow, with subject lines last despite being tested most often.
Can I test on open rates?
No. Privacy features load tracking images without a human involved, and our tracked sends average 35.8% opens, so an open rate difference between variants often measures provider behaviour rather than recipient interest.
Why do subject line wins fail to repeat?
Because the effect is small and the measurement is noisy. A 3% versus 4% comparison is commercially meaningful and statistically weak at a few hundred sends, so most reported wins are chance restated as insight.
How do I avoid fooling myself in a test?
Commit to the sample size before sending, change one variable, run both arms in the same week from the same mailboxes, and judge on positive replies rather than total replies or opens.
Is sequence length worth testing?
Yes, and it is under tested. Roughly 42% of replies arrive from follow-ups in published benchmarks, so changes to sequence length and spacing move a larger share of the outcome than most copy changes do.
