Cold email A/B testing guide (what to test and how)
Most cold email tests cannot produce an answer at the volume they run at. The audit that tells you which, and what to change when the answer is no.
Most cold email A/B tests cannot produce an answer, and the reason is arithmetic rather than discipline. The volume required to separate two reply rates scales with the inverse square of the gap between them, and the gap produced by rewording a sentence is small. Before designing a test, run the audit below: it tells you whether the comparison you are planning can resolve at the volume you have, and if it cannot, what to test instead so the quarter is not spent generating numbers that mislead you later.
What an A/B test on cold email is actually comparing
Two proportions. That is the whole statistical object: the fraction of arm A that replied against the fraction of arm B that replied. Everything difficult follows from three questions about that object.
What is the unit of randomisation? The thing you flip a coin over. Usually a lead.
What is the unit of analysis? The thing you count. Also usually a lead — but if you count replies at account level while randomising at lead level, the arithmetic no longer holds.
What is the population? Not "my list". The population is the set of records that reached the step being tested, which at step one is your list and at step four is a heavily filtered residue of it.
Get those three aligned and the standard two-proportion arithmetic applies. Misalign any one and the test produces a number that looks like evidence and is not. Two related pieces cover neighbouring ground properly: what a good reply rate even is deals with building the measurement instrument before you compare anything with it, and subject lines covers the specific mechanics a subject test needs. This post is about whether the test can resolve at all.
The volume arithmetic, and the scaling law that matters more than the table
Comparing two proportions at a two-sided 5% significance level with 80% power, using the normal approximation, gives these requirements per arm:
- Separating a true 3% from a true 4.5% — a 50% relative lift — needs about 2,500 delivered messages per arm, so 5,000 in total.
- Separating a true 3% from a true 3.6% — a 20% relative lift — needs about 13,900 per arm.
- Separating a true 3% from a true 3.3% — a 10% relative lift — needs about 53,200 per arm.
The table is less useful than the rule behind it. Required sample scales with one over the square of the difference. Halving the effect you want to detect roughly quadruples the volume you need. That single relationship decides your whole testing programme, because it means the question "what should I test?" is really the question "what can produce a big enough difference?"
An outbound team sending 3,000 cold emails a month can resolve a 50% relative lift in about six weeks of full volume with nothing else running. It cannot resolve a 20% lift inside a year. So when someone proposes testing two phrasings of a call to action, the correct response is not "good idea, let us try it" — it is to ask what difference that change could plausibly produce, and then to note that a difference that small is not measurable here.
Your list is clustered by company, and that inflates the false-positive rate
This is the defect that survives every other correction, because it is invisible in the send counts.
Sample-size arithmetic assumes independent observations. Two contacts at the same company are not independent. They share a mail gateway with one filtering policy, one domain reputation view of you, one internal Slack channel where "did you get this weird email too?" gets typed. Their outcomes correlate.
Survey sampling has a standard correction for this, the design effect, from Leslie Kish's 1965 work on the subject. In its simplest form it is one plus the average cluster size minus one, multiplied by the intraclass correlation. Concretely:
- Three contacts per company with a correlation of 0.10 gives a design effect of 1.2 — you need 20% more sends than the naive arithmetic said.
- Five contacts per company with a correlation of 0.20 gives a design effect of 1.8 — nearly double.
Two things to take from this. First, if you are multi-threading accounts, multiply the sample size you calculated by the design effect before you start. Second, and more usefully, randomise at the company rather than at the lead. Assign every contact at one company to the same arm. It costs you nothing, it removes the within-company contamination where two colleagues receive different variants and compare them, and it makes the account the honest unit of both randomisation and analysis.
A test at step three is not a test on your list
Steps do not sample the same population. Step one goes to everyone. Step three goes to the people who did not reply, did not bounce and did not unsubscribe at steps one and two. That is a selected group, and the selection is not random — it is precisely the people your first two messages failed to move.
Three consequences worth designing around:
The baseline rate differs by step, so the sample size differs by step. Later steps typically convert at a lower rate against a smaller pool, which means they need proportionally more volume to resolve the same relative effect. A test that was adequately powered at step one is under-powered at step four twice over.
A winner at step three may be a winner only for the residue. The variant that works on people who ignored your first two emails is not necessarily the variant that works on a fresh list. Do not promote a step-three winner to step one without retesting it there.
Cross-step tests are confounded by attrition. If arm A stops more people at step one — because it is better, or because it triggers more unsubscribes — then arm A's step-three population is different from arm B's, and any step-three comparison is comparing two different groups of people rather than two messages.
The clean design is one test at one step at a time, with everything downstream held fixed. It is slower. It is also the only version that answers a question.
Autocloz's free plan covers 5 users and 10 mailboxes with the warmup and DMARC monitoring included, which is enough to run one properly-powered test at a time on your own domains — start free and check your monthly volume against the arithmetic above before you design the programme.
What to test, ranked by whether it can move enough
Order your testing programme by expected effect size, not by how easy the change is to make. Easy changes are usually small changes, and small changes are unmeasurable.
- Who is on the list. Changing the segment changes the reply rate by multiples, not by percentages. A test of "founders at 10–50 person agencies" against "heads of marketing at 200-person retailers" produces a difference large enough to see in hundreds of sends rather than tens of thousands. This is the highest-yield test available and almost nobody runs it as a test.
- What you are offering. An audit versus a demo versus a benchmark report is a change in what the reader is being asked to want. Large effect, and cheap to test because the copy around it barely changes.
- The ask. A twenty-minute call versus "is this worth a look?" versus a reply-with-a-word question changes the friction of responding, which is a first-order driver of reply rate.
- Channel mix. Whether a LinkedIn touch precedes the email changes the conditions under which the email is read. Larger effect than most copy changes, and harder to run cleanly because the two channels have different rate limits.
- Message structure. Length, whether the first line references the recipient, whether there is a link. Moderate effect, resolvable at high volume.
- Wording. Subject phrasing, opening line variants, sign-off. Smallest effect, largest sample requirement. Test these last and only if you have the volume, and be honest that most published wins at this level are underpowered results that will not replicate.
The inversion is the point: the tests that are easiest to set up are the ones least likely to resolve. Two related pieces are worth reading alongside this — personalisation at scale covers making the first four levers operational, and the subject line tester is a reasonable pre-flight check for the sixth even though it is not a substitute for a test.
A worked quarter, with the volume budget written down
Illustrative arithmetic, not a measured case study. Substitute your own numbers.
A team sends 4,000 cold emails a month, so 12,000 in a quarter. Baseline reply rate 3%. What programme fits?
Option A — one wording test. Testing a subject variant that might produce a 20% relative lift needs about 13,900 per arm, so 27,800 total. The quarter supplies 12,000. The test cannot run. Not "will run slowly" — cannot run.
Option B — one segment test. Splitting the quarter across two segments, 6,000 each, and expecting a difference of the order of 3% against 6% needs roughly 750 per arm to resolve. The quarter supplies eight times that. The test resolves inside three weeks, with volume left for a second test after it.
Option C — one offer test at fixed segment. 3% against 4.5% needs about 2,500 per arm, so 5,000 total. Fits in six weeks, leaving six weeks for the follow-on test on the winner.
The plan that follows from the arithmetic: weeks one to three, segment test. Weeks four to nine, offer test on the winning segment. Weeks ten to twelve, hold everything and re-measure the baseline to confirm the gains persisted. Zero wording tests, because at 12,000 sends a quarter they are not measurable and running them would consume the volume the resolvable tests need.
Write the budget down before the quarter starts. The failure mode this prevents is running four tests at once, none of them powered, and ending the quarter with four numbers and no knowledge.
Diagnosing a result that is telling you nothing
Symptoms, and what each one means.
- The winner flips when you re-run it. Almost always an underpowered test. Check the per-arm send count against the arithmetic above; if it is short by an order of magnitude, the first result was noise.
- Arm counts are badly unequal. Either the assignment is not uniform or one arm was added mid-flight. If a variant was added after sends had already gone out, the earlier sends were assigned under a different arm count and cannot be attributed to the current list.
- The open-rate difference is large and the reply-rate difference is zero. Suspect machine opens before you conclude the subject line worked. A variant that trips a security gateway differently gets a different pixel-load profile from the same humans.
- The winner has a higher bounce rate. You are looking at a list-quality difference, not a copy difference. Check whether the arms drew from the same source.
- The result appeared after you started checking daily. Repeated looks inflate the false-positive rate well above the nominal level, because each look is another opportunity to cross the line by chance. Fix the sample size before starting; do not fix it afterwards.
- The effect is real in the test and gone in production. Check whether the test ran on one segment and production runs on three. A variant tuned to the segment you tested on is not a variant that generalises.
What Autocloz's A/B testing does not do
Specific limits, because a testing tool you misunderstand is worse than none.
There is no significance test. The per-variant results endpoint reports sends, opens, clicks, replies and bounces per arm, computes each as a rate, and names the arm with the highest rate as the winner. Eligibility to win requires at least one send. Not a hundred, not a thousand — one. The winner label is a description of which arm currently has the highest observed rate, and treating it as a statistical claim is the single most likely way to misuse it. Read the send counts next to it, every time.
The winner metric differs by dimension. For a subject test the winner is picked on open rate; for a body test, on reply rate. Given that open rate is now heavily contaminated by machine fetches, read the reply column yourself rather than accepting the subject-test winner as named. The product does filter machine opens out of these counts using the same rule the campaign stats use, which removes blank-user-agent prefetches and gateway scanners, but filtering is a reduction in contamination rather than its elimination.
Body variants are configurable and have never dispatched. The body_variants field exists on every sequence step, and the picker that would use it has no production call sites anywhere in the codebase. The analytics side is deliberately left on its original attribution shape for that dimension precisely because there is no dispatcher behaviour to match — changing it would invent numbers for sends that never happened. Subject-line A/B on email is the dimension that actually runs.
Editing the variant list after sends have gone out breaks attribution. Arm assignment is recomputed at read time from the current list, so adding or removing a variant moves the index basis under sends that already happened. Freeze the variant list for the duration of a test.
Historic sends from before 27 July 2026 cannot be re-derived. Two defects were fixed that day — the primary subject was not in the rotation at all, so a one-variant test ran 100% variant and no control, and assignment was random per send rather than sticky per recipient. Sends made under the old behaviour did not record their arm, so they are bucketed by the current rule and their numbers are approximate. Any test that was already running across that date should be restarted rather than read.
Per-step analytics are capped. Attribution reads the most recent sends per step up to a limit of 25,000, and the response flags when it has been capped. Above that, the numbers describe a recent window rather than the campaign's whole history.
None of that makes the feature useless. It makes it a measurement surface rather than a decision engine, which is the right way to hold any A/B tool — the arithmetic stays your job. For the reporting surface itself, campaign analytics is where the per-variant numbers live, and if you are weighing tools primarily on their experimentation model, the comparison against Instantly sets out where the approaches differ.
Frequently asked
How much volume does a cold email A/B test actually need?
More than most outbound teams have. Comparing two proportions at a 5% two-sided significance level with 80% power, separating a true 3% reply rate from a true 4.5% one needs roughly 2,500 delivered messages per arm. Separating 3% from 3.6% needs closer to 14,000 per arm. Required volume scales with the inverse square of the difference you are trying to detect, so halving the effect you care about roughly quadruples the sends you need.
Should I measure an A/B test on open rate or reply rate?
Reply rate, for anything you intend to act on. Open rate is inferred from a tracking pixel, and pixel loads are now routinely triggered by mail clients and security gateways rather than by people, so an open-rate difference can be a difference in how two variants were machine-processed. Reply is a human action with no machine analogue, which is why it survives as a measurement even though it is a smaller number.
What should I test first in cold email?
Test the thing with the biggest possible effect, because effect size determines whether a test can resolve at all. Ordered roughly by the size of change they can produce: who is on the list, what the offer is, which channel mix reaches them, how the message is structured, and only last which words are used. A wording test is the hardest to resolve and the least likely to matter.
Can I run more than one A/B test at the same time?
You can run tests on different steps at once, but understand that a test at step three measures only the people who did not reply or bounce at steps one and two, which is a selected population rather than your list. If you must run several tests, keep each one on a different step, fix the sample size before you start, and treat a result that appears in only one of several simultaneous tests with the scepticism it deserves.
Does Autocloz tell me whether an A/B result is statistically significant?
No. Its per-variant results report sends, opens, clicks, replies and bounces per arm with machine opens filtered out, and it names a winner as the arm with the highest rate among those with at least one send. There is no significance test, no confidence interval and no minimum sample before a winner is named, so the winner label is descriptive rather than a statistical claim and should be read alongside the send counts.
Why do two contacts at the same company make a test less reliable?
Because they are not independent observations. People at one company share a mail gateway, a filtering policy and often an opinion, so their outcomes correlate. Standard sample-size arithmetic assumes independence, and correlated observations inflate the false-positive rate. The correction is the design effect from survey sampling — one plus the number of contacts per company minus one, times the correlation between them — applied as a multiplier on the sample size you calculated.