Performance

Is your email list big enough to A/B test?

TL;DR

The 1,000-per-variant rule is an open-rate number, and Apple Mail Privacy Protection has made open rate too noisy to test on. The metric that survives is clicks, but click rates sit near 2.5%, so detecting a realistic 20% lift needs roughly 16,000 subscribers per variant, not 1,000. If your list is under about 20,000 and you test clicks correctly, one send cannot reach statistical significance. Below is the per-metric sample-size table and what to do instead.

Is your email list big enough to A/B test? The answer hinges on one thing nobody puts in the calculator: which metric you plan to measure. Every ESP help doc and sample-size tool repeats the same figure, 1,000 subscribers per variant. That number is not wrong. It just describes open rate, and open rate is the one metric you can no longer trust. Test clicks instead, and the real threshold jumps to something closer to 17,000 subscribers per variant. For most publishers, that quietly rules out the single-send A/B test they were planning.

The short answer, by the metric you test

Sample size is not one number. It scales with your base rate, so the rarer the event you are measuring, the more people you need to see a difference that is not just noise. Opens happen 30 to 40% of the time. Clicks happen 2 to 3% of the time. That gap is the whole story, and it is the part the rule of thumb skips.

Here is what it takes to detect a 20% relative lift at 95% confidence and 80% power, using the standard two-proportion sample-size calculation. A 20% relative lift means moving a 35% open rate to 42%, or a 2.5% click rate to 3.0%. That is an ambitious but realistic target for a single subject line or layout change.

Metric you test Typical base rate 20% lift target Per variant Full test (both arms)
Open rate35%42%~750~1,500
Click-to-open rate12%14.4%~3,100~6,200
Click-through rate2.5%3.0%~16,800~33,600

Read the last column, not the first. The familiar 1,000-per-variant advice lines up with the open-rate row, which is exactly the metric Apple broke in 2021. The moment you switch to the metric that still means something, you need ten to twenty times the audience. Nobody puts that on the calculator's landing page, because it makes the calculator far less fun to use.

Why the 1,000-per-variant rule quietly fails

Two things went wrong with the rule of thumb. The first is that it was always a base-rate-specific number that got repeated as if it were universal. A 1,000-subscriber test can comfortably separate a 35% open rate from a 42% one, because opens are frequent enough that random variation settles down fast. The same 1,000 subscribers cannot tell a 2.5% click rate from a 3.0% one. At that scale you might see 25 clicks in one arm and 30 in the other, and five clicks is well inside the range of pure luck.

The second problem is worse. Open rate stopped being a metric you can test on at all. Apple Mail Privacy Protection pre-fetches the tracking pixel in every message sent to Mail users, whether or not a human ever looks at it. In our own corpus we see reported opens running 40 to 44% above the real figure once you strip out proxy-triggered opens, which we broke down in our analysis of Apple MPP and open rates. When close to half your opens are machine-generated, an A/B test on open rate is measuring which subject line Apple's servers happened to fetch first. Which is nothing.

What you should test instead, and the list size it demands

Test clicks. Click-through rate and click-to-open rate are the two signals Apple cannot fake, because a proxy server does not click links. Across the newsletters we analyze, a healthy click-to-open rate lands between 10% and 20%, with 6% as the floor where something is clearly wrong, which is the pattern we mapped in the click-to-open rate benchmarks. Raw click-through rate against the full list is lower, usually 2 to 3%. Those are the numbers you can trust, and they are also the numbers that need a big audience to test.

Run the calculation for a 2.5% click rate and a realistic 20% relative lift and you land between 14,000 and 21,000 subscribers per variant, depending on your exact base rate. Split a list in half to run the test and you need roughly 30,000 subscribers total before a single send can return a statistically significant result. That is the sentence most sample-size posts avoid printing. If your list is under about 20,000 and you are testing clicks honestly, one email cannot get you to 95% confidence. It is not a matter of waiting a bit longer or squinting at the dashboard until the winner looks convincing. The math is not there. This is the number I most often have to talk publishers out of ignoring, because their ESP will happily declare a winner anyway.

If your list is small, accumulate the test across sends

There is one honest way to run a valid click test on a small list: stop treating a test as a single send. If you send the same two variants across multiple issues to randomized halves of your list and pool the results, the sample accumulates. A 2,000-person list that sends weekly reaches a 16,000-per-variant sample in about eight sends, assuming your list size and click rate hold steady over those two months.

The catch is that this only works when the thing you are testing stays constant and everything around it does too. A structural change like a single-column layout against a two-column one can be held fixed across eight issues. Subject-line wording cannot, because the specific words change every week, so there is nothing stable to pool. Accumulation buys you significance on durable design decisions, not on the per-issue copy choices most people actually want to test. Name that tradeoff before you start, or you will collect eight weeks of data that answers no single question.

Too small to split-test? Score the variants first

When your list cannot reach significance in one send, the next best move is to pick the stronger variant before it goes out. The Newsletrix subject line tester runs your options head to head and scores them on the factors that correlate with clicks, so you ship the better line instead of guessing.

Test your subject lines →

What to do when the math is against you

Most publishers sit under 20,000 subscribers, so most publishers cannot cleanly A/B test clicks in a single send. That is fine. It just means the goal shifts from proving a small winner to making good decisions without a p-value to hide behind.

The first move is to pre-score variants instead of split-testing them. A model trained on which subject lines and layouts drive clicks gives you a directional read in seconds, with no send required. It is a proxy, not a randomized trial, and it will occasionally be wrong on a line that breaks the pattern in a good way. But a decent proxy beats an underpowered test that reports a false winner with fake confidence, which is the real alternative on a small list. If you have followed our walkthrough on A/B testing subject lines, this is the shortcut for the sends where you do not have the audience to do it properly.

The second move is to test bigger swings. Sample size shrinks fast as the effect you are chasing grows. A 20% relative lift needs about 16,000 subscribers per variant at a 2.5% click rate; a 50% lift needs closer to 3,000. So stop testing one word in the subject line and start testing whole framings: a completely different lead story, or a plain-text issue against a designed one. Big, structural differences produce big effects, and big effects are the only ones a small list can detect. The tradeoff is real. You learn that plain text beat HTML, not which of six layout tweaks carried the win. On a small list, that is the trade you have to make.

You can also watch what larger competitors test, since they have the volume to find effects you cannot. Tracking a rival's subject-line and layout patterns over time, the kind of teardown that tools like the alternatives to Litmus support, lets you borrow conclusions from senders with 200,000-subscriber lists. It is not the same as testing on your own audience, but it points you at the structural bets worth copying before you burn a send on them.

How we size a test before running one

Before we run any test on a newsletter, we do the boring step first: work out whether the list can support it. Take the base rate for the metric that matters, decide the smallest lift worth acting on, and run the two-proportion calculation. If the required per-variant sample is larger than half the list, the test is dead on arrival and we do not send it. That five-minute check saves weeks of collecting data that will never separate signal from noise. Set the threshold against your own trend rather than a borrowed benchmark, which is a habit worth building from the open-rate benchmark data too.

The number people forget is time. Significance is about the count of clicks, not the count of days. A 12,000-subscriber list testing clicks will not cross a 20%-lift threshold in a week, a month, or a quarter on single sends, because the per-send click volume is too low to ever get there. The right question is never how long to wait. It is whether the list produces enough clicks per send that pooling a handful of issues gets you there, and if not, whether pre-scoring or a bigger swing is the smarter call. Answer that up front and you stop running tests that were always going to end inconclusive.

Frequently asked questions

How many subscribers do I need to A/B test a newsletter?

The number changes with the metric you test. For open rate at 95% confidence and 80% power, detecting a 20% relative lift takes roughly 750 subscribers per variant. For click-through rate, the metric that survives Apple Mail Privacy Protection, you need closer to 15,000 to 21,000 per variant because clicks are far rarer than opens. Because you split the list to run the test, a valid click test needs a total list near 30,000 subscribers.

Can I A/B test with under 1,000 subscribers?

For open rate, no, and it would not tell you much anyway since Apple Mail Privacy Protection inflates opens by 40 to 44%. For clicks, a single send is far too small. The workable option on a list that size is to run the same two variants across many issues and pool the results, which only works if you are testing a stable design choice rather than per-issue copy. For one-off decisions, pre-score the variants instead of split-testing them.

Why is 1,000 per variant not enough for click tests?

Sample size scales with how rare your outcome is. Opens happen 30 to 40% of the time, so 1,000 subscribers produce enough opens for random noise to settle. Clicks happen around 2 to 3% of the time, so 1,000 subscribers might produce 25 clicks in one arm and 30 in the other, and that five-click gap is well within chance. To separate a real 20% difference in a 2.5% click rate you need roughly 16,000 subscribers per variant, not 1,000.

Does Apple Mail Privacy Protection affect A/B testing?

Yes, and it breaks open-rate testing specifically. Apple Mail Privacy Protection pre-fetches the tracking pixel for Mail users before any human opens the message, so reported opens run well above real opens. In our corpus that inflation runs 40 to 44%. An A/B test on open rate now largely measures Apple's proxy behavior rather than reader interest, which is why clicks are the only reliable metric to test.

How long until an email A/B test reaches significance?

Significance depends on the number of opens or clicks collected, not the number of days elapsed. A large list can reach a valid click test in one send; a small list may never reach it on single sends because the per-send click volume is too low. If your list is under about 20,000 and you are testing clicks, plan to pool the same test across multiple issues or switch to pre-scoring, because waiting longer on one send will not fix a sample-size problem.

Related reading

Get started

Stop guessing. Start winning.

Join newsletter creators using AI-powered competitor intelligence to ship better content, faster.

No credit card required  ·  Cancel anytime  ·  All features on every plan