A/B test sample size is the part of testing nobody wants to talk about, because the math says most small stores cannot run most of the tests they are running. There, we said it. A test’s job is to detect a difference, and detection has a price paid in visitors. Small lift, small traffic: the test cannot pay, and it will hand you a coin flip dressed up as a result.
The intuition is simple. Conversion data is noisy; any two identical pages will show different numbers over a week just by chance. To claim your variant “won,” its lift has to stand out above that noise, and the noise only shrinks with volume. The smaller the true lift, the more visitors you need before you can tell it apart from luck, and the relationship is brutally nonlinear: halving the lift you want to detect roughly quadruples the traffic you need. A subtle button tweak that might move conversion a few percent relative is, on a store doing a hundred orders a month, statistically invisible. You could run it for a quarter and learn nothing except that winter happened.
Run the math before the test
Before you build anything, spend two minutes in a sample size calculator: enter your baseline conversion rate and the smallest lift you would care about, and it tells you the visitors needed per variant. Divide by your weekly traffic to the tested page. If the answer is “nineteen weeks,” you have your verdict: that test is not runnable on your store. That is not a failure, it is a free answer that just saved you a month.
What small stores should do instead
- Test bigger swings. Big changes produce detectable effects on small traffic. New offer structure, different hero, restructured product page: these can clear the noise floor. Button colors cannot.
- Test upstream, where the events are. This is why we told you in tip 33 to watch add-to-cart rate: more events per visitor means the same traffic buys you more statistical power than purchase-based tests get.
- Test on your busiest pages. The homepage and your top collection can support tests your long-tail product pages never will.
- Below the traffic floor, stop pretending. Make the change, watch your baselined metrics for two full weeks, and be honest that this is a before/after judgment, not an experiment. A humble before/after beats a fake A/B test, because at least you know how much to trust it.
The discipline from tip 7 still stands: one variable, a metric chosen in advance, no peeking. Today’s rule sits in front of it: first check whether the test can be won at all. A 5 percent lift you cannot detect is a test you should refuse to run.