A/B testing on low traffic: what to do instead
On low traffic the danger is not an inconclusive test, it is a flattering one. What underpowered tests do to your numbers, and what to do instead.
Small sites do not fail at A/B testing by getting no result. They fail by getting a result that is far too flattering, then building a quarter on it.
Can you A/B test on low traffic?
You can run the test. You cannot trust the number it returns. Below roughly 850 conversions per test, most experiments never resolve, and the few that do overstate the win. On a small site the real danger is not a blank result. It is a flattering one.
Here is the shape of it. Take 8,000 visitors a month converting at 2%, split evenly across a four week test. That is 4,000 visitors and 80 sales per variant. The conversions one test needs to settle even a 20% swing is around 850. You are not close, and no amount of patience with a four week window closes that gap.
Most advice stops there and says run it longer. That treats a weak test as a slow test. It is neither. It is a different test, one that lies in a specific direction.
What happens to a test that cannot reach power?
Power is the chance of spotting a real effect. At 4,000 visitors per variant on a 2% conversion rate, a genuine 10% improvement has about a 9.6% chance of being detected. Nine times in ten, a change that truly works finishes the test looking like nothing.
That part founders half expect. They call it inconclusive and move on. The expensive part is the one time in ten it comes back green.
Why does a small win read bigger than it is?
Because only an unusually large swing can clear the significance bar on thin data. The ordinary size of a real win never gets through. So the wins you are allowed to see are drawn from the exaggerated tail, and the measured lift you read off the dashboard is inflated.
- +43% the lift you would read at 4,000 visitors per variant
- +26% the lift you would read at 10,000 per variant
- +11% the lift you would read at 80,000 per variant
The true lift in all three of those is 10%. Nothing about the change is different. The only thing that varies is how much traffic was behind the measurement, and the thinner it gets the more the surviving winners exaggerate. Those figures come from simulating 600,000 tests per row at a 2% baseline and a real 10% improvement, taking the median measured lift among the tests that reached 95% significance.
A small site does not get quiet wrong answers. It gets loud ones.
This is not a quirk of web analytics. John Ioannidis stated it as a general law of research in Why Most Published Research Findings Are False: the smaller the studies in a field, the less likely its findings are to be true, and the smaller the true effects, the worse it gets. Both corollaries describe your website precisely.
Want the ranked list before you test anything? Revslip finds it on your URL.
What should you do instead of testing?
Stop testing small changes, because you cannot afford to measure them, and start shipping changes big enough to be visible. A bundle of five real fixes can move conversion 30% or more, and a 30% swing needs around 390 conversions to confirm rather than 3,200.
- Fix what is provably broken. A form field that fails on mobile does not need an experiment. It needs fixing.
- Bundle, do not drip. Ship five changes at once. You lose the ability to attribute, and you buy an effect large enough to see.
- Pick the biggest leak first. Order by what each problem costs a month, not by what is quickest to build.
- Set the stopping rule before you start. Decide the date and the sample in advance, then do not look early.
That last one matters more than it sounds. Evan Miller showed that if you stop a test the moment it goes green, on a change that does nothing at all, you will call it a winner 26.1% of the time rather than 5%. Watching a weak test daily is the fastest way to manufacture a result.
How do you decide without a p-value?
You decide on mechanism. If you can name why a change should work, point at the specific friction it removes, and show the drop-off it targets, that reasoning carries more weight on a small site than a p-value built on 160 sales. Significance was never the only evidence available.
You can build that list yourself in an afternoon, and the method is written up in how to audit my own website.
Revslip checks roughly 200 conversion signals and ranks each by what it costs a month, which is a mechanism-first approach rather than a measurement-first one.
How do you measure a change you did not test?
Compare matched periods and say out loud what you cannot rule out. Four weeks before against four weeks after, same days of the week, same campaigns, seasonality noted. Then describe the result as a direction with a caveat rather than a proven number, and never forecast off it.
The honest phrasing is "conversion went from 1.4% to 1.9% over four matched weeks, and we changed four things, and traffic mix was steady". That sentence is weaker than a p-value and more useful than a fabricated one. The longer method, including how many weeks you have to wait before the number means anything, is in how to track whether a fix worked.
Where is this advice weakest?
In two places, and the first one is ours. Bundling trades away attribution: when five changes ship together and conversion moves, you do not learn which one did it, so you cannot build a playbook from it. That is a genuine loss, not a rounding error.
Our own conflict, stated plainly Revslip sells diagnosis, so "diagnose instead of test" is the conclusion we would reach anyway. The 134 businesses we have audited also arrived already convinced something was wrong, which is not a neutral sample of the web.
The second weakness is that sequential and Bayesian methods genuinely do better than the fixed-horizon maths above. They reach decisions on less data by planning for repeated looks instead of punishing them. They shrink the requirement. They do not abolish it, and a site at 160 conversions a month stays out of reach on any method.
Common questions
Three things founders ask once the threshold sinks in: how long is too long to leave a test running, whether switching to a different statistical engine rescues the situation, and whether testing is worth doing at all at their size. Short answers first, reasoning after.
How long should I run a test on a small site?
Long enough to hit the sample you decided on in advance, or not at all. A test stretched to three months collects a season change, a pricing update and a traffic shift, which is no longer one experiment.
Do Bayesian tools fix this?
They change what the output means rather than creating information. A Bayesian result on 160 sales is dominated by the prior you chose. Useful, honest, and not the same thing as evidence about your page.
Is testing worth doing at all below the threshold?
Yes, for changes you expect to be large, and as a safety net to catch a new page that is much worse. Underpowered tests are poor at confirming small wins and reasonable at catching disasters.
If you would rather start from the ranked list than the experiment queue, run a free audit.
Power and effect-inflation figures computed from 600,000 simulated tests per traffic level at a 2% baseline, two-sided at 95%, reproducible with any binomial simulation. Repeated-significance figure from Evan Miller, How Not To Run an A/B Test. Small-study corollaries from Ioannidis, PLOS Medicine, 2005. Neither source sells conversion services or testing software.