A/B testing has become the reflex move of CRO, so much so that we forget it was designed by and for very high-volume sites. The reality of most of the SMBs I work with: 5,000 to 30,000 visits per month, a few dozen to a few hundred conversions. At that scale, the classic A/B test, the one the e-commerce giants run, does not work as-is: it would take months to reach a reliable result, and the site will have changed before then. Should you give up on experimentation? No, but you need a different method: calculate before launching, prioritize ruthlessly, use statistics built for the job, and above all learn to recognize the decisions that do not need a test at all.
On a low-traffic site, the scarce resource is not the test idea. It is the testing window. You only get eight to ten of them per year on any given page. Each one must be spent like a budget.
The statistical wall: why your test will detect nothing
Let's start with the calculation almost nobody runs before launching. An A/B test detects an effect only if it has enough observations to separate signal from noise. Take a page converting at 2% and suppose a variant improves it by 10% in relative terms (2% → 2.2%), already a fine win. Evan Miller's sample size calculator delivers the verdict: you need roughly 78,000 visitors per variant, over 150,000 in total, to detect that effect at the usual standards (95% confidence, 80% power). With 10,000 monthly visits on the page, the test would run for more than a year.
That brutal calculation structures everything that follows. Only two levers shorten a test: increasing the effect you are hunting for (a +30% effect is detected about ten times faster than a +10% one) and increasing the frequency of the measured event (a click toward the cart is ten times more frequent than a purchase). The whole low-traffic CRO method flows from these two levers: test big, and measure higher up the funnel. The third reflex, quietly lowering the statistical bar, is not a lever; it is a way of telling yourself stories.
Prioritizing with ICE: test rarely, test big
Since testing windows are scarce, prioritization becomes the most profitable move in CRO. The grid I use is ICE scoring, popularized by Sean Ellis: each hypothesis is scored out of 10 on Impact (how much conversion would move if the hypothesis is true), Confidence (what evidence supports it: GA4 data, heatmaps, customer verbatims, support tickets) and Ease (implementation cost). The product of the three scores ranks the list; you test from the top.
On a low-traffic site, I add a twist to the grid: Impact is eliminatory. A hypothesis with high Ease but low Impact (changing a button color, rewording a label) can score a decent total, but it produces a +2% or +3% effect your traffic will never detect. It does not deserve a testing window: either ship it directly (if it is common sense), or drop it. The windows are reserved for radical changes: a new value proposition above the fold, a restructured quote form, a redesigned pricing page, a fully reordered funnel. It is counter-intuitive, but it is mathematical: the less traffic you have, the bolder your tests must be: a small site cannot statistically afford cautious tweaking.
So where does Confidence come from? From upstream research: that is exactly the role of heatmaps and session recordings, which show where visitors hesitate, click into thin air or abandon, and from clean journey data, which requires a reliable analytics foundation. A hypothesis backed by three converging sources (heatmap + verbatim + GA4 funnel) earns its window; a workshop hunch, rarely.
Sequential testing: decide earlier without cheating
The classic statistical protocol (known as "fixed horizon") has a hard rule: you compute the sample size in advance, and you only look at the result at the end. Yet everyone peeks along the way, and stops the test the first time it turns "significant". That peeking wrecks the test's validity: by checking every day, you multiply the opportunities for a false positive, and the real error rate can exceed 25% instead of the advertised 5%, as explained in "How Not To Run an A/B Test", the reference piece on the subject.
The right answer is not to forbid yourself from looking, but to use sequential statistics, designed to be monitored continuously: they adjust the decision threshold at every observation, allowing you to stop a test early when the effect is large, without inflating the error risk. It is the approach modern platforms have adopted (Optimizely's Stats Engine documented it publicly), and it is especially valuable on low traffic: combined with bold tests, it returns its verdict within weeks when the effect is real and massive, instead of imposing a fixed horizon sized for a small effect.
And the "before/after" test, shipping the change and comparing conversion rates month over month? I use it, but with open eyes: it is a measurement, not an experiment. Without simultaneous random assignment, seasonality, a SEA campaign kicking off or an article starting to rank is enough to explain the difference. Before/after is reserved for changes whose expected effect is far larger than the metric's usual noise, and it is read on isolated channels (compare direct traffic's conversion rate before and after, not the blended rate).
Step down a level: measure where the volume lives
Second lever: the metric. If the final sale is too rare to power a test, optimize on the micro-conversion that precedes it and occurs 5 to 20 times more often: click toward the form, add to cart, checkout start, appointment booking. The logic is the same one I apply in ticketing when the Purchase pixel is unavailable: you steer on a proxy, provided you have verified that its exchange rate toward the final conversion is stable. And a test that lifts form starts by 25% without moving submissions reveals something precious in itself: the problem is not the funnel entrance. It is the form.
Third lever: aggregation. Testing a template rather than a page (the 40 product pages sharing one layout, a firm's 12 service pages) pools the traffic and lifts a page that is individually invisible above the statistical threshold. It is the only way to test properly on content sites and mid-sized catalogs.
When not to test
The A/B test is a risk-reduction tool, not a validation ritual. Three situations where the right decision is not to test:
- Obvious flaws get fixed, not tested. A form bug on mobile, a page loading in six seconds, an unfindable price, a phone field that is mandatory for no reason: user research and session recordings are enough to establish the problem. Testing the fix of an objective defect means spending a testing window to confirm that repaired beats broken.
- Cheaply reversible, low-stakes decisions get shipped. Changing a button label, reordering an FAQ: if rollback costs nothing and the stakes are low, ship it, note the date, watch the metric. The marginal learning of a test is not worth its opportunity cost.
- Below the threshold, research beats experimentation. When the power calculation announces more than 6 to 8 weeks of testing even on a micro-conversion, the best use of your time is not waiting: it is five user interviews, one usability testing session, a re-read of customer verbatims, which will produce more sound decisions than an underpowered test.
| Conversions/month on the page | Recommended approach |
|---|---|
| < 100 | No A/B testing: user research, fix the obvious flaws, before/after on major changes |
| 100 to 500 | Radical tests only, on micro-conversions, sequential statistics, one test at a time |
| 500 to 1,000 | Bold tests on the main conversion, aggregated templates, strict ICE prioritization |
| > 1,000 | Classic testing program: finer iterations become possible, several test areas in parallel |
The most expensive mistake
Stopping the test "because it turned green". A fixed-horizon test checked daily will almost always pass through a zone of illusory significance, especially in the first days, when samples are tiny and variance is huge. The team celebrates a +18%, ships the variant, and three months later the conversion rate has not moved: the gain never existed. If your tool does not use sequential statistics, the rule is simple: set the duration in advance, respect it, and always include full weeks to neutralize weekly cycles.
In the crucible: distilling certainty from scarcity
Low-traffic CRO is not discount CRO. It is more demanding CRO. Where the big site can afford to test everything and let volume settle the matter, the modest site must distill: a few high-impact hypotheses, grounded in real research, measured where the volume lives, with statistics that do not lie. The lead is the queue of underpowered micro-tests that will never teach you anything; the gold is three decisions a year: large ones, settled cleanly. At this scale, every test is an ingot: you do not pour it into a dubious mold.
If your site converts but you do not know where to start (what to test, what to ship without testing, what to instrument first), that is exactly the work I do within my CRO expertise: funnel audit, ICE prioritization of hypotheses, a testing protocol sized to your volume, alongside the analytics foundation so every decision rests on a reliable measurement.
A funnel to optimize, modest traffic?
I audit your conversion journey, calculate what your traffic can genuinely test, and prioritize the actions: those that deserve a test, and those you should simply ship.