Think Build Implement Repeat
Web Development

A/B Testing With Small Numbers, Honestly

Last updated:

The uncomfortable arithmetic

To detect a modest improvement with any confidence you need a substantial number of conversions per variant, not visits. A site with thirty enquiries a month cannot run a meaningful test on enquiry rate — it would take many months, during which everything else changes.

Running the test anyway produces a result, and the result is noise. Acting on it is worse than not testing, because it feels evidence-based.

What to do instead

  1. Make changes that are obviously better. Clearer headline, visible pricing, fewer form fields, faster pages. These do not need testing; they need doing.
  2. Watch session recordings. Five recordings of real people attempting your conversion path teach more than a underpowered test.
  3. Ask customers. Five conversations with recent buyers about what nearly stopped them is cheap and specific.
  4. Measure the trend over quarters rather than testing variants over weeks.
Most small sites have several obvious improvements outstanding. Testing is what you do when you have run out of those, and very few businesses have.

Test higher up the funnel

You may not have enough enquiries to test on, but you may have enough clicks, scroll depth or form starts. Testing on a more frequent event gets you a usable signal sooner, provided the event genuinely predicts the outcome.

Be careful: optimising form starts can reduce completions if the change attracts less serious visitors. Check the downstream number even when testing an upstream one.

When testing becomes worthwhile

  • Hundreds of conversions per month, not per year
  • A stable site where other changes are not confounding the result
  • A specific hypothesis rather than a general urge to improve
  • Discipline to run the test to completion rather than stopping when it looks good

Stopping early is the classic error

Watching a test and stopping when one variant is ahead produces a false positive most of the time. Early differences are noise, and noise is at its largest when the sample is smallest.

Decide the duration and sample size before starting, and do not look at the result as a decision until then.

Frequently asked questions

How many conversions do we need?

Enough that a modest difference would be distinguishable from chance — typically hundreds per variant for a small effect. Use a sample size calculator with your own baseline before committing.

Can we test on a smaller effect size?

Detecting a smaller effect needs a larger sample, not a smaller one. If you can only detect large effects, test only changes you expect to be large.

What about testing with AI-generated variants?

Generating variants is cheap and easy; the constraint remains traffic. More variants make the sample problem worse, not better.

Is personalisation an alternative?

It has the same statistical problem in a more complicated form, plus maintenance. For most small sites, one good page beats several personalised ones.

Keep reading

Told you should be A/B testing?

Tell us your monthly conversion count. We will tell you honestly whether testing can conclude anything for you yet.

Book a free 30-minute call Get a project estimate WhatsApp us

Related services

What we build for problems like this one

Web DevelopmentCustom Software Development