The short answer
With low conversion volumes, differences between variants are mostly noise. Running until you have enough conversions to distinguish signal from chance takes longer than most people allow, and stopping early is why so many tests produce contradictory results.
Decide the stopping rule before you start, and hold to it.
Why early decisions mislead
- Early results swing wildly on small numbers
- The first days include a learning period
- Day of week affects who sees the ads
- Stopping when you like the answer biases every test you run
- Multiple simultaneous changes make attribution impossible
The fourth is the one that quietly degrades an account over time. Stopping whenever a variant is ahead guarantees a series of decisions based on noise.
What to change
| Change | Effect size | Data needed |
|---|---|---|
| Landing page | Often large | Moderate |
| Offer or message | Often large | Moderate |
| Creative | Large on social | Moderate |
| Headline wording | Small to moderate | More |
| Bid adjustments | Usually small | A lot |
Test the top rows first. Small changes need far more data to detect and are rarely where the improvement is.
Set the rule in advance
- State what you are testing and what would count as a win.
- Decide the minimum conversions before you look.
- Decide the maximum duration, so a test does not run forever.
- Run it without changing anything else.
- Decide against the criterion, including when the answer is no difference.
Point five includes the outcome people find hardest: most tests show no meaningful difference, and recording that is more useful than finding a winner that was not there.
Accept what low volume means
For a business with a handful of enquiries a week, many tests cannot be run meaningfully. That is a real constraint, not a failure of method.
In that situation, make changes based on reasoning and known good practice, monitor the overall trend, and avoid claiming a test proved something it could not.