Think Build Implement Repeat
London, UK +44 7367 067226
WhatsApp FOLLOW f in X
  1. Home
  2. Blog
  3. Multi-Armed Bandits Instead of A/B Testing
AI & Machine Learning

Multi-Armed Bandits Instead of A/B Testing

Bandits shift traffic towards what is working while the test runs. Where that is genuinely better, and where a plain A/B test remains the right answer.

Updated 2 min readBy SpiderHunts Technologies

Free estimateNo obligation

Get a free estimate

Tell us what you need. A senior engineer reads every enquiry.

Takes under a minute. We never share your details.

  • Free consultation
  • No commitment
  • NDA on request

Prefer to talk? Book a free 30-minute call →

Quick answer — TL;DR

Bandits reduce the cost of testing by moving traffic to better options during the experiment. They suit many short-lived options where you want performance rather than a clean measurement. For a decision you need to defend or understand, a straight A/B test is still better.

The cost of a fixed split

A standard A/B test sends half the traffic to each option for the whole test. If B is clearly worse, half your traffic gets the worse experience until the test ends - a known and accepted cost.

A bandit adjusts the split as evidence accumulates, sending more traffic to whatever is performing better while still exploring. Total performance during the test is higher.

The trade you are making

A/B testBandit
Traffic splitFixedAdapts to results
Performance during testLowerHigher
Clean effect estimateYesHarder - traffic is confounded with time
Handles many optionsPoorlyWell
Easy to explainYesLess so

The third row is the real cost. Because allocation changes over time, a bandit's results are entangled with when each option was shown. If something else changed mid-test - a campaign, a season, a news event - unpicking it is considerably harder.

Where bandits fit well

  • Many options - a dozen subject lines or creatives, where a fixed split gives each too little traffic
  • Short-lived content - a promotion running two weeks, where the decision does not outlive the test
  • Continuous optimisation - ranking, layout, offers, where there is no final answer to reach
  • Expensive exploration - where showing a bad option costs real money per impression

Where a plain A/B test is still right

Where the decision matters, is permanent, and will be questioned, you want a clean estimate of the effect and a defensible method.

  • A pricing change you will live with for a year
  • Anything with regulatory or contractual implications
  • A result that needs to persuade sceptical stakeholders
  • Where you need to understand why, not only which
  • Where the effect is small and needs careful measurement

In those cases the traffic cost of a fixed split is a reasonable price for a result you can rely on and explain.

Practical cautions

Bandits converge on the option that looks best early, and early results are noisy. A variant that got lucky in the first hours can capture traffic before the truth emerges, so approaches that keep exploring matter.

They also assume the environment is stable. If the best option changes - a seasonal shift, a different audience arriving - a converged bandit keeps serving yesterday's winner. Continuing to explore a small share of traffic guards against that and is worth the modest cost.

A bandit optimises performance. An A/B test produces an answer. Decide which you need first.

FAQ

Frequently asked questions

The questions readers ask us after this guide.

Still have a question?

Ask us directly — a senior engineer will get back to you.

Ask about your project

Do bandits need less traffic than A/B tests?

Not necessarily to reach the same statistical confidence. They produce better performance during the test, which is a different benefit.

Can I get a clean effect size from a bandit?

It is harder, since allocation varies with time. Methods exist but they are more complex and more fragile than a fixed split.

What if the best option changes over time?

Keep exploring a small share of traffic, or use an approach that discounts older evidence. A fully converged bandit will not notice a change.

Are bandits suitable for email?

Yes, particularly for subject lines, where you can send to a portion first and adapt. Bear in mind the audience is not identical across sends.

Keep reading

More on AI & Machine Learning

Start here

Want machine learning project details from us?

Tell us what you are trying to predict and roughly what data you hold. We will come back with an honest view on whether machine learning is the right tool, what the work would involve and a realistic cost range. If a spreadsheet would do the job, we will say so.

  1. You tell us what you needTwo minutes on the form, or a message on WhatsApp.
  2. A senior engineer reviews itAnd comes back with questions, a realistic range and an honest view on fit.
  3. Free 30-minute scoping callWe talk through scope, options and a realistic estimate — with no obligation.
Free estimateNo obligation

Talk to someone who builds this

Send a short brief and we will come back with an honest view and a realistic range.

Takes under a minute. We never share your details.

  • Free consultation
  • No commitment
  • NDA on request

Prefer to talk? Book a free 30-minute call →