Think Build Implement Repeat
London, UK +44 7367 067226
WhatsApp FOLLOW f in X
AI & Machine Learning

Measuring Marketing Impact With Machine Learning Attribution

Last updated:

Every platform says it drove the sale

Add up the conversions claimed by Google Ads, Meta, your email platform and your affiliate network and you will often get more sales than the business made. Each tool counts the customers it touched. None of them counts only the customers it changed.

Machine learning attribution promises to sort this out by modelling each channel's contribution from the data. It is a genuine improvement on last-click. It is also still a model of correlations, and correlations in marketing data are full of traps. The way to measure marketing impact honestly is to combine attribution with experiments that test it.

How machine learning attribution works, briefly

There are two broad families. Multi-touch attribution looks at individual customer journeys and assigns credit to touchpoints, using methods such as Shapley values or Markov chains to estimate how conversion probability changes when a channel is removed. Marketing mix modelling works at an aggregate level, using weekly spend and sales over time, plus factors such as seasonality and pricing, to estimate each channel's effect.

Multi-touch models are increasingly hampered by missing journey data as tracking erodes. Mix models have come back into fashion because they need no user-level tracking, and open-source tools have made them far more accessible. We cover them in more depth in marketing mix modelling.

Where attribution models mislead

  • Selection bias. Retargeting and branded search reach people already about to buy. Models credit them heavily.
  • Correlated spend. If you always raise TV and paid social together before Christmas, a mix model cannot separate their effects.
  • Missing data. Channels that are hard to track, such as podcasts, word of mouth or offline, get under-credited.
  • Too little variation. If a channel's spend barely changes, the model has nothing to learn its effect from and leans on its assumptions.
  • Assumptions presented as findings. Prior settings for ad decay and saturation can drive the result more than the data does.
An attribution model tells you where credit is consistent with the data. An experiment tells you what happens when you change something.

What incrementality testing is

Incrementality testing measures the extra sales caused by marketing by comparing a group exposed to it with a comparable group that is not. The difference is the lift. Common designs:

MethodHow it worksGood forWatch out for
Audience holdoutRandomly exclude a share of an audience from a campaignEmail, retargeting, CRM audiencesHoldout must be truly random and untouched
Platform lift studyAd platform randomises exposure and reports liftPaid social and some searchMinimum spend requirements, platform marks its own homework
Geo experimentSwitch spend on or off in matched regionsChannels with no user-level controlNeeds enough regions and sales per region
Time-based switch-offPause a channel for a periodQuick sanity checksSeasonality and other changes confuse results

For businesses selling across the UK, geo experiments often use groups of postcode areas or TV regions. For businesses selling mainly in one city, geo tests may not be feasible, and audience holdouts become the main tool.

An illustrative test on a modest budget

Imagine an online retailer spending around 15,000 pounds a month, with a large share on branded search and retargeting, both of which look excellent in the platform reports. The team suspects much of those sales would happen anyway.

  1. Choose retargeting first, because it is easiest to hold out at audience level
  2. Exclude a random 20 percent of the retargeting audience for six weeks
  3. Compare purchase rate and revenue per person between exposed and holdout groups using the store's own order data, not platform reporting
  4. Calculate incremental cost per sale and compare it with the cost per sale the platform claimed
  5. Feed the result into the mix model as a calibration, and set retargeting spend on the incremental figure

A result showing lift well below the platform's claims is common for retargeting, but it is not guaranteed, which is exactly why you run the test. Branded search can be tested next by pausing it in selected regions, watching whether organic clicks absorb the traffic.

Combining the two

Experiments are slow and cover one question at a time. Models are always on but uncalibrated. The practical approach is a loop: run the mix model, identify the channel where the model is least certain or the spend is largest, test it, and use the experiment to calibrate the model's estimate for that channel. Over a year you build a model anchored in several real experiments.

For smaller budgets, skip the full model and simply run one incrementality test per quarter on the biggest line item. Our note on attribution for small marketing budgets covers lighter options.

When this is overkill

If you spend a few thousand a month across two channels, a spreadsheet of spend against sales, plus one simple holdout test, will teach you more than a machine learning model. Mix models need enough history, typically two years of weekly data, and meaningful variation in spend.

Where SpiderHunts helps is building the data foundation, analysing tests against real order data rather than platform dashboards, and fitting mix models as part of machine learning projects when the budget justifies it. We are also happy to tell a business that its best next step is a single, well-run holdout test.

Frequently asked questions

What is incrementality in marketing?

Incrementality is the extra outcome, such as sales or sign-ups, that happens because of a marketing activity, compared with what would have happened without it. It is measured with controlled experiments such as holdout groups or geo tests.

Is marketing mix modelling better than multi-touch attribution?

They answer questions differently. Mix models work without user-level tracking and cover offline channels, but need long histories. Multi-touch models give journey-level detail but suffer from tracking gaps. Calibrating either with experiments makes it far more trustworthy.

How long should an incrementality test run?

Long enough to collect enough conversions in both groups and to cover full buying cycles, often four to eight weeks. Longer consideration cycles need longer tests.

Can small businesses run incrementality tests?

Yes. Email and retargeting holdouts are cheap and simple. Geo experiments and mix models need more scale, but a single audience holdout on your largest campaign is within reach of most businesses.

Why do ad platforms report more conversions than I actually have?

Each platform credits conversions from anyone who saw or clicked its ads within its attribution window, and several platforms may count the same sale. Their reports measure association, not the extra sales they caused.

Keep reading

Not sure which marketing actually drives sales?

Tell us your channels, rough spend and where sales are recorded. We will suggest the simplest test that would tell you something you do not already know.

Book a free 30-minute call Get a project estimate WhatsApp us

Related services

What we build for problems like this one

AI AgentsMachine LearningAI Integration