Measuring Marketing Impact With Machine Learning Attribution
Last updated:
Every platform says it drove the sale
Add up the conversions claimed by Google Ads, Meta, your email platform and your affiliate network and you will often get more sales than the business made. Each tool counts the customers it touched. None of them counts only the customers it changed.
Machine learning attribution promises to sort this out by modelling each channel's contribution from the data. It is a genuine improvement on last-click. It is also still a model of correlations, and correlations in marketing data are full of traps. The way to measure marketing impact honestly is to combine attribution with experiments that test it.
How machine learning attribution works, briefly
There are two broad families. Multi-touch attribution looks at individual customer journeys and assigns credit to touchpoints, using methods such as Shapley values or Markov chains to estimate how conversion probability changes when a channel is removed. Marketing mix modelling works at an aggregate level, using weekly spend and sales over time, plus factors such as seasonality and pricing, to estimate each channel's effect.
Multi-touch models are increasingly hampered by missing journey data as tracking erodes. Mix models have come back into fashion because they need no user-level tracking, and open-source tools have made them far more accessible. We cover them in more depth in marketing mix modelling.
Where attribution models mislead
- Selection bias. Retargeting and branded search reach people already about to buy. Models credit them heavily.
- Correlated spend. If you always raise TV and paid social together before Christmas, a mix model cannot separate their effects.
- Missing data. Channels that are hard to track, such as podcasts, word of mouth or offline, get under-credited.
- Too little variation. If a channel's spend barely changes, the model has nothing to learn its effect from and leans on its assumptions.
- Assumptions presented as findings. Prior settings for ad decay and saturation can drive the result more than the data does.
An attribution model tells you where credit is consistent with the data. An experiment tells you what happens when you change something.
What incrementality testing is
Incrementality testing measures the extra sales caused by marketing by comparing a group exposed to it with a comparable group that is not. The difference is the lift. Common designs:
| Method | How it works | Good for | Watch out for |
|---|---|---|---|
| Audience holdout | Randomly exclude a share of an audience from a campaign | Email, retargeting, CRM audiences | Holdout must be truly random and untouched |
| Platform lift study | Ad platform randomises exposure and reports lift | Paid social and some search | Minimum spend requirements, platform marks its own homework |
| Geo experiment | Switch spend on or off in matched regions | Channels with no user-level control | Needs enough regions and sales per region |
| Time-based switch-off | Pause a channel for a period | Quick sanity checks | Seasonality and other changes confuse results |
For businesses selling across the UK, geo experiments often use groups of postcode areas or TV regions. For businesses selling mainly in one city, geo tests may not be feasible, and audience holdouts become the main tool.
An illustrative test on a modest budget
Imagine an online retailer spending around 15,000 pounds a month, with a large share on branded search and retargeting, both of which look excellent in the platform reports. The team suspects much of those sales would happen anyway.
- Choose retargeting first, because it is easiest to hold out at audience level
- Exclude a random 20 percent of the retargeting audience for six weeks
- Compare purchase rate and revenue per person between exposed and holdout groups using the store's own order data, not platform reporting
- Calculate incremental cost per sale and compare it with the cost per sale the platform claimed
- Feed the result into the mix model as a calibration, and set retargeting spend on the incremental figure
A result showing lift well below the platform's claims is common for retargeting, but it is not guaranteed, which is exactly why you run the test. Branded search can be tested next by pausing it in selected regions, watching whether organic clicks absorb the traffic.
Combining the two
Experiments are slow and cover one question at a time. Models are always on but uncalibrated. The practical approach is a loop: run the mix model, identify the channel where the model is least certain or the spend is largest, test it, and use the experiment to calibrate the model's estimate for that channel. Over a year you build a model anchored in several real experiments.
For smaller budgets, skip the full model and simply run one incrementality test per quarter on the biggest line item. Our note on attribution for small marketing budgets covers lighter options.
When this is overkill
If you spend a few thousand a month across two channels, a spreadsheet of spend against sales, plus one simple holdout test, will teach you more than a machine learning model. Mix models need enough history, typically two years of weekly data, and meaningful variation in spend.
Where SpiderHunts helps is building the data foundation, analysing tests against real order data rather than platform dashboards, and fitting mix models as part of machine learning projects when the budget justifies it. We are also happy to tell a business that its best next step is a single, well-run holdout test.
Frequently asked questions
What is incrementality in marketing?
Is marketing mix modelling better than multi-touch attribution?
How long should an incrementality test run?
Can small businesses run incrementality tests?
Why do ad platforms report more conversions than I actually have?
Not sure which marketing actually drives sales?
Tell us your channels, rough spend and where sales are recorded. We will suggest the simplest test that would tell you something you do not already know.