Why RFM refuses to die
RFM sorts customers by how recently they bought, how often, and how much they spend. It has been in use for decades, it can be built in a spreadsheet, and any marketer can explain it to a board.
It also encodes something genuinely predictive. Recency in particular carries a great deal of information about whether someone will buy again, which is why a three-variable method competes with models using hundreds of features.
Any proposal to replace it should start by measuring it honestly as a baseline. A model that beats RFM by a margin too small to change a mailing decision has not earned its running cost.
Where RFM genuinely runs out
- You need a probability, not a bucket. 'Segment 3' does not let you rank a mailing by expected return; 'a 12% chance of ordering in 30 days' does.
- The signal is not in the three variables. Browsing behaviour, product mix, service history, contract terms and seasonality all carry information RFM cannot represent.
- Different customers need different actions. RFM tells you who is valuable, not what would change their behaviour.
- Your categories have very different repurchase cycles. A customer buying annually and one buying weekly cannot share a recency scale sensibly.
A fair comparison
Comparing the two properly means holding the decision constant. Pick the actual campaign - who gets the mailing, who gets the call - and evaluate both methods on that.
| Question | RFM answers | A model answers |
|---|---|---|
| Who are our best customers? | Well | Similarly well |
| Who will buy next month? | Roughly, by proxy | Directly, with a probability |
| Who is about to lapse? | Only once they already have | Before, given the right features |
| Who will respond to this offer? | Not really | Yes, with response history |
That third row is where models usually pay. Recency detects lapse after the fact by definition; a model using changes in behaviour can flag it earlier, while there is still time to act.
The hybrid most businesses should run
In practice the sensible arrangement is often both. Keep RFM as the reporting language everyone understands, and add model scores where a specific decision needs ranking.
This avoids the migration where a business switches wholesale to model segments nobody can interpret, loses the shared vocabulary, and quietly reverts within a year.
If nobody can explain the segment, nobody will design a campaign for it.
What to measure after launch
Judge either method on incremental response, not on how well the scores correlate with spend. High-value customers buy anyway; the question is whether contacting them changed anything.
That means holding out a control group from the start. Without one you cannot tell a good model from a good customer base, and the most common outcome is an enthusiastic report that measures nothing. Our piece on uplift modelling goes into targeting the persuadable rather than the likely.