Machine Learning for Media and Publishing
Last updated:
Referral traffic is less reliable, so readers you know matter more
A specialist publisher with two million monthly visits and 15,000 paying subscribers used to plan around search and social traffic. That traffic now swings more than it did: search results include generated answers, AI assistants summarise articles without sending a visit, and social platforms change what they show without notice.
The business that holds up best is the one that knows its readers. Registered users, newsletter subscribers and paying members are audiences a publisher controls, and they generate exactly the behavioural data machine learning needs. That is where we would put the effort first.
Where machine learning helps a publisher
- Propensity to subscribe. Scoring anonymous and registered readers by how likely they are to pay, based on visit frequency, topics read and depth of reading.
- Dynamic paywall decisions. Deciding whether a given reader sees a hard wall, a registration wall or another free article.
- Subscriber churn prediction. Spotting paying readers whose engagement is falling before the renewal date arrives.
- Recommendations. Suggesting related articles and newsletters that deepen engagement rather than chasing clicks.
- Tagging and archive organisation. Classifying thousands of back-catalogue articles into a consistent taxonomy so topic pages and recommendations work.
- Headline and send-time testing. Using bandit methods that shift traffic towards better options while a test is still running.
Dynamic paywalls, done carefully
A fixed metered paywall treats a first-time visitor from a search result and a loyal reader on their fortieth visit the same way. A propensity model lets you vary that: loyal, engaged readers are asked to subscribe sooner, while casual visitors see a registration prompt or a few more free articles so they come back.
- Start with a simple rules-based segmentation and measure it
- Build a propensity model on registered readers, where the data is richest
- Test model-driven paywall decisions against the rules on a random share of traffic
- Watch subscriptions, registrations and total advertising revenue together
- Keep a permanent control group so you know the model is still earning its place
Measuring only subscriptions is the classic mistake. An aggressive wall can lift conversions this month while cutting the traffic and registrations that produce next year's subscribers.
Churn among subscribers is often quieter than expected
Most subscribers do not cancel angrily. They stop opening the newsletter, read less often and then cancel at renewal, sometimes months later. A churn model trained on reading behaviour, payment method, tenure and plan type can flag them while there is still time to re-engage.
The interventions are the hard part: a personalised digest of topics they used to read, a plan change rather than a discount, a reminder of a benefit they never used. Discounts work, but offering them to everyone who looks at risk teaches subscribers to threaten to leave. Our post on building a churn prediction model covers the modelling detail.
The cheapest subscriber to keep is the one you noticed drifting in March, not the one who clicked cancel in September.
Tagging the archive is the unglamorous win
Most publishers have years of content tagged inconsistently by different editors: the same topic under four names, and half the articles untagged. That quietly breaks topic pages, recommendations and any attempt to understand what content drives subscriptions.
Classification models, and increasingly language models with a fixed taxonomy and a validation step, can retag an archive of tens of thousands of articles in days. An editor reviews a sample, fixes the taxonomy where the model reveals gaps, and new articles get suggested tags on publish. It is a modest project that makes every other model work better.
When machine learning is not worth it for a publisher
| Publisher profile | Better first step |
|---|---|
| Under a few thousand subscribers | Clean analytics, simple meter tests, better newsletters |
| No reader login or registration | Build registration before any propensity model |
| Mostly advertising revenue, little first-party data | Improve consented data collection first |
| Large archive, poor tagging | Archive classification before recommendations |
Advertising yield optimisation is also mostly handled by ad platforms now. Building your own bidding models is rarely worth it below very large scale.
Newsletters are the most underused data
For many mid-sized publishers, the newsletter list is the largest group of known readers and the best predictor of who will eventually pay. Open and click history, the topics readers click on, how quickly they open after a send and whether they forward emails all say far more than an anonymous page view.
A simple model trained on newsletter behaviour can decide which readers get a subscription offer, which get a second newsletter recommendation and which should be left alone for a while. It also helps with list health: readers who have not opened in months drag down deliverability, and quietly suppressing them improves the results for everyone else.
- Score newsletter readers by likelihood to subscribe and time offers accordingly
- Recommend additional newsletters from the topics each reader actually clicks
- Flag disengaged addresses for a re-permission campaign before removal
- Measure revenue per thousand sends alongside open rate
How we would work with a publisher
At SpiderHunts we usually start by joining reading behaviour to subscription records, which lives in separate systems at almost every publisher. That joined view answers useful questions before any model exists, such as which topics new subscribers read in their final week before paying.
From there we build the propensity or churn model, connect it to the paywall or email platform and set up the control groups. That work sits within our data science service, with the product integration handled as part of the same project.
Frequently asked questions
How do publishers use machine learning?
Does a dynamic paywall increase subscriptions?
How many subscribers do you need for a churn model?
Can AI tag old articles accurately?
Trying to turn readers into subscribers who stay?
Tell us about your audience data, your paywall and your content system. We will tell you which model would move revenue first, and which ones are not worth the effort at your size.
Related services
What we build for problems like this one