Think Build Implement Repeat
London, UK +44 7367 067226
WhatsApp FOLLOW f in X
  1. Home
  2. Blog
  3. Checking Model Performance by Subgroup
AI & Machine Learning

Checking Model Performance by Subgroup

A single accuracy figure hides groups the model handles badly. How to find them, which slices to check, and what to do when performance is uneven.

Updated 2 min readBy SpiderHunts Technologies

Free estimateNo obligation

Get a free estimate

Tell us what you need. A senior engineer reads every enquiry.

Takes under a minute. We never share your details.

  • Free consultation
  • No commitment
  • NDA on request

Prefer to talk? Book a free 30-minute call →

Quick answer — TL;DR

Overall accuracy is an average that conceals variation. Break performance down by the slices that matter commercially and ethically - customer size, region, channel, product type, tenure - and treat a weak subgroup as a finding rather than a rounding error.

The average is not the experience

A model reported as 88% accurate is 88% accurate across whatever mix the test set happened to contain. Individual groups may be far better or far worse, and the customers in the worse groups experience the model as unreliable.

This matters commercially as well as ethically. If the group the model handles worst happens to be your highest-value customers, the headline figure is actively misleading about the model's business value.

Which slices to check

  • Commercially important groups - largest customers, highest-margin products, strategic accounts
  • Data-poor groups - new customers, recently added products, small branches
  • Structurally different groups - a different channel, a different country, a different contract type
  • Protected characteristics, where the decision affects people - required in some contexts and good practice in most
  • Time - performance in the most recent period against earlier ones

The data-poor category is usually where the worst performance hides, for an obvious reason: the model had least to learn from. It is also frequently where growth is.

Doing it without fooling yourself

Slicing many ways means many chances to find an apparently poor group by chance. Small slices produce noisy estimates, and a group of forty cases can look alarming for no reason.

  1. Set a minimum slice size below which you report uncertainty rather than a point estimate.
  2. Show confidence intervals, not bare percentages, so small groups visibly carry wide ranges.
  3. Decide the slices before looking, so you are testing rather than hunting.
  4. Repeat the check on a later period before acting - a genuine weakness persists.

What to do about a weak subgroup

CauseResponse
Too few training examplesCollect more, or weight the group up in training
Genuinely different behaviourConsider a separate model for that group
Missing features for that groupFind features that exist for them
Inherently harder to predictSet a different threshold, or exclude and handle manually

That last row is legitimate and under-used. A model that abstains for a group it handles badly, routing those cases to a person, is more honest than one that produces unreliable scores everywhere.

Put it in the documentation

Subgroup performance belongs in whatever documentation accompanies the model, alongside the headline figure. Anyone deciding whether to rely on it needs to know where it is weak.

It also protects the project. A known, documented weakness that someone chose to accept is a managed risk; the same weakness discovered later by a customer is an incident.

An average accuracy figure is a promise to nobody in particular.

FAQ

Frequently asked questions

The questions readers ask us after this guide.

Still have a question?

Ask us directly — a senior engineer will get back to you.

Ask about your project

How many subgroups should we check?

The ones where a difference would change a decision. Decide the list in advance rather than slicing until something looks bad.

What if a small group looks much worse?

Check whether the sample is large enough to be meaningful, then re-check on a later period. Persistent weakness is real; a single noisy estimate is not.

Is this the same as fairness testing?

Fairness testing is a subset, focused on protected characteristics and usually with legal weight. The general technique is the same.

Should we build separate models per group?

Sometimes, where behaviour genuinely differs and each group has enough data. It adds maintenance, so it needs to earn its place.

Keep reading

More on AI & Machine Learning

Start here

Want machine learning project details from us?

Tell us what you are trying to predict and roughly what data you hold. We will come back with an honest view on whether machine learning is the right tool, what the work would involve and a realistic cost range. If a spreadsheet would do the job, we will say so.

  1. You tell us what you needTwo minutes on the form, or a message on WhatsApp.
  2. A senior engineer reviews itAnd comes back with questions, a realistic range and an honest view on fit.
  3. Free 30-minute scoping callWe talk through scope, options and a realistic estimate — with no obligation.
Free estimateNo obligation

Talk to someone who builds this

Send a short brief and we will come back with an honest view and a realistic range.

Takes under a minute. We never share your details.

  • Free consultation
  • No commitment
  • NDA on request

Prefer to talk? Book a free 30-minute call →