Think Build Implement Repeat
London, UK +44 7367 067226
WhatsApp FOLLOW f in X
  1. Home
  2. Blog
  3. Building a Labelling Workflow for Business Data
AI & Machine Learning

Building a Labelling Workflow for Business Data

Labelled examples are the constraint on most projects. How to write guidelines, measure agreement and keep quality up once the novelty wears off.

Updated 2 min readBy SpiderHunts Technologies

Free estimateNo obligation

Get a free estimate

Tell us what you need. A senior engineer reads every enquiry.

Takes under a minute. We never share your details.

  • Free consultation
  • No commitment
  • NDA on request

Prefer to talk? Book a free 30-minute call →

Quick answer — TL;DR

Label quality sets the ceiling on model quality. Written guidelines with worked edge cases, two people labelling an overlapping sample, and a measured agreement rate are the three things that separate a usable dataset from an expensive one.

The bottleneck nobody budgets for

Supervised learning needs examples with known answers. For most business problems those do not exist in usable form, and someone has to create them. This is routinely underestimated in both time and difficulty.

It is also delegated badly. Labelling gets handed to whoever is least busy, without guidelines, and the resulting inconsistency limits the model permanently. No amount of modelling recovers from labels that contradict each other.

Write the guidelines before labelling anything

Labelling guidelines are a short document that defines each category and, more importantly, rules on the awkward cases. The awkward cases are the whole point - the obvious ones need no guidance.

  1. Define each label in a sentence, in the business's own vocabulary.
  2. Give two or three real examples per label, taken from your actual data.
  3. Write explicit rules for the boundaries - when something could be two labels, which wins.
  4. Provide a way to mark 'unclear' rather than forcing a guess, and review those regularly.
  5. Version the document, because it will change and labels applied under different versions are not comparable.

That last point is often overlooked. If the guidance changes halfway through, earlier labels need revisiting or at minimum flagging.

Measure agreement, and take it seriously

Have two people label the same sample independently and measure how often they agree. This single number tells you more about project feasibility than any amount of discussion.

AgreementWhat it means
HighThe task is well defined; a model can plausibly learn it
ModerateGuidelines need work, or categories overlap
LowThe task as defined may not be learnable - rethink before spending

Low agreement is not a failure of the labellers. It is information: the categories are ambiguous, or the task genuinely requires context the data does not contain. Finding that out in week one is worth a great deal.

Keeping quality up over time

Labelling is repetitive and quality decays. A few practical measures help more than exhortation.

  • Keep an overlapping sample throughout, not just at the start, so drift is visible
  • Insert occasional items with known answers as a quality check
  • Cap session length - accuracy falls off after a couple of hours of continuous labelling
  • Give labellers a route to raise cases the guidelines do not cover, and actually update the guidelines
  • Review the 'unclear' pile periodically; it is where new categories reveal themselves

Who should do it

Domain experts produce better labels but are expensive and busy. A common and workable pattern is for experts to write the guidelines and label a gold-standard set, with a larger volume labelled by trained non-experts and measured against that gold standard.

Outsourcing is viable for general tasks and poor for ones needing your specific domain knowledge. If a labeller needs to know how your business categorises a product return, an external vendor will struggle regardless of their quality process.

Two people who cannot agree on a label have told you the model will not either.

FAQ

Frequently asked questions

The questions readers ask us after this guide.

Still have a question?

Ask us directly — a senior engineer will get back to you.

Ask about your project

How many labelled examples do we need?

It depends on how many categories, how distinct they are and how much variety exists. A few hundred per category is a common starting point, with rare categories needing relatively more.

Can a language model do the labelling?

It can produce a first pass that humans correct, which is often much faster. You still need a human-labelled sample to measure it against, or you cannot tell whether it is right.

What if experts disagree with each other?

Then the task definition needs work. Bring them together on the disputed cases and write the resolution into the guidelines.

Should we label everything or a sample?

A well-chosen sample, usually stratified so rare but important categories are represented. Labelling everything is rarely necessary and often unaffordable.

Keep reading

More on AI & Machine Learning

Start here

Want machine learning project details from us?

Tell us what you are trying to predict and roughly what data you hold. We will come back with an honest view on whether machine learning is the right tool, what the work would involve and a realistic cost range. If a spreadsheet would do the job, we will say so.

  1. You tell us what you needTwo minutes on the form, or a message on WhatsApp.
  2. A senior engineer reviews itAnd comes back with questions, a realistic range and an honest view on fit.
  3. Free 30-minute scoping callWe talk through scope, options and a realistic estimate — with no obligation.
Free estimateNo obligation

Talk to someone who builds this

Send a short brief and we will come back with an honest view and a realistic range.

Takes under a minute. We never share your details.

  • Free consultation
  • No commitment
  • NDA on request

Prefer to talk? Book a free 30-minute call →