Think Build Implement Repeat
London, UK +44 7367 067226
WhatsApp FOLLOW f in X
  1. Home
  2. Blog
  3. Automatic Spend Classification for Procurement
AI Integration

Automatic Spend Classification for Procurement

You cannot negotiate what you cannot see. How to categorise messy supplier spend automatically, and why the taxonomy matters more than the model.

Updated 2 min readBy SpiderHunts Technologies

Free estimateNo obligation

Get a free estimate

Tell us what you need. A senior engineer reads every enquiry.

Takes under a minute. We never share your details.

  • Free consultation
  • No commitment
  • NDA on request

Prefer to talk? Book a free 30-minute call →

Quick answer — TL;DR

Spend classification is a text classification problem on short, abbreviated, inconsistent descriptions. The model is the easy part; agreeing a taxonomy that matches how you actually negotiate is the work that decides whether it is useful.

The problem with knowing what you buy

Ask a mid-sized business what it spends on packaging across all suppliers and the answer usually takes days to assemble. Invoice lines are abbreviated, supplier names vary, and the same item is coded differently by different sites.

Without that view, consolidation opportunities stay invisible. The purpose of spend classification is not tidiness - it is being able to walk into a negotiation knowing your total volume.

Why the descriptions are so hard

Invoice line text is some of the messiest data in a business. It is written for a human who already knows the context, often by a system with a character limit.

  • Abbreviations that vary by site - 'crrgtd bx 5ply' and 'corrugated box 5 ply'
  • Supplier part numbers with no descriptive content at all
  • Free-text notes mixed into the description field
  • The same physical item under several different codes across entities
  • Multi-line invoices where the useful description is on a header line

This is why simple keyword rules plateau quickly. They work for the obvious cases and leave a long tail that is exactly where unmanaged spend hides.

Get the taxonomy right first

The category structure decides whether the output is useful. A taxonomy inherited from the finance chart of accounts usually groups spend by how it is reported rather than how it is bought, which is the wrong cut for negotiation.

Taxonomy built forGroups byUseful for
Financial reportingCost centre and accountStatutory accounts
ProcurementWhat is bought and from which marketConsolidation, negotiation
OperationsWhere it is consumedBudget ownership

You may need more than one view, which argues for tagging each line with several attributes rather than forcing a single hierarchy.

A practical build sequence

  1. Normalise supplier names first - fuzzy matching and a manual review of the top suppliers by value. This alone often reveals consolidation opportunities.
  2. Label a stratified sample by hand, covering high-value lines and a genuine sample of the long tail.
  3. Train a classifier on the description text plus supplier and account code as features.
  4. Route low-confidence lines to a review queue rather than forcing a category.
  5. Feed reviewed lines back into training, so the long tail improves over time.

Weighting by value rather than line count matters throughout. Ninety per cent of lines classified correctly is a poor result if the misclassified ten per cent carries most of the spend.

Measuring it the way procurement will

Report coverage by value, not by row count, and report the unclassified residual prominently. A shrinking residual is the honest measure of progress.

The business case is not classification accuracy - it is the consolidation and negotiation that follow. Tracking spend brought under management is what tells you whether the project paid.

You cannot negotiate a volume you cannot count.

FAQ

Frequently asked questions

The questions readers ask us after this guide.

Still have a question?

Ask us directly — a senior engineer will get back to you.

Ask about your project

Can this run on invoice data alone?

Usually yes, though purchase order data where it exists is cleaner and improves results. Supplier master data helps considerably.

How accurate does spend classification need to be?

Accurate enough by value that category totals are trustworthy for negotiation. Perfect line-level accuracy is rarely necessary.

Should we use a standard taxonomy?

A standard scheme gives comparability, but only if it reflects how you buy. Many businesses use a standard at the top levels and their own categories below.

What about spend in multiple currencies and entities?

Normalise to one currency at a consistent rate for analysis and keep the original for audit. Entity differences in coding are usually the bigger problem.

Keep reading

More on AI Integration

Start here

Want machine learning project details from us?

Tell us what you are trying to predict and roughly what data you hold. We will come back with an honest view on whether machine learning is the right tool, what the work would involve and a realistic cost range. If a spreadsheet would do the job, we will say so.

  1. You tell us what you needTwo minutes on the form, or a message on WhatsApp.
  2. A senior engineer reviews itAnd comes back with questions, a realistic range and an honest view on fit.
  3. Free 30-minute scoping callWe talk through scope, options and a realistic estimate — with no obligation.
Free estimateNo obligation

Talk to someone who builds this

Send a short brief and we will come back with an honest view and a realistic range.

Takes under a minute. We never share your details.

  • Free consultation
  • No commitment
  • NDA on request

Prefer to talk? Book a free 30-minute call →