The problem with knowing what you buy
Ask a mid-sized business what it spends on packaging across all suppliers and the answer usually takes days to assemble. Invoice lines are abbreviated, supplier names vary, and the same item is coded differently by different sites.
Without that view, consolidation opportunities stay invisible. The purpose of spend classification is not tidiness - it is being able to walk into a negotiation knowing your total volume.
Why the descriptions are so hard
Invoice line text is some of the messiest data in a business. It is written for a human who already knows the context, often by a system with a character limit.
- Abbreviations that vary by site - 'crrgtd bx 5ply' and 'corrugated box 5 ply'
- Supplier part numbers with no descriptive content at all
- Free-text notes mixed into the description field
- The same physical item under several different codes across entities
- Multi-line invoices where the useful description is on a header line
This is why simple keyword rules plateau quickly. They work for the obvious cases and leave a long tail that is exactly where unmanaged spend hides.
Get the taxonomy right first
The category structure decides whether the output is useful. A taxonomy inherited from the finance chart of accounts usually groups spend by how it is reported rather than how it is bought, which is the wrong cut for negotiation.
| Taxonomy built for | Groups by | Useful for |
|---|---|---|
| Financial reporting | Cost centre and account | Statutory accounts |
| Procurement | What is bought and from which market | Consolidation, negotiation |
| Operations | Where it is consumed | Budget ownership |
You may need more than one view, which argues for tagging each line with several attributes rather than forcing a single hierarchy.
A practical build sequence
- Normalise supplier names first - fuzzy matching and a manual review of the top suppliers by value. This alone often reveals consolidation opportunities.
- Label a stratified sample by hand, covering high-value lines and a genuine sample of the long tail.
- Train a classifier on the description text plus supplier and account code as features.
- Route low-confidence lines to a review queue rather than forcing a category.
- Feed reviewed lines back into training, so the long tail improves over time.
Weighting by value rather than line count matters throughout. Ninety per cent of lines classified correctly is a poor result if the misclassified ten per cent carries most of the spend.
Measuring it the way procurement will
Report coverage by value, not by row count, and report the unclassified residual prominently. A shrinking residual is the honest measure of progress.
The business case is not classification accuracy - it is the consolidation and negotiation that follow. Tracking spend brought under management is what tells you whether the project paid.
You cannot negotiate a volume you cannot count.