Think Build Implement Repeat
London, UK +44 7367 067226
WhatsApp FOLLOW f in X
AI Apps

Building Systems That Match Things to Other Things

Last updated:

Where matching beats search

Keyword search misses the candidate whose CV says “client account management” when the brief says “customer success”, and the supplier whose datasheet describes the same specification in different words.

Semantic matching compares meaning rather than strings, which is exactly the gap that costs businesses opportunities.

Rank, do not decide

Present a ranked shortlist with reasons, never a decision. The people using it know things the data does not: who is genuinely available, who fell out with that manager, which supplier was late last time.

Systems that filter automatically reliably discard good options that looked unpromising on paper, and nobody ever sees what was lost.

The matching stack

  1. Exact identifiers first where they exist — product codes, registration numbers. Reliable and often missing.
  2. Normalised text matching on the structured attributes you do have
  3. Semantic similarity for descriptions that differ in wording but not meaning
  4. A human review band in the middle, with the decision remembered so it is never asked twice

Realistic expectations

On messy real-world catalogues or CV data, expect 60–80% matching automatically with high confidence, and the remainder needing review. Anyone promising fully automatic matching across untidy data has not tried it on yours.

The remembered decisions are what shrink the review band over the first few months.

What it costs

A matching application over one data set with a review interface typically £20,000–£50,000. Catalogue size and data quality drive the number far more than the algorithm does.

Frequently asked questions

How much data do we need?

Enough items to make the matching worthwhile — this is not a training problem, so a few thousand records is plenty. Quality of the descriptions matters more than volume.

Can it explain a match?

It can show which parts of each description drove the similarity, which is usually enough for a human to judge. It is not a full explanation and should not be presented as one.

Does it work across languages?

Multilingual matching works reasonably with modern embeddings. Verify on your own data rather than assuming, particularly for technical vocabulary.

How do we stop it recommending the same things repeatedly?

Diversity rules on top of similarity, plus tracking what has already been shown. Pure similarity ranking produces a narrow, repetitive shortlist.

Keep reading

Search missing things a person would find?

Send us a sample of what you match and a few examples keyword search misses. We will tell you whether semantic matching would help.

Book a free 30-minute call Get a project estimate WhatsApp us

Related services

What we build for problems like this one

AI AgentsCustom Software DevelopmentSaaS Development