Think Build Implement Repeat
London, UK +44 7367 067226
WhatsApp FOLLOW f in X
  1. Home
  2. Blog
  3. Machine Learning Terms That Change What You Decide
AI & Machine Learning

Machine Learning Terms That Change What You Decide

Not a full glossary - the dozen terms that appear in proposals and reports and genuinely change what you should do about them.

Updated 2 min readBy SpiderHunts Technologies

Free estimateNo obligation

Get a free estimate

Tell us what you need. A senior engineer reads every enquiry.

Takes under a minute. We never share your details.

  • Free consultation
  • No commitment
  • NDA on request

Prefer to talk? Book a free 30-minute call →

Quick answer — TL;DR

A short list of terms whose meaning changes a decision: baseline, leakage, drift, calibration, precision and recall, threshold, overfitting and holdout. Knowing these is enough to read a proposal or a model report critically.

Enough to read a report critically

Full glossaries are long and mostly cover terms you will never need to act on. This is the shorter list - terms that, if you understand them, change what you ask and what you decide.

Each entry below says what it means and why it matters to you rather than to the person building the model.

The terms about whether it works

TermWhat it meansWhy you care
BaselineHow well the current method performsWithout it, no improvement can be proved
HoldoutData kept back to test onIf it was used for tuning, the score is optimistic
OverfittingLearned the training data rather than the patternLooks excellent, fails live
LeakageUsed information unavailable at prediction timeThe score is unachievable; ask what fields
DriftThe world changed since trainingMeans ongoing monitoring is required

If a report gives an accuracy figure and does not mention the baseline or the holdout, those are the first two questions to ask.

The terms about what it does

  • Precision - of the cases it flagged, how many were right. Matters when acting on a flag is costly.
  • Recall - of the cases that mattered, how many it caught. Matters when missing one is costly.
  • Threshold - the cut-off turning a score into an action. A business decision, not a technical default.
  • Calibration - whether a 70% score means it happens 70% of the time. Essential if you multiply the score by money.
  • Class imbalance - the thing you are predicting is rare. Makes accuracy misleading; ask for precision and recall instead.

Precision and recall trade against each other. Any claim to have improved both substantially deserves a question about what changed to allow it.

The terms about running it

TermWhat it meansWhy you care
InferenceMaking a prediction with a trained modelThis is the recurring cost
RetrainingRebuilding on newer dataNeeds a schedule, a budget and an owner
FeatureAn input the model usesEach one is a dependency that can break
PipelineThe path from raw data to predictionWhere most production failures happen

Two phrases worth pushing back on

'The model is 95% accurate' - on what data, against what baseline, and with what class balance? Accuracy on an imbalanced problem can be achieved by always predicting the majority.

'It learns and improves over time' - only if there is a retraining process with someone running it. Models do not improve on their own, and the phrase frequently conceals that nobody has planned for retraining.

Ask what the baseline was. It is the single question that most often changes the conversation.

FAQ

Frequently asked questions

The questions readers ask us after this guide.

Still have a question?

Ask us directly — a senior engineer will get back to you.

Ask about your project

Do I need to understand the algorithms?

No. Understanding what was measured, against what, and how it will be maintained is far more useful.

What is the most useful single question?

What does the current process achieve, measured the same way? Without it no number means anything.

Is high accuracy always good?

No. On rare events it can be achieved by predicting the common case every time. Ask for precision and recall.

What should a model report contain?

The baseline, the test data, the metric and why it was chosen, performance by important subgroup, and known weaknesses.

Keep reading

More on AI & Machine Learning

Start here

Want machine learning project details from us?

Tell us what you are trying to predict and roughly what data you hold. We will come back with an honest view on whether machine learning is the right tool, what the work would involve and a realistic cost range. If a spreadsheet would do the job, we will say so.

  1. You tell us what you needTwo minutes on the form, or a message on WhatsApp.
  2. A senior engineer reviews itAnd comes back with questions, a realistic range and an honest view on fit.
  3. Free 30-minute scoping callWe talk through scope, options and a realistic estimate — with no obligation.
Free estimateNo obligation

Talk to someone who builds this

Send a short brief and we will come back with an honest view and a realistic range.

Takes under a minute. We never share your details.

  • Free consultation
  • No commitment
  • NDA on request

Prefer to talk? Book a free 30-minute call →