Two failure modes, opposite directions
Untrained users fail in one of two ways. Some treat the score as fact, following it even where they can see it is wrong, because the system said so. Others dismiss it after one visible error and revert to their own methods permanently.
Both come from the same gap: nobody explained what the number is or how good it has been. Neither is unreasonable given that absence.
What the session needs to cover
- What the number means. Is 70 a probability, a rank or a score out of a hundred? People will assume probability unless told.
- How accurate it has been. Real figures on real cases, including where it does worst.
- When to disregard it. Named situations where the model has no basis - new customers, unusual orders, anything it has not seen.
- How to flag a problem. A route for 'this looks wrong', and evidence that someone reads it.
Half an hour covering those four, with examples from the team's own work, achieves more than an hour on how the model works.
Probabilities need practice
People reason about probability poorly without practice, and that is not a criticism - it is well documented and applies to everyone. A 20% chance of failure feels like 'unlikely, ignore', when across two hundred cases a week it means a great deal.
Frequencies land better than percentages. 'Of every ten flagged like this, about two turn out to be genuine' is understood immediately, where 20% invites rounding to zero.
| Says | Heard as | Better phrasing |
|---|---|---|
| 20% risk | Unlikely, ignore | Two in ten - one or two of today's ten will be real |
| 90% accurate | Basically always right | About one in ten of these will be wrong |
| High confidence | Certain | It has been right about this often before |
Show the track record where they work
Nothing builds calibrated trust like visible history. A small panel showing how the model performed on last month's cases, in the same screen as the predictions, does more than any training session.
Include where it did badly. A model presented as uniformly excellent will lose credibility permanently at the first visible failure; one presented with known weaknesses survives them, because the failure was expected.
Close the loop
People stop reporting problems when nothing visibly happens. If someone flags a prediction and never hears back, they will not flag the next one, and you lose the most valuable feedback available.
Acknowledge reports, say what was found, and where a change results, tell the person who raised it. That is a small process with a disproportionate effect on both model quality and adoption.
Say 'two in ten', not '20%'. One of those gets acted on correctly.