The alert problem
Rule-based monitoring generates alerts from thresholds - transaction size, velocity, jurisdiction, pattern. Set them tightly and the team drowns; set them loosely and the control weakens. Most firms end up with far more alerts than analysts.
The result is a backlog, and a backlog is itself a control failure. Genuine issues sit in a queue behind hundreds of routine transactions that happened to cross a threshold.
Prioritise, do not replace
The realistic role for a model here is triage: keep the rules that define your risk appetite, and use a model to rank the resulting alerts by how likely they are to warrant investigation.
That framing matters for more than modesty. Rules are explainable, auditable and defensible to a regulator. Replacing them with a model that produces alerts for reasons nobody can articulate creates a different and larger problem.
- Rules continue to define what generates an alert - the control stays intact
- The model orders the queue so analysts see the most likely cases first
- Nothing is closed automatically without human review unless your regulator has agreed to it
- Every ranking decision is recorded with the factors that drove it
Training data is the hard part
A model needs labelled outcomes, and here the labels are weak. An alert closed as no further action means an analyst was not persuaded - not that nothing was happening. Confirmed cases are rare and slow to arrive.
This has practical consequences. The model learns to predict analyst decisions, which embeds any inconsistency in past investigations. Sampling reviews to check consistency before training is worth the effort, and the findings are useful regardless.
What has to be demonstrable
Whatever is built must be explainable at the level of an individual decision, documented, tested for bias across customer groups, and monitored for drift with results retained.
| Requirement | What it means in practice |
|---|---|
| Explainability | Show the factors behind each ranking, for an individual alert |
| Auditability | Version the model and retain what was live on any given date |
| Fairness testing | Check performance across customer segments and geographies |
| Ongoing monitoring | Detect drift, and evidence that you did |
These requirements tend to favour simpler, more interpretable models than a pure accuracy contest would select. That is a reasonable trade in a regulated control.
A cautious deployment path
Run the model in shadow first: score alerts but do not change the queue, and compare its ranking against what analysts actually found. That produces evidence for the regulator conversation and catches problems before they matter.
Only then change the working order, and keep sampling low-ranked alerts indefinitely so you can prove the model is not systematically missing a category. That sampling is the control on the control.
In a regulated process, an unexplainable improvement is not an improvement.