Medical Image Triage: What Machine Learning Can and Cannot Do
Last updated:
Triage is the realistic job
A radiology department with a backlog of several days reads studies broadly in order of arrival and urgency category. A chest X-ray with an unexpected serious finding can sit in the queue behind dozens of normal ones. A dermatology service receives hundreds of referral photos a week and most turn out to be benign.
The most established use of machine learning in medical imaging is not diagnosis. It is prioritisation: flag the studies most likely to contain something urgent so a clinician looks at them first, and help sort the rest. A tool that reorders a worklist sounds modest. Getting a suspected bleed on a head scan read hours earlier is not.
This post is written for clinic owners, digital health founders and operations leads, not as clinical guidance. The decisions described here belong with clinicians and regulatory specialists.
What imaging machine learning can do well
- Worklist prioritisation. Moving studies with likely critical findings to the top of the reading queue.
- Detection support. Highlighting regions of interest, such as possible nodules or fractures, for the clinician to review.
- Measurement. Consistent, repeatable measurements that are tedious by hand.
- Quality checks. Flagging images that are poorly positioned, cropped or unreadable, so they can be retaken before the patient leaves.
- Routing referrals. Helping sort image-based referrals into urgency bands for clinical review.
What it cannot do
| Limitation | Why it matters |
|---|---|
| Replace the clinician's read | Models look for what they were trained on; clinicians notice the unexpected finding next to it |
| Generalise automatically | Performance can drop on different scanners, protocols, image quality or patient populations |
| Use context it does not have | History, symptoms and prior imaging change the meaning of a finding |
| Explain itself reliably | Heat maps show where the model looked, not a clinical reason |
| Stay accurate without monitoring | Equipment changes and case mix shifts can quietly degrade performance |
A triage tool that misses a finding is dangerous only if people start treating 'not flagged' as 'normal'. That is a workflow risk, not a model property.
The subtle risk is automation bias. If a normal-looking study is labelled low priority, a busy clinician may read it faster and with less attention. Good deployments are explicit that the absence of a flag means nothing clinically.
Regulation comes first, not last
Software intended to inform diagnosis or treatment is typically regulated as a medical device. In the UK that means the medical device regulations overseen by the MHRA; in the EU, the Medical Device Regulation, with the EU AI Act adding requirements for AI systems that are, or are part of, regulated medical devices. In the US, the FDA regulates such software. The details depend on intended purpose, and a small wording change in how a tool is described can change its classification.
- Buying a tool: check it holds the right approvals for your market and your intended use
- Building a tool: plan for a quality management system, clinical evaluation and regulatory submission from the start
- In every case: clinical safety risk management and a named clinical lead
- Patient data used for development needs a proper legal basis and governance, never a quick export
This is why most providers should buy approved imaging tools rather than build them. Custom development is better aimed at the workflow and data around them.
Validate locally before trusting it
A tool's published accuracy comes from someone else's data. Before relying on it, a service should test it on its own images.
- Assemble a retrospective set of local studies with confirmed outcomes, including difficult and rare cases
- Run the tool silently, without showing results to clinicians
- Compare its flags with the confirmed outcomes, by scanner, site and patient group
- Agree acceptable performance and what happens when it is not met
- Go live with monitoring, audit of misses and a route for clinicians to report concerns
Shadow running is not glamorous, and it is the step that catches the tool that works beautifully on one scanner and poorly on the older one in the satellite clinic.
Where software teams genuinely help
The imaging model is often bought. What is rarely available off the shelf is the integration: moving images from the imaging archive to the tool and results back into the reporting system, building worklist views clinicians actually use, logging outputs for audit, and producing monitoring reports on performance over time.
That is the work SpiderHunts takes on within custom healthcare software projects, alongside the administrative side covered in AI in healthcare administration, where the regulatory burden is lighter and quick wins are more common. We do not build diagnostic models for clinical use without a regulated partner and clinical leadership in place.
There is plenty of useful imaging-adjacent work that sits outside diagnosis. Checking that referral photos meet a minimum quality before they reach a clinician, matching images to the right patient record, or summarising the operational picture of a reporting backlog for managers all save time without claiming to interpret anything clinically. Those are often the better first projects, precisely because they can go live in weeks and they build the data plumbing any later clinical tool will need.
Frequently asked questions
Can AI read X-rays and scans instead of radiologists?
Is medical imaging AI regulated as a medical device?
Why does imaging AI perform differently at different hospitals?
Should a private clinic build its own imaging AI?
What is automation bias in medical imaging?
Exploring imaging AI for your service?
Tell us what you are considering and where it would sit in the clinical pathway. We will help you ask the right validation and integration questions before you commit.
Related services
What we build for problems like this one