Quality
- An evaluation set of real cases, running automatically on every change
- A measured accuracy figure, above the threshold agreed before the build
- Confidence thresholds set from a shadow run, not from intuition
Failure behaviour
- Decided behaviour for outage, rate limit, timeout and bad output
- A tested switch that disables the AI and falls back to manual
- Quarantine for repeated failures, with the input preserved
Test the switch before launch, with the business owner operating it. A control nobody has used is a control nobody will reach for.
Cost and monitoring
- A hard cap per day and per user, with a decided behaviour at the cap
- Alerts on quality drift, cost per item and output volume
- Alerts routed to people who can act, not to a shared channel nobody reads
Access and records
- Credentials in your accounts, with expiry monitored
- Audit trail recording input, context, model version, output and reviewer
Ownership
- A named business owner for quality
- A named technical contact for changes
- A runbook covering each alert and the first three checks
- A ninety-day review booked, with the baseline attached
Eleven items. Most take hours rather than days, and the cost of skipping them is paid at the worst possible moment.