The tolerance is different
An internal tool that gets something wrong costs an apology in a meeting. A customer-facing one that gets something wrong may create a commitment you have to honour, or a screenshot that circulates.
So we build them to a different standard, and we usually recommend proving the capability internally first.
The five guardrails
- Strict grounding — answers only from your own material, with refusal rather than inference
- Visible escalation to a person, always available, never buried
- Disclosure that it is an AI assistant, at the start
- Rate limiting and abuse handling, because it will be probed
- Logging of every interaction, for dispute resolution and improvement
It must be able to say it does not know
A model that invents a refund policy has created a commitment. We test refusal behaviour as carefully as we test answers, because it is the property that keeps customer-facing AI safe.
Escalation design decides satisfaction
The biggest failure in customer-facing AI is not wrong answers, it is trapping people. Escalate on the second failed attempt, on any sign of frustration, on anything about money or cancellation, and always show a route to a person.
When it escalates, hand over the full context. Making the customer repeat everything is where goodwill is lost, and it is entirely avoidable.
What to measure
- Resolution rate — no re-contact within seven days
- Satisfaction on escalated cases, which reveals a bad system faster than anything
- Time to human when escalation happens
- Re-contact rate, which catches confidently wrong answers that deflection counts as wins