On a Tuesday afternoon, your support agent tells a customer they can return a custom order after the deadline. They can’t: that policy doesn’t exist. Someone on the team spots it while reading the conversation, and the reaction is nearly universal: open the logs and work out why it said that. It’s the most natural move and the one that prolongs the damage the most.
The thesis in one line: when an AI agent gets something wrong with a customer, the right order is contain, scope the damage, notify, and only then diagnose. Almost everyone starts at the end, and while they diagnose, the agent keeps answering.
What to do when an AI agent fails: contain before you diagnose
Diagnosing takes hours; containing takes minutes. While you hunt for the cause, the agent is still in production with the same instructions, the same data and the same tools that produced the mistake, so every new conversation is another chance to repeat it. Containing isn’t fixing: it’s shrinking what the agent can do until you know what happened. And it has to be reversible, because you’ll do it in a hurry and without a diagnosis.
The options, from least to most drastic:
- Lower its autonomy. Switch it to “drafts, a person sends.” The service stays up and the new damage stops. It only works if you defined a scale of agent autonomy levels beforehand.
- Take away the tool that caused the damage. If it promised a refund, remove its refund permission; if it wrote to the CRM, make it read-only. This is the practical payoff of having thought through an AI agent’s permissions.
- Hand the affected topics to a person, with a fixed, honest message: “A member of our team will take it from here.”
- Switch it off entirely. A last resort, but a legitimate one. It requires a named kill switch and someone authorized to press it without asking three people for permission.
The rule: press the smallest button that stops the damage. A switched-off agent is also an incident, just a visible one.
How many other customers got the same bad answer
This is the question almost nobody has prepared for, and it decides how serious the incident really is. An isolated error and a pattern are handled differently, and from one customer’s conversation you can’t tell which you have. Scoping the damage means answering three questions with data:
- How many were exposed: conversations where the agent said the same thing or something equivalent, from the first time it could have happened until the moment of containment.
- How many acted on it: those who requested the return, accepted a deadline or made a decision based on that answer.
- How many hold a commitment involving money or a deadline: the only ones who need an individual reply today.
None of that is possible unless conversations are stored and searchable. Log retention is usually decided on storage cost or privacy grounds, without this moment in mind; decide it knowing it’s what lets you count the people affected, as explained in how long your agent’s log should last. A good starting rule: the review window opens at the agent’s last change (instruction, data source or model), not at the first complaint.
Notify: who, when and with what message
Notifying comes before diagnosing because the affected customer will find out anyway, and it’s better they hear it from you. You don’t need the cause to say what matters: what happened, what is valid and what isn’t, and what you’re going to do. Three recipients, three moments:
| Who | When | What you tell them |
|---|---|---|
| Whoever decides at the company (management or service owner) | As soon as you confirm an error with impact | What the agent said, since when, to how many people, and what you’ve contained |
| Customers with a commitment involving money or a deadline | As soon as possible, individually, once the scope is known | That the information was wrong, what the right answer is and how it gets resolved |
| Everyone else exposed | Only if they acted on the answer or the topic is sensitive | A short correction, no drama |
Two things that don’t work: waiting for the root cause before writing to the customer, and sending a statement about “a technical issue.” Tell them what the agent said and what is correct. Whether to honor what the agent promised by mistake is a business decision, made deliberately by the right person; what can’t happen is the agent making it by default.
Do you have to notify anyone if only one customer was affected?
Yes: that customer and whoever decides at the company. What changes is everyone else: with a case confirmed as isolated after scoping the damage, an individual correction is enough. But “isolated” is a conclusion you prove with the search above, not an assumption you start from.
Only then: diagnose without reopening the problem
With the damage stopped and the affected customers identified, diagnosis no longer runs against the clock. Look for the cause in this order, cheapest to most expensive: the instruction (was it ambiguous or contradictory?), the information source (an outdated document, a page that changed?), the tool (did it return a wrong value?) and, last, the model. It’s almost never the model, and it’s the hypothesis everyone jumps to first.
One risk remains: fixing and re-enabling on the same day. The fix is one more change and needs testing before it touches production, with the conversation that failed as the first test. Then autonomy is restored step by step, not all at once.
The first two hours, summarized
| Window | What you do | What you don’t do |
|---|---|---|
| Minutes 0-15 | Contain: lower autonomy, remove the tool or hand off to a person | Look for the cause |
| Minutes 15-60 | Scope: count exposed, affected and money commitments | Assume it was an isolated case |
| Minutes 60-90 | Notify: decision-maker first, then customers with commitments | Wait for the root cause before communicating |
| Minutes 90-120 | Start diagnosing with the damage stopped | Re-enable without testing the fix |
What has to be ready before it happens
Two hours is only possible if the pieces existed beforehand: a named kill switch and someone authorized to use it, a searchable conversation log, a list of who to notify and a draft message. If this is the first time you’re thinking about it at six in the evening, it will take more than two hours. The on-call rota, the severities and the one-page runbook that organize all this are in who responds when an automation goes down, and if you’d rather not build or sustain it yourself, it’s what we do with AI agent incident management.
The line to take away: an agent getting something wrong is normal; an agent that keeps getting it wrong while you work out why is not.