Skip to content
Implementa.
AutomationPlaybook··6 min

What to do when an AI agent fails with a customer: the first two hours

What to do when an AI agent fails with a customer: contain, scope the damage, notify, and only then diagnose. Almost everyone starts at the end and prolongs the harm. The order of the first two hours, the question nobody has prepared for, and what has to be ready before it happens.

Senior AI Operations Implementer

AI Operations Pod

On a Tuesday afternoon, your support agent tells a customer they can return a custom order after the deadline. They can’t: that policy doesn’t exist. Someone on the team spots it while reading the conversation, and the reaction is nearly universal: open the logs and work out why it said that. It’s the most natural move and the one that prolongs the damage the most.

The thesis in one line: when an AI agent gets something wrong with a customer, the right order is contain, scope the damage, notify, and only then diagnose. Almost everyone starts at the end, and while they diagnose, the agent keeps answering.

What to do when an AI agent fails: contain before you diagnose

Diagnosing takes hours; containing takes minutes. While you hunt for the cause, the agent is still in production with the same instructions, the same data and the same tools that produced the mistake, so every new conversation is another chance to repeat it. Containing isn’t fixing: it’s shrinking what the agent can do until you know what happened. And it has to be reversible, because you’ll do it in a hurry and without a diagnosis.

The options, from least to most drastic:

  • Lower its autonomy. Switch it to “drafts, a person sends.” The service stays up and the new damage stops. It only works if you defined a scale of agent autonomy levels beforehand.
  • Take away the tool that caused the damage. If it promised a refund, remove its refund permission; if it wrote to the CRM, make it read-only. This is the practical payoff of having thought through an AI agent’s permissions.
  • Hand the affected topics to a person, with a fixed, honest message: “A member of our team will take it from here.”
  • Switch it off entirely. A last resort, but a legitimate one. It requires a named kill switch and someone authorized to press it without asking three people for permission.

The rule: press the smallest button that stops the damage. A switched-off agent is also an incident, just a visible one.

How many other customers got the same bad answer

This is the question almost nobody has prepared for, and it decides how serious the incident really is. An isolated error and a pattern are handled differently, and from one customer’s conversation you can’t tell which you have. Scoping the damage means answering three questions with data:

  • How many were exposed: conversations where the agent said the same thing or something equivalent, from the first time it could have happened until the moment of containment.
  • How many acted on it: those who requested the return, accepted a deadline or made a decision based on that answer.
  • How many hold a commitment involving money or a deadline: the only ones who need an individual reply today.

None of that is possible unless conversations are stored and searchable. Log retention is usually decided on storage cost or privacy grounds, without this moment in mind; decide it knowing it’s what lets you count the people affected, as explained in how long your agent’s log should last. A good starting rule: the review window opens at the agent’s last change (instruction, data source or model), not at the first complaint.

Notify: who, when and with what message

Notifying comes before diagnosing because the affected customer will find out anyway, and it’s better they hear it from you. You don’t need the cause to say what matters: what happened, what is valid and what isn’t, and what you’re going to do. Three recipients, three moments:

WhoWhenWhat you tell them
Whoever decides at the company (management or service owner)As soon as you confirm an error with impactWhat the agent said, since when, to how many people, and what you’ve contained
Customers with a commitment involving money or a deadlineAs soon as possible, individually, once the scope is knownThat the information was wrong, what the right answer is and how it gets resolved
Everyone else exposedOnly if they acted on the answer or the topic is sensitiveA short correction, no drama

Two things that don’t work: waiting for the root cause before writing to the customer, and sending a statement about “a technical issue.” Tell them what the agent said and what is correct. Whether to honor what the agent promised by mistake is a business decision, made deliberately by the right person; what can’t happen is the agent making it by default.

Do you have to notify anyone if only one customer was affected?

Yes: that customer and whoever decides at the company. What changes is everyone else: with a case confirmed as isolated after scoping the damage, an individual correction is enough. But “isolated” is a conclusion you prove with the search above, not an assumption you start from.

Only then: diagnose without reopening the problem

With the damage stopped and the affected customers identified, diagnosis no longer runs against the clock. Look for the cause in this order, cheapest to most expensive: the instruction (was it ambiguous or contradictory?), the information source (an outdated document, a page that changed?), the tool (did it return a wrong value?) and, last, the model. It’s almost never the model, and it’s the hypothesis everyone jumps to first.

One risk remains: fixing and re-enabling on the same day. The fix is one more change and needs testing before it touches production, with the conversation that failed as the first test. Then autonomy is restored step by step, not all at once.

The first two hours, summarized

WindowWhat you doWhat you don’t do
Minutes 0-15Contain: lower autonomy, remove the tool or hand off to a personLook for the cause
Minutes 15-60Scope: count exposed, affected and money commitmentsAssume it was an isolated case
Minutes 60-90Notify: decision-maker first, then customers with commitmentsWait for the root cause before communicating
Minutes 90-120Start diagnosing with the damage stoppedRe-enable without testing the fix

What has to be ready before it happens

Two hours is only possible if the pieces existed beforehand: a named kill switch and someone authorized to use it, a searchable conversation log, a list of who to notify and a draft message. If this is the first time you’re thinking about it at six in the evening, it will take more than two hours. The on-call rota, the severities and the one-page runbook that organize all this are in who responds when an automation goes down, and if you’d rather not build or sustain it yourself, it’s what we do with AI agent incident management.

The line to take away: an agent getting something wrong is normal; an agent that keeps getting it wrong while you work out why is not.

Shall we get it shipping?

If this resonated, 30-minute conversation with no commitment. We tell you what fits, what doesn't and the approximate price.

See cases
What to do when an AI agent fails with a customer: the first two hours · Implementa