A pilot that classifies customer requests is validated in January. It works: real cases, human review, committee sign-off, into production. In June the complaints start: requests routed to the wrong department, replies that miss the point. The team does the sensible thing and checks the prompt, the model and the configuration. Everything is exactly as it was in January. “Nobody changed anything,” they conclude, and the conversation slides toward “AI just isn’t reliable.”
The thesis in one line: when an agent gets worse and nobody touched the code, its inputs changed, and January’s test only proved it worked for January’s people, January’s forms and January’s calendar. “Nobody changed anything” isn’t the system’s defense; it’s the diagnosis.
Why an AI agent gets worse over time when the code is the same
Because an agent isn’t a program that always does the same thing: it’s a function of what it receives, and what it receives changes. A validation measures behavior on a specific sample of inputs. If reality drifts away from that sample, the number you saw in the test doesn’t travel with it. It’s the same reason accuracy alone isn’t enough in production: accuracy isn’t a property of the model, it’s a property of the system plus its data, and data moves.
Four ways inputs change without warning
- The calendar. Campaigns, quarter-end, peak season, the holiday month. An agent validated on a quiet month’s traffic meets a different kind of request in another month: more urgent, more ambiguous.
- New people who write differently. A new channel, a customer from another country, a segment that never used to show up. Longer messages, a second language mixed in, different abbreviations. The test sample didn’t contain them.
- An upstream system that changes. A form that adds an optional field, a CRM that changes a date format, a vendor that renames a column. Nobody tells the agent because nobody knows the dependency exists.
- The business changes and the prompt doesn’t. A new product, new prices, a different returns policy. The instructions still describe the company as it was in January.
Input drift and model drift aren’t fixed the same way
Keep them apart, because the first question in any diagnosis is which of the two you have. Model drift comes from the vendor: they update or retire a version and the same prompt behaves differently. You contain it with a pinned version and a battery of cases that runs on every change, which we cover in the guide on how to switch models without breaking your automations.
Input drift happens even if you pin everything. You can freeze the model, the prompt and the infrastructure for a year: if the world feeding it changes, the agent gets worse anyway. That’s why the usual reaction, trying another model, tends to be a mistake, as we argue in switching models won’t fix your process: you move the problem somewhere else and pay for a migration on top.
A one-minute question separates the two: did the system change, or did what goes into it change? If the model, the prompt and the integrations are as they were and performance is dropping, start with the inputs.
How to catch input drift before your customer does
A signal that arrives through complaints arrives late: by then you’re weeks into the errors. The four checks that get ahead of it are cheap, but they have to exist before you need them:
- Keep the snapshot from validation. The sample of inputs the agent was approved on is your reference. If you didn’t keep it, there’s nothing to compare against.
- Compare what comes in with that snapshot every week. You don’t need fancy statistics: message length, language, split by category, empty fields, share of cases the agent has never seen. A sharp change in any of them is the alarm.
- Watch exceptions by input type. The rate of cases the agent hands to a person, or that a person corrects, rises before the complaints do, and broken down by type it tells you where reality is moving.
- Re-run your evals with recent cases. A fixed January battery only tells you what you already knew. Add new real cases every month, labeled by a person, as described in evals: how to know your AI actually works.
The first step, by the way, only exists if you took a baseline. It’s the same principle as measuring the performance of your automations: the number you didn’t take before can’t be taken afterwards.
What to do once the input has already moved
There are three responses, and the choice depends on how far things have moved and what an error costs:
- Update the reference. Fold the new cases into the validation and retune instructions and examples. It’s the answer when the change is stable and the agent can learn to cover it.
- Widen the scope with its own path. If a new, distinct kind of input has appeared (another language, another channel), treat it as a new case with its own validation, not a variant of the old one.
- Lower the autonomy in the meantime. What the agent no longer recognizes goes to a person, and the agent moves from deciding to proposing. It’s the mechanism behind the levels of agent autonomy: dropping one step is reversible; switching it off isn’t always.
None of this works without an owner. Nobody in particular catches input drift because it isn’t anybody in particular’s job: not the person who built the agent, who is on another project by now, and not the people who use it. You need a named person, a cadence (weekly, to start) and the authority to change the agent when the data says so. The rest of the plan is in the guide on maintaining AI automations. And if you’d rather not build or sustain it yourself, that’s what we do when we monitor AI in production: instrument the inputs, the exceptions and the reference cases, and act when they move.
How to tell if it’s happening to you right now
- Nobody can show you the sample of inputs the agent was approved on.
- The last time new cases were added to the validation was before launch.
- Exceptions are counted in total, not by input type.
- Ask who watches all this and the answer is “the team,” not a person.
If two of the four are true, your agent has probably already moved and you’re just missing the instrument to see it.
The line to take away: an agent doesn’t get worse because it gets old; it gets worse because the world keeps changing and nobody checked again that they were still speaking the same language.