Skip to content
Implementa.
InfrastructureAI Agents··8 min

Your AI agent doesn't have a performance problem: it has an auditable-evidence problem

When something goes wrong with an agent, the question you get asked is not whether it was working well: it is what you can show. The auditable evidence for an AI agent is not the dashboard and not the log, and the market sells three different layers as if they were one. How to tell them apart, and what to demand in writing before you sign.

Senior AI Infrastructure Implementer

AI Infrastructure Pod

The thesis in one line: the day an agent does something that has to be explained, nobody will ask you about its accuracy rate. They will ask what happened in that one case, and you will have exactly what you decided to keep six months earlier.

The scene repeats itself. A customer complains, an auditor asks, a boss wants to know. Somebody opens the agent dashboard and shows a 96 % accuracy rate. And the answer from the other side is always the same: "great, but I am asking you about Marta's case, the one from 14 March." The 96 % does not answer that question. It does not answer it because it measures something else.

That mismatch has a name and it is not a technical fault: somebody bought one layer believing they were buying three.

Performance and auditable evidence are not the same question

Performance is an aggregate, present-tense question: is the system doing well, what does it cost, where is it degrading? You answer it with a curve and it helps you operate. Auditable evidence is a singular, past-tense question: why did this decision come out this way, for this customer, on this day? You answer it with a case file and it helps you defend yourself.

They are different questions with different answers, and the trap is that the first one is visible and the second one is not. A dashboard gets shown in a demo; a reconstructed case file only shows up when somebody needs it, which is always too late to build it.

The three layers the market sells as one

When a vendor says "our agents are auditable," they could mean three things that do not substitute for one another. Separate them before you compare offers, because nearly every vendor covers one well and the other two in passing.

LayerWhen it actsWhat question it answersWhat it does NOT do
GuardrailsBefore: preventiveCan it do this?Leaves no trace of why what did happen, happened
ObservabilityDuring: operational, aggregateIs it doing well today?Does not reconstruct a specific case
Audit trailAfter: probative, singularWhy did THIS case come out this way?Prevents nothing; it only proves

The first layer is a design decision: what the agent can touch and how far it goes on its own. It gets settled before you connect anything — it is covered in what permissions to give an AI agent and in the levels of agent autonomy — and its value is that certain things never happen. But a guardrail that works is invisible: it does not produce proof, it produces an absence of incident.

The second is the one nearly everybody buys first, and rightly so: without it an agent breaks in silence. That is monitoring AI in production, and it answers in aggregate. Its unit is the metric, not the case.

The third is the one almost nobody has built, because it demands uncomfortable decisions: what gets kept, under which identifier, and for how long. That is AI decision traceability, and it is the only one that helps you on the day of the complaint. All three are necessary. But only one answers the question you are going to be asked.

Why your log is not proof

Here comes the reasonable objection: "we log everything." That is nearly always true, and nearly always useless. An operational log and a probative trail are separated by two properties that are not technical but governance ones.

  1. Integrity. A probative trail has to be able to show it has not been touched since it was written. If anyone with access to the system can edit or delete it without leaving a mark, it proves nothing: it is one party's version of events.
  2. A retention period somebody chose. A log rotates. It gets deleted after thirty days because of storage, or because nobody thought about it. A trail has a deliberately chosen period, written down, aligned with how long complaints actually take to surface in your business — which is rarely thirty days.

And here is the clash worth anticipating: keeping data longer pushes against minimising personal data. You do not solve that with a tool, you solve it with a written policy per type of system: what gets kept whole, what gets kept pseudonymised, what gets kept only as a reference. That decision belongs to the business and to legal, not to the vendor. If nobody has taken it, the default policy is whatever the person who set the default rotation chose.

What European law already requires, and of whom

It is worth being precise, because this gets exaggerated in both directions. The European AI Act does not require you to log everything any AI does. It does require it for systems classified as high-risk, and there the obligation runs through two places.

Article 12 requires the system to technically allow the automatic recording of events over its lifetime, with logging capabilities that serve to identify risk situations and to monitor its operation. That falls on whoever builds it.

The part almost nobody has read is Article 26(6), and that one falls on whoever deploys: a deployer of a high-risk system must keep the logs the system generates automatically, to the extent those logs are under their control, for a period appropriate to the intended purpose and of at least six months, unless other Union or national law says otherwise, in particular data protection law. Source: Article 26: Obligations of Deployers of High-Risk AI Systems, EU Artificial Intelligence Act, accessed 11 September 2026.

Two qualifications so as not to overshoot. First: the application date for these obligations has moved, and the same source puts Article 26 in force on 2 December 2027 for high-risk systems under Annex III and 2 August 2028 for those under Annex I, per Article 113. Second, and more important: if your agent is not high-risk, none of this binds you. But the phrase "to the extent such logs are under their control" already tells you where the industry standard is heading, and a large customer or an insurer can ask you for it long before a regulator does.

Three things to demand in writing before you sign

You do not need a twenty-page annex. Three questions answered in writing tell you whether you have evidence or a dashboard.

  1. Which fields get recorded, exactly. Not "we log activity." The list: end-to-end case identifier, version of the model and of the instructions that applied, what was retrieved, which tools were called, which policy was applied and who reviewed it. If the answer is an adjective instead of a list, there is no trail.
  2. Where they live and for how long. In which system, under whose control, with what retention period and what happens when it expires. A default period nobody chose is a decision taken by omission, and you will discover it on the day you need it.
  3. Who can export them and in what format. This is the one most often forgotten and the most expensive: the day you switch vendors, do you take the trail with you or does it stay on their platform? If it is not written down, assume it stays.

None of the three is a hard technical question. They are questions a serious vendor answers in an email, and the ones that do not get answered in an email tell you more than the ones that do.

What to do this week

Pick an agent that is already working and run the six-month test on a real case. That is all it takes: in half an hour you will know which of the three layers you are in and which one you are missing. If you cannot reconstruct the case, the list above is your pending conversation with whoever sold it to you.

And if, while reconstructing it, you find you do not even know which version of the instructions was running that day, that is an earlier problem and it gets fixed first: the instructions of an AI agent are not a prompt, they are a document with an owner, a version and a date. Without that version written down, the most complete trail in the world leaves you halfway.

The honest summary: performance tells you whether the system is worth it; auditable evidence tells you whether you can stand behind it in front of somebody. The first is easy to buy. The second gets decided beforehand and cannot be built backwards.

Shall we get it shipping?

If this resonated, 30-minute conversation with no commitment. We tell you what fits, what doesn't and the approximate price.

See cases
Your AI agent doesn't have a performance problem: it has an auditable-evidence problem · Implementa