Skip to content
Implementa.
AI AgentsInfrastructure··12 min

Who reviews the reviewer: the AI supervisor agent that audits your agent

Once an agent produces hundreds of outputs a day, full human review stops being review and becomes a rubber stamp. What an AI supervisor agent is, what it checks, where it sits, why it cannot share a model or a context with the agent it watches, and the three cases where you still need a person.

Senior AI Operations Implementer

AI Operations Pod

The thesis in one line: past a certain volume, reviewing everything by hand stops being review and turns into a rubber stamp. If you want to keep watching what your agent does, the second line has to be automatic — and it has to be a different agent, not the same one asked a different question.

Your review process was designed when the agent did twelve things a day. One person read all twelve, fixed two, and learned something along the way. It worked. Eight months later the agent does four hundred and the person is still the same person, with the same calendar and the same meetings. Nobody changed the procedure: they still mark every item "reviewed". What changed is what that word means.

There is no bad faith here, just arithmetic. Four hundred outputs at two minutes of honest reading is over thirteen hours. Nobody has thirteen hours, so what happens is what always happens: read the first line, check the formatting holds, approve. The control still exists on the org chart and has stopped existing in reality.

What an AI supervisor agent is and what exactly it checks

An AI supervisor agent is a second agent whose only job is to look at what the first one produced before the action ships, check it against a written policy and return a verdict: pass, block, or escalate to a human. It does not generate. It does not rewrite. It does not opine on whether this could read better. It decides one thing: whether this can go out.

What it checks is deliberately boring, which is exactly why it works. Four closed questions with verifiable answers:

  • That what it claims exists. Every figure, name, date or reference in the output has to be pointable at a source that was actually retrieved. If a number shows up that is in no retrieved document, it is not a number: it is an invention wearing the shape of a number.
  • That the action is inside what is allowed. Which tool it called, on which record, for what amount, on whose behalf. The policy says what it may touch; the supervisor checks that what it is about to touch is on that list and not another.
  • That the case looks like something the policy anticipated. Language, customer type, jurisdiction, deal size. A case outside the anticipated set is not necessarily wrong: it is uncovered, which is a different thing and gets handled differently.
  • That no mandatory step got skipped. If the procedure says you check the history before answering a complaint, the supervisor checks whether it did. It is the dumbest of the four and the one that trips most often.

Notice what is NOT on that list: "that the answer is good". Average system quality gets measured with evaluations over saved cases, before you ship, and that is a separate discipline we cover in running evals so you know whether your AI actually works. A supervisor does not measure averages. It looks at this case, right now, and decides whether it ships.

Why can the agent not just review itself?

Because the bias is measured and it has a name. In Self-Preference Bias in LLM-as-a-Judge (Wataoka, Takahashi and Ri) the authors propose a metric to quantify it and find that GPT-4 scores its own generations significantly higher; their hypothesis is that models favour what feels familiar to them, measured as lower perplexity. Source: Self-Preference Bias in LLM-as-a-Judge, arXiv, 29 October 2024 (revised June 2025). This is a technical result, not a market number: it holds the same in Chicago as in Madrid.

Translated into your process: asking an agent to audit itself is asking it to doubt the sentence that felt like the most natural thing in the world to it two seconds ago. It will, in some share of cases, with great confidence and very little use.

The shared-model problem is the one everyone quotes, but it is not the worst one. The worst one is shared context. If the error is born from a badly retrieved document, a reviewer reading that exact same badly retrieved document will confirm the error enthusiastically and mark it verified. Independence here is not a moral principle: it is the mathematical condition for the second look to add information the first one did not have.

In parallel, or between the agent and the destination system?

This is the architecture call people argue about most and defend worst, because it gets framed as a technical question when it is a question about which error you can afford. There are only two positions.

In parallel (observer)In the middle (gatekeeper)
What happens if it failsThe action already shipped; you find out laterThe action does not ship; the process stops
Latency addedNone: it runs afterwardsOne extra check on every operation
Which error it preventsNone; it documents them and lets you fixThe irreversible ones, which are the ones that matter
Cost of a false positiveLow: one alert too manyHigh: you just blocked good work
Where it makes senseReversible actions, high volume, break-in phaseMoney, third parties, regulated data

Our rule is short: the supervisor sits in the middle only where the action is irreversible or leaves the company. Everywhere else it sits alongside. And you start in parallel even if you plan to end up in the middle, because for the first two weeks its job is not to block: it is to show you how often it would have blocked and on what grounds. Making it a gatekeeper on day one is installing a brake with an unknown false-positive rate.

This is not a technology decision, it is an autonomy decision, and it rides on the one you already made when you set the autonomy levels of an AI agent: the looser the first line runs, the better the case for putting the second one in its path.

Pass, block or escalate: there is no fourth exit

A supervisor with more than three verdicts is a supervisor with opinions. Here are the three, and what each one forces you to build:

  1. Pass. The action ships and it is on record that it went through the supervisor, against which policy version, with which result. A "pass" with no trail is worth nothing: six months out, the distance between "it was reviewed" and "we can prove it was reviewed" is the whole distance. That record is the one we mean by traceability of AI decisions.
  2. Block. The action does not ship. It goes back to the first agent with the specific reason — which check failed, on which data — and gets one retry. Fail again and it goes up. A silent block nobody counts is a leak: work stops, the customer waits, and nobody knows until they call.
  3. Escalate. It goes to a person with the case already assembled: what the agent proposed, which check failed, what the policy says, and what it suggests doing. Escalating is not forwarding. It is arriving with the work done so the decision costs a minute instead of twenty.

The metric that matters for a supervisor is not how many actions it blocked: it is how many it escalated and how long those took to clear. Escalate 40 % and you have not built a second line, you have relocated the bottleneck. Escalate 0.5 % and either your system is flawless or your policy is so loose it checks nothing. The healthy range looks a lot more like the first digit than the second, and you tune it by reading escalations one by one for the first few weeks.

The policy it compares against does not appear out of nowhere: it is the one you already wrote when you defined the instructions of an AI agent, with one non-negotiable condition — written separately, versioned separately. If the supervisor inherits the same document, it inherits its blind spots too, and an inherited blind spot is one that never gets caught.

The three cases where you still need a person

An automatic supervisor does not remove the human. It changes their diet. Instead of skimming four hundred items, they read twenty properly. And there are three situations where those twenty are not up for negotiation.

  1. The amount crosses a threshold. Not because the agent is more likely to be wrong at $40,000 than at $400, but because the cost of being wrong is no longer absorbed by the process: it is absorbed by the P&L. The threshold is a business decision, written in currency, revisited quarterly. Where that person sits without becoming the brake on everything else is the human-in-the-loop design of an automation.
  2. The responsibility is legally non-delegable. There is no design freedom here, because the law names people. The EU AI Act, in article 14, requires for the biometric identification systems in Annex III point 1(a) that no action or decision be taken unless the identification has been separately verified and confirmed by at least two natural persons with the necessary competence, training and authority. Source: Article 14 — Human oversight, Regulation (EU) 2024/1689. In Spain, the data protection authority adds the four-eyes principle — double verification by two different people — for automated processes with major impact on rights, in its guidance on agentic AI (version 1.2, February 2026). Two natural persons means two natural persons: no supervisor, however good, counts as one.
  3. The case looks like nothing you anticipated. This is the uncomfortable one, because you cannot write it into a rule in advance. A supervisor only knows how to check against what the policy contemplated; faced with something genuinely new, the only correct move is to admit it is not covered and hand it over. That ability to say "this is not mine to decide" is the same one the first line needs, and it is worked out in what an AI agent does when it does not know. The difference is that a supervisor has to fail upward: when in doubt, escalate.

And where you need the person, you need the function, not the favour: someone with written criteria, a queue, a committed response time and a measure of their own accuracy. That is human oversight of AI at scale, and it does not compete with the automatic supervisor — it is what makes it usable. The machine filters the boring 95 % so human judgment lands whole on the 5 % that deserves it.

What got stopped, what it was worth, what evidence remains

This is where the conversation usually goes wrong, because the supervisor gets pitched as a security matter and dies in the security committee. It is a business matter. Two numbers to size it, both from Gartner and both global in scope: more than 40 % of agentic AI projects will be cancelled by the end of 2027 on rising costs, unclear business value or inadequate risk controls (Gartner, 25 June 2025); and guardian agent technologies will account for at least 10 % to 15 % of the agentic AI market by 2030 (Gartner, 11 June 2025).

Read them together and the comedy shows: the industry is simultaneously killing projects for lack of control and standing up an entire category to sell you that control as a product. What neither number tells you is how much gets stopped in your house. Gartner does not have that number. You do, in your own history, and it is the only one worth deciding on.

Which is why a supervisor dashboard is not technical. It is five business numbers: how many actions were stopped this week, of what type, what they were worth, how many turned out to be false alarms, and how long each escalation took to clear. With those five you can have the real conversation about giving it more autonomy or less. Without them you can only argue by vibes, and in that argument the loudest voice always wins.

One warning about what this is not: it is not security. A supervisor checks against a business policy; it does not defend you from someone deliberately trying to trick the system. That is a different problem with different tools, a different team and a different budget, and we treat it separately in the AI agent risks you do not see in the demo. Confusing the two is the most expensive way to end up with neither.

What I would do on Monday

One question and one experiment. The question, with numbers and no decoration: how many outputs does your agent produce a day, and how many does one person actually read, end to end? If multiplying that second figure by two minutes exceeds the free time that person has, you already know your review is a rubber stamp and you can stop reading here.

The experiment fits in two weeks and never touches production. Take the last five hundred outputs your agent produced, exactly as stored. Put a second agent — different model, different prompt, retrieving sources on its own — to check just three things: that the figures exist in the source, that the action was permitted, and that no mandatory step got skipped. Count how many it would have blocked and go read them one by one. If two or three of them should never have shipped, your business case is already written, and you did not write it: your own history did.

The question in the title has an unheroic answer. You review the reviewer, yes — once a week, on a sample. The rest of the time it gets reviewed by a different machine, with a different head and different sources, against a policy you wrote on a day when you were calm. It is not elegant. It is the only thing that holds when volume multiplies by forty and headcount does not.

Shall we get it shipping?

If this resonated, 30-minute conversation with no commitment. We tell you what fits, what doesn't and the approximate price.

See cases
Who reviews the reviewer: the AI supervisor agent that audits your agent · Implementa