The thesis in one line: there are two kinds of AI KPIs and only one pays the invoice. There are the ones that shine in the demo —users who signed up, prompts fired, the committee’s applause— and there are the ones that move the real work: cases resolved without anyone touching them, hours that stop being spent, errors that drop. The first group is theatre; the second is operations. And the rule to tell them apart is brutally simple: if you don’t measure it every week against a baseline, it’s theatre.
The AI KPIs in the enterprise that matter (and the ones that are pure theatre)
When a company shows off its AI project, it almost always shows the comfortable metric: how many people used it, how many queries got fired, how happy the team was after the training. Those are numbers that climb on their own over time and say nothing about whether the work gets done better. The uncomfortable number —the one that decides whether the project pays off— is a different one: which concrete task used to come out one way and now comes out better, faster or with fewer errors. The difference between a project that pays and an expensive demo lives between those two groups of metrics.
This isn’t a contrarian consultant’s take. The MIT NANDA report «The GenAI Divide: State of AI in Business 2025» analysed 300 deployments and found that 95% of AI pilots had no measurable impact on the P&L, even though over 80% of companies had already tried tools like ChatGPT or Copilot (full report, 2025). Translated: almost everyone measures adoption and almost nobody measures outcome. That gap is the theatre.
The theatre rule: no baseline and no weekly cadence, it’s smoke
A metric with no baseline isn’t a metric: it’s a decoration. «We resolve 200 tickets a month with AI» means nothing if you don’t know how many you resolved before, how long it took, and how many came back wrong. The number that counts isn’t the absolute one, it’s the delta against how it was done without the tool. That’s why day one of any AI project taken seriously goes to measuring the before —the baseline—, not to running the demo.
And a metric you only look at in the quarterly meeting doesn’t count either. Adoption is decided in weeks: if a team tries the tool, isn’t convinced and goes back to the manual way, by the time the committee three months out arrives the project is already dead and nobody noticed in time. A weekly cadence against a baseline is what separates a dashboard from a postcard.
The four operational metrics that do count
If you’re going to measure one thing, measure the work. These four are the ones that tell you whether the AI is doing something or just switched on:
- Cases resolved without human intervention. Out of every hundred tasks that come in —tickets, emails, invoices, queries—, how many come out resolved and right without a person having to touch them. It’s the queen metric of any automation: it measures work done, not activity. If you ship an AI accounting agent, this is the number that decides whether it saves you a headcount or gives you one more thing to review.
- Hours saved against the baseline. Not «hours the tool claims it saves», but the real subtraction: what the process took before minus what it takes now, times its frequency. It’s the metric that translates straight into money and the one the committee understands without a translator.
- Error rate. An AI that does triple the work with double the errors saves you nothing: it shifts the cost of doing to the cost of reviewing and fixing. Measuring the percentage of outputs that have to be redone is what stops the self-delusion of «it’s so fast» while someone cleans up behind it.
- Real per-person usage, not active licenses. How many people actually use it each week in their work, role by role. It’s the metric that exposes the adoption gap —people with access who never touch it— before it turns into a dead project. You pay for the license; usage is earned, and you only see it if you measure it per person.
These aren’t GEO KPIs —don’t confuse them—
There’s one place where this distinction gets blurred on purpose: AI visibility. Whether your brand shows up in ChatGPT or Google AI is also measured, but that’s a different discipline with its own dashboard —mention rate, citation, share of voice—, and we break it down in the GEO KPIs. That dashboard measures whether you’re seen; the one in this article measures whether your operation works. They’re two different boards, and presenting them as if they were the same is the favourite trick of anyone who wants to show the pretty number. If what you need is to take visibility to leadership without smoke, that has its own format too: the GEO report for leadership.
How to build the dashboard without smoke
- Measure the baseline before you switch anything on. Day one is for timing and counting the current process, not for the demo. Without that before, any after is an anecdote.
- Pick one concrete, visible task, not «AI in the company». Cases resolved for THAT process, hours for THAT process. A dashboard for one task that pays convinces more than ten generic indicators.
- Put the four operational metrics on a weekly cadence, with an owner. Someone looks at the delta each week and acts when it drops, just like any other number in the business.
- Keep the operations board separate from the visibility board. GEO in one place, operations in another. Mixing them only serves to hide the one going badly behind the one going well.
None of this is exotic: it’s treating AI like any other part of the business you want to pay off —with a baseline, a delta and someone watching—. If you’d rather that dashboard be built and sustained by whoever also fixes the process underneath, operations automation is exactly that: we don’t ship you a vanity-metrics panel, we leave you measuring work done.