Solution · AI Operations
Training your team to run AI isn't a ChatGPT course. It's building them a function.
With AI in production there are three roads: hire the team, have someone else run it, or have the people you already employ run it. The third is the fastest to start and the worst executed, because it gets mistaken for a one-day workshop. We design the function, train your people on your real systems, and stay until they hold it up on their own.
The problem
AI is already in production. Running it isn't anybody's actual job.
- Production AI is held up by two people who never signed up for the work: they do it on top of their own job, with no shift, no written criteria and no mention of it in their review.
- When an output looks off, everyone decides alone whether to fix it, ignore it or escalate it. With no shared criteria, the same case gets three different answers depending on who's around.
- Exceptions — what the system can't resolve — pile up in a chat channel instead of a queue with an owner and a response time, and get handled in order of who pushes hardest.
- You ran an AI training months ago. Your people can write prompts and still don't know what to do when the system drifts on a Tuesday morning.
- Nobody looks at consumption until the invoice lands, and by then the conversation is an accounting one instead of an operational one.
- Every time the provider moves the model there's a strange week: quality shifts, nobody knows whether it's the model, the prompt or the data, and the call gets made on gut feel.
Cost of staying the same
You paid for the system and not for the operation. While the AI gets it right, that gap is invisible; the day it drifts, a customer finds it and you pay for it in full. And there's a second cost, slower and more expensive: since running it isn't anybody's job, improving it isn't either — so the system sits exactly where you left it on launch day while the market keeps moving. The other two exits aren't free either: opening a role drops you into a hiring process for a title that's badly inflated, and handing the whole function outside leaves your operational judgment in somebody else's house.
The solution
We design the function, train your people to run it, and let go in phases
- 1We start with the map, not the classroom: which AI systems are live, what decisions they make, what happens when they fail, and who is covering that gap today without it showing up anywhere.
- 2We define the function and split it: who reviews outputs and at what sampling rate, who works the exception queue, who escalates and to whom, who maintains the quality checklists and who reviews consumption. With names, shifts and response times — not a theoretical org chart.
- 3We train on your system, with your cases and your real failures. No textbook exercises: your people practice on the outputs your AI produced last week, including the ones that went wrong, and learn to tell a model failure from a data failure.
- 4We put the criteria in writing: what gets fixed, what gets accepted, what gets escalated and what stops everything. That document is what turns one person's judgment into a team standard, and it's also what stops a handover from resetting the learning.
- 5We add the two pieces that separate running it from watching it: versioning the prompts — so a change is tested before it ships and can be reverted — and reading costs in operation, so the team catches the odd spend before finance does.
- 6We stay in production and let go in phases: first we operate with them watching, then they operate with us watching, then alone — with the function's metrics on their own board and a standing review with us.
What changes
What you stop losing
Running the AI stops being a favor and becomes a job with an owner, a shift and written criteria, inside people who already know your operation, your customers and your exceptions.
Mechanism
The criteria stop living in two people's heads: they get written down, trained on real cases, and survive holidays, sick leave and handovers.
Mechanism
The system starts improving again: once somebody has the mandate to touch it and a safe way to test and revert a change, the AI stops freezing at the point where it was deployed.
Mechanism
What we measure: % of outputs reviewed at the agreed sampling rate, exception queue response time, % of blockers escalated down the right path, quality checklists run on time, and consumption anomalies caught by the team before finance.
What we measure
Spec sheet
- Work it removes
- letting running the AI be an invisible favor from two people: no shift, no written criteria, no exception queue and nobody watching quality or cost
- Typical setup
- 6–10 weeks
- Input
- your AI systems already in production, the team you have today, and everything done by hand when something drifts
- Output
- an operating function split across your people — review, exceptions, escalation, quality checklists, prompt versioning and cost reading — with written criteria, its own metrics, and the ability to hold it up without us
- Works with
- OpenAIAnthropicAzure OpenAIGoogle Vertex AILangSmithLangfuseJiraLinearServiceNowNotionMicrosoft 365Google Workspace
- Can connect to
- Tu cola de incidencias y tu proceso de guardia actualesTus registros de producción, de donde salen los casos reales de formaciónTu gestión de prompts y tu panel de consumoTu marco de gobierno de IA y tu inventario de sistemas
- What we measure
- % of outputs reviewed at the agreed sampling rateexception queue response time% of blockers escalated down the right pathquality checklists run on timeconsumption anomalies caught by the team before finance
- Good fit for
- companies with AI already in production and a team that knows the operation — CIO, COO or head of AI — who would rather build the capability inside than open a role or hand their operational judgment to a third party
- Not a fit for
- anyone with nothing in production yet — there's no function to run, there's a system to build first — or anyone after general AI training for the whole workforce: that's adoption, and it's a different service
Frequently asked questions
They're the three corners of the same triangle, and the choice isn't ideological — it's time, cost and control. Building the team means hiring the missing roles and giving them a mandate: that's the road if AI is going to be core to your business and you can wait out a hiring process. Outsourcing means someone else runs the function on your stack: that's the road if you need it working now and would rather pay for capability that already exists. Training is the third: the people who already know your operation, your customers and your exceptions learn to run the system, and the judgment stays in-house. It's the cheapest in money and the most expensive in leadership attention, because it requires someone inside to genuinely change jobs, not attend a workshop. Plenty of companies land on a mix: we train the team and run it with them while the capability matures.
No, and the difference shows on day one. A course teaches you to use the tool: write better prompts, learn the features, get more out of the chat. This teaches you to hold up a system that's already making decisions with customers on the other end: how outputs get sampled, what to do with one that smells wrong, when to stop the whole thing, how to work an exception queue with a response time, how to test a change before shipping it, how to read consumption. And it doesn't end with a certificate: it ends with a function split across named people, written criteria and metrics your team looks at daily. If what you need is the first thing, that's adoption and it's a different service — we do it too, we just don't call it the same.
Most of the function isn't engineering — it's operational judgment. Telling a good output from a bad one in your business, knowing which case is worth stopping for and which gets fixed and moves on, understanding which exception is a one-off and which is a pattern that needs fixing upstream. That's done better by someone who's watched your processes for years than by someone who just arrived with a fashionable title. The two technical pieces — versioning prompts and reading costs — get taught concretely, on your tools, and don't require writing code. Where a technical profile is genuinely needed, we say so and write it into the split instead of pretending anyone can cover it.
That's why the deliverable isn't the trained person: it's the written function. The criteria, the split, the response times, the checklists and the metrics are documented and live, so a replacement arrives to read and practice rather than to rebuild. We also always split across more than one person and define a backup for each piece — a function that depends on one name isn't a function, it's a dependency. And if the gap is big, we cover it ourselves while the replacement ramps up, without the operation stopping.
Want it running in your business?
You’ve pinned the problem. We ship the fix and leave it measured.