Solution · AI Operations
The model you bought is going away. The only question is whether you find out before your customer does.
Picking a model isn't a one-time decision: providers retire versions on notice, update them underneath you and move prices, and your system stays in production through all of it. We build the function that turns switching models into a boring operation — your own test bench, real comparison, and a way back — instead of a leap of faith.
The problem
Your model provider is a supplier. You're treating it like a constant.
- The model you run was picked one afternoon, eighteen months ago, because it was the good one at the time. Nobody has revisited the decision and nobody knows whether it's still right.
- A deprecation email arrived with a date on it. Nobody knows how many of your systems depend on that exact version, so nobody knows how much work sits behind that date.
- Somebody tried a new model on five hand-written examples, thought it looked better, and shipped it. The comparison was an impression, not a measurement.
- One morning quality changed with nobody touching anything. The team argued whether it was the model, the prompt or the data, decided on gut feel, and moved on.
- The same expensive model handles the task that needs judgment and the one that just sorts an email into three buckets, because deciding which goes where was never anyone's job.
- If the provider goes down for half an hour, you go down for half an hour. There's no second path and nobody has ever tested whether there could be.
- Switching models feels like a six-week project, so it doesn't happen — and an eighteen-month-old decision stays frozen out of fear, not judgment.
Cost of staying the same
You have a critical supplier whose product changes under your feet and no way to measure the effect. That gets paid for three ways. First, surprise: a retirement date lands and the migration happens in a rush, with no test bench, with the risk spread across your customers. Second, silent drift, which is worse: the system doesn't error, it just starts answering slightly differently — longer, in another format, refusing things it used to accept — and since nobody measures, you find out through a complaint. Third, opportunity cost: because switching is scary, you don't, so you keep paying premium-model prices for cheap tasks and keep running a worse model than the one already available.
The solution
We turn a model switch into a measured, reversible, boring operation
- 1The inventory first, which almost nobody has: which of your systems call which model at which exact version, what task each call performs, and what would break if that version vanished tomorrow. Without it, any deprecation notice is an emergency instead of a scheduled task.
- 2We build your golden case bench: a frozen set of real cases from your operation — the normal ones, the weird ones and the ones that went wrong — with the result you consider correct. It's the asset that turns model comparison from an impression into a measurement, and it's yours even if you fire us tomorrow.
- 3We define per-task routing: which calls genuinely need the good model and which get solved by a cheaper or faster one with no drop in outcome. That decision gets made with the bench in front of you, not with the intuition of whoever talks loudest.
- 4We build the migration procedure: candidate against the bench, side-by-side comparison on your real work, gradual rollout to a slice of traffic, live output comparison, and only then the full switch. With the rollback prepared and rehearsed before you start, not improvised on the bad day.
- 5We install drift detection. Providers also update without changing the version name, and that doesn't error: format shifts, length shifts, JSON adherence loosens, and the line where the model says no moves. The bench runs on a schedule so that drift gets caught by an alert and not by a customer.
- 6And we put it on the calendar. Your providers' retirement dates land as alerts with runway — not as an email somebody archived — and the model decision gets reviewed on an agreed cadence, with cost per task and measured quality on the table.
What changes
What you stop losing
Anthropic commits to a minimum of 60 days' notice before retiring public models, and OpenAI to a minimum of six months for its generally available models. That notice is only useful if somebody knows which of your systems depend on the affected version: the inventory turns a date into a plan.
Providers' public deprecation policies (Anthropic, OpenAI), accessed 2026-08
Switching models stops being a leap of faith: the candidate gets measured against your own cases, using your definition of correct, before it touches a single customer.
Mechanism
Silent drift stops being discovered through a complaint. When the provider updates underneath you, outputs change without erroring — format, length, JSON adherence, where it refuses — and the scheduled bench turns that into an alert.
Mechanism
Cost drops by design rather than by cutting: when each task goes to the model it actually needs, the saving comes from not paying for judgment where sorting was enough — not from using worse AI.
Mechanism
What we measure: systems with an inventoried model version, candidate quality against the golden case bench, deviation caught by scheduled runs, cost per task before and after routing, rehearsed rollback time, and days of runway on every announced retirement date.
What we measure
Spec sheet
- Work it removes
- your company's model decision freezing out of fear, and every provider retirement, update or price change getting handled in a rush, unmeasured, with customers in the middle
- Typical setup
- 4–8 weeks
- Input
- your AI systems in production, the model versions they call today, and real cases from your operation — including the ones that went wrong
- Output
- a live inventory of what depends on which version, a golden case bench of your own, per-task routing, a migration procedure with a rehearsed rollback, and alerts for drift and retirement dates
- Works with
- OpenAIAnthropicGoogle Vertex AIAzure OpenAIAmazon BedrockMistralLangSmithLangfuseBraintrust
- Can connect to
- Your production logs, where the bench cases come fromYour evaluation layer and prompt management, if they already existYour model gateway or router, so switching is configurationYour cost-per-model and cost-per-task dashboardYour change process and on-call calendar
- What we measure
- % of systems with an inventoried model version and a named ownercandidate quality against the golden case benchbehavioral deviations caught by scheduled runscost per task before and after routingrehearsed rollback timedays of runway on every announced retirement date
- Good fit for
- companies with AI already in production and real dependence on one or more model providers — CIO, CTO or head of AI — who need the model decision to be reviewable, measurable and reversible rather than a blind commitment
- Not a fit for
- anyone still in pilot with no real traffic yet — there's no bench to build there, you have to reach production first — or anyone who wants us to name the best model on the market: that depends on your work, which is exactly why it gets measured instead of opined on
Frequently asked questions
And we still say it: switching models because something underperforms is almost always moving the problem, because the bottleneck is usually a badly defined process or dirty data. This is the opposite of that. Here you're not switching to fix anything: you're switching because the provider is retiring the version you run, because they updated it underneath you, because an option appeared that does the same for a fraction of the cost, or because the decision you made eighteen months ago isn't the one you'd make today. Those changes are coming whether you want them or not. What we build is the ability to absorb them without drama and with evidence — and, along the way, the thing that tells you when the model isn't the problem, which is exactly when you shouldn't touch it.
It's a frozen set of real cases from your operation with the result you consider correct: the normal ones, the weird ones, the ambiguous ones and above all the ones that went wrong. Dozens to a few hundred well-chosen cases is usually enough; you don't need thousands. It matters most because it's the only thing that turns «this model seems better» into a comparable number: you run the candidate through the same set as the current model and see exactly where it improves and where it degrades, on your work rather than on a generic exam. And it's the most valuable thing you take away from the project, because it doesn't depend on any provider: it works just as well next year with models that don't exist yet. That's why we build it with you and it stays in your house.
Because they measure something else. A public benchmark tells you how a model behaves on a standardized exam that looks nothing like your operation: not your documents, not your vocabulary, not your edge cases, not your definition of correct. It's useful information for ruling out clearly weaker candidates and for reading where the market is going, and we use it for that. But a ranking can't make the decision: a model can score better in the open and perform worse on one specific task with specific documents, and the reverse happens too. The provider, on top of that, has an obvious stake in the comparison. Your case bench doesn't.
By running the bench on a schedule against the same endpoint and comparing against the baseline. That's the only reliable method, because a silent update doesn't produce an error: it produces drift. Tool-call formatting shifts a little, JSON adherence loosens, responses get longer or shorter, and the boundary of what the model refuses moves. None of that breaks the integration, so classic monitoring — availability, latency, error rate — reports all green while quality slides. With scheduled bench runs, that drift surfaces as an alert with concrete cases attached, and then the conversation is «this changed on Tuesday and here are the cases» instead of «I feel like it's been answering oddly lately».
Less than any alternative, and it's easy to verify: the deliverables are yours and don't depend on us. The version inventory, the golden case bench, the routing rules, the migration procedure and the retirement calendar are documented and live in your systems. We build on your stack and on whatever evaluation tooling you already have; if you have none, we choose with you and explain why, with no proprietary layer you'd later have to escape. The normal way this ends is your team running it — that's training your team, a different service — or us running it while that capability matures. Both exits are written down from the start.
Want it running in your business?
You’ve pinned the problem. We ship the fix and leave it measured.