The conversation always opens with a tool name. Somebody saw a demo, somebody read a thread, and the question that reaches the meeting is whether we self-host Langfuse or pay for a platform. It is the wrong question, and it shows up three weeks later: tracing is in, the dashboard exists, and nobody can say whether the agent answers worse than it did last month.
The thesis in one line: AI agent observability is not a tool, it is four layers you decide separately, and open pieces cover all four today. What breaks is not the tool — the open ones work — but the two things nobody puts in the comparison table: the data convention is still in development and its documentation has moved, and “open” means at least three licences that do not let you do the same things.
The AI agent observability tools you can self-host: four layers, not one
The starting mistake is treating this as a single purchase. “Observability” gets used as if it were one product, and in an agent it is four different jobs that fail in different ways and almost never all matter on the same day. Splitting them is what turns a brand choice into an engineering decision.
| Layer | What question it answers | What happens without it |
|---|---|---|
| Correlated traces | What did the agent do, in what order, with which tools and what input at each step? | You debug blind: you have the final output and no way to know which step went sideways |
| Pre-production evaluation | Does this prompt or model change beat the baseline, or lose to it? | Every change is a bet, and the only test is waiting for somebody to complain |
| Cost attribution | What does each case, each agent and each tool cost — not the aggregate invoice? | You see the provider total at month end and can neither attribute nor cut it |
| Drift detection | Is it still answering like it did a month ago, with the same model underneath? | Quality slides slowly, classic monitoring reports all green, and a customer finds it first |
The first two are covered today by any of the serious open pieces. The third depends on having instrumented at the right granularity from the start — adding it later means reinstrumenting. The fourth is not a feature you switch on: it is a cadence somebody has to run, and it is the one most likely to end up ownerless. Why drift produces a slide rather than an error we work through in what happens when the AI model changes underneath your automations; what matters here is which layer should have caught it.
The first thing that breaks is not the tool: it is the convention
For the four layers to talk to each other, they all have to call the same things by the same name: what a model call is, what an agent invocation is, what a tool execution is. That is a semantic convention, and it is the part of the stack almost nobody checks before choosing. It is also the least settled.
Two concrete facts, both checkable at the primary source. First: OpenTelemetry’s GenAI conventions — the gen_ai.* attributes — are still marked as in development, not stable. Second, and the one that will actually bite you: that documentation no longer lives where your search engine will send you. The historical page in the semantic-conventions repository now says only that the content has moved to the GenAI semantic conventions repository and that the old one is no longer maintained.
On top of that there is a second convention in circulation: OpenInference, the semantic layer Arize maintains and its open tool runs on. Not a problem in itself — exporters translate — but it is why two pieces both “OpenTelemetry-compatible” can arrive with span trees that do not fit. Pick one and make it the house standard; do not let each instrumenting team choose.
“Open” means three different licences, and only one lets you do what you think
This is where the open stack looks far more like the argument that already exists in automation than anyone admits. It is exactly the trap we take apart in open source AI automation tools: “€0 licence” is not “€0 cost”, and “open source” is not always what the words suggest. Agent observability works the same way, with one twist of its own.
| Piece | What its own docs say | What it means in practice |
|---|---|---|
| Langfuse | Open source and self-hostable with Docker on your own infrastructure; some add-on features require a licence key | For running your own fleet the core is enough. Its self-hosting feature list marks three as enterprise — organization creators, instance management API and UI customization |
| Arize Phoenix | Released under Elastic License 2.0; self-hosting on your own infrastructure or cloud account is free and fully permitted, with no feature gates | Used in-house, you get no clipped features. ELv2 is not OSI-approved and its core restriction is offering it to third parties as a managed service |
| OpenTelemetry | A CNCF project; the GenAI conventions are in development | The plumbing is the most solid and least arguable part of the stack. The instability is in the vocabulary, not the transport |
The operational read: if you build this for your own operation, all three licences let you, and the legal conversation is short. If at some point you plan to resell the dashboard to your clients as a service — an agency, a managed service provider — ELv2 is precisely what forbids it, and that is worth knowing before you build a business on top, not after.
The point where the open stack stops paying off (and it is not a number of agents)
The expected answer here is a figure: up to X agents self-host, past X buy. That figure does not exist, and anyone handing it to you is selling the side that suits them. The threshold is not about volume. It is about guarantees.
The vendor says it in its own documentation, and it is the most honest line in the whole category. In Langfuse’s deployment options table, the Docker Compose start is described as a single VM without high availability, scaling or backups, and production self-hosting — Kubernetes, AWS, Azure, GCP — lists responsibility in one column only: your infrastructure.
That is the real threshold, and it has nothing to do with how many agents you run. The open stack stops paying off the day the dashboard becomes critical infrastructure: when somebody outside notices the outage, when the trace is the evidence you need to answer a customer or an auditor, or when retaining six months of traces becomes a requirement rather than a preference. That day you are no longer picking a tool, you are deciding who gets up to fix it. Which is the same question — asked about automations — as who answers when an automation goes down.
It also pays to read any vendor’s scale numbers at arm’s length. Langfuse states in its documentation that it processes more than 90 billion observations a month and that 21 of the Fortune 50 use it. Those are vendor figures about its own product, not an independent measurement and not a result of ours; they tell you the piece takes load, not what it will take in your house.
When NOT to build the open stack
Four situations where this project is a detour, not progress. All four are real and all four arrive dressed as a good idea:
- You have one agent and it is not instrumented. You do not need an observability platform, you need tracing. Instrument first with the convention pinned and decide where you ship it after; the reverse order has you choosing a tool with zero data about what you need to see.
- Nobody is going to look at the dashboard. A dashboard with no on-call is not observability, it is infrastructure spend with charts. If there is no named person on a rota, build that first.
- Your problem is quality, not health. If the complaint is “it answers badly”, no trace fixes it: what you need is a rubric and a case bank, which is a different job — we cover it in evaluating AI agent quality.
- The agent writes to a system of record and has no access matrix yet. Observability tells you what it did; it does not stop it being able to. That order lives in what permissions to give an AI agent, and it comes first.
What you measure in month one
The sign the stack is genuinely built is not that the dashboard loads. It is four questions you could not answer before and now answer in a minute, with the trace in front of you:
- Of this month’s failures, how many were caught by an alert and how many by a person outside. It is the only metric that says whether the dashboard works.
- What the most expensive case cost, and why. If the answer is “it cannot be broken down”, the attribution layer is missing, token chart or no token chart.
- How long it takes you to reconstruct what happened in a specific case three weeks ago. If it is hours, the traces are there but the correlation is not.
- How many prompt or model changes shipped without going through the baseline. That number should be zero, and in month one it almost never is.
Note that none of the four is a metric about the tool. They are metrics about the operation, which is why the open stack does not finish on deployment day: what Docker gets you is the cheap half. When the team cannot carry that continuous work, the conversation is monitoring AI in production as a function, not as a dashboard.
The line for the next meeting
Next time somebody asks which agent observability tool we should stand up, the useful answer is not a name. It is: tell me which of the four layers you are missing, which convention you are pinning, and who gets up when it falls over. With those three answers the tool choice makes itself, and it is almost always the open one.