AI Operations · Role
AI Reliability Engineer (LLMOps)
The AI Reliability Engineer (LLMOps) keeps AI systems standing: they monitor, control cost and latency, and respond when an agent fails or drifts. The SRE of the AI layer.
What they do
- Instrument agent observability: cost, latency, errors and quality.
- Define alerts, fallbacks and per-agent spend limits.
- Diagnose production incidents and prevent regressions.
Skills it needs
Salary by country and seniority
| Country | Seniority | P25 | Median | P75 | |
|---|---|---|---|---|---|
| Germany | All | €73.000 | €73.000 | €73.000 | Stima da fonti pubbliche |
| Spain | All | €70.000 | €70.000 | €70.000 | Stima da fonti pubbliche |
| France | All | €61.225 | €75.300 | €80.350 | Stima da fonti pubbliche |
| Italy | All | €46.033 | €46.033 | €46.033 | Stima da fonti pubbliche |
| United States | Mid | €114.927 | €114.927 | €114.927 | Stima da fonti pubbliche |
| United States | Senior | €142.042 | €142.042 | €142.042 | Stima da fonti pubbliche |
| United States | All | €142.042 | €142.042 | €142.042 | Stima da fonti pubbliche |
Figures in euros, gross annual — each with its source.
Calculate a salaryAlso known as
The market uses several titles for the same role:
How it changes by seniority
- Junior — Instruments logs and dashboards and handles alerts by runbook.
- Mid — Defines alerts, fallbacks and per-agent spend limits; diagnoses incidents.
- Senior — Owns the agent fleet’s reliability and prevents regressions at scale.
When to hire
When you have several agents in production and their cost or reliability is slipping.
When not to
If you still have a single pilot with no real traffic: nothing to make reliable yet.
Interview questions
1. A production agent degrades in quality without anyone touching the code. How do you diagnose it?
What to look for: Thinks drift (provider model, input data), not just bugs; uses observability.
2. What do you instrument in an LLM system to operate it with confidence?
What to look for: Traces inputs/steps/cost/latency and alerts on the abnormal; dashboards for the owner, not just logs.
3. The provider is retiring the model you use. What’s your plan?
What to look for: Re-runs evals with the new model, versions and prepares rollback; doesn’t trust "it’ll behave the same".
4. How do you set SLAs on something probabilistic like an LLM?
What to look for: Defines targets by accuracy/latency/cost and escalates to a human when they slip.
Scorecard
- Drift detection — Separates code failure from a changing world; catches degradation in time.
- LLM observability — Instruments inputs, steps, cost and latency; useful alerting, not noise.
- Model-change management — Versions, re-evaluates and prepares rollback for provider deprecations.
- Operable reliability — Defines measurable targets and escalation paths for something probabilistic.
Red flags
- — Only logs the final output: when it fails, there’s no way to know why.
- — Switches models without re-running the evals.
- — Alerts on everything until nobody heeds the alerts.
Hiring or assessing this profile?
Generate the full interview script and scorecard, tailored to the seniority you’re hiring for.
Open the interview scorecard