Vai al contenuto
Se non funziona, non paghi. 30 giorni.
Implementa.

AI Operations · Role

AI Evaluation Engineer

The AI Evaluation Engineer decides whether an AI system is good enough for production. They design the evals, measure quality objectively and block whatever misses the bar.

What they do

  • Build evaluation sets and per-use-case quality metrics.
  • Wire automated evals into the pipeline (a gate before deploy).
  • Catch quality regressions and hallucinations before users do.

Skills it needs

Eval and test-dataset designQuality metrics and gradingError analysisAutomated testingPrompt and model tuning

Salary by country and seniority

CountrySeniorityP25MedianP75
United StatesAll€117.040€117.040€117.040Stima da fonti pubbliche

Figures in euros, gross annual — each with its source.

Calculate a salary

Also known as

The market uses several titles for the same role:

AI Quality EngineerLLM Evaluation EngineerAI Test EngineerModel Quality Engineer

How it changes by seniority

  • Junior Builds evaluation sets and runs evals against a given design.
  • Mid Wires automated evals as a pipeline gate and defines per-use-case metrics.
  • Senior Designs the evaluation strategy and catches regressions before users do.

When to hire

When 'seems to work' is no longer enough and you need measured quality to trust it.

When not to

If volume is low and occasional manual review still covers you.

Interview questions

  1. 1. You have to decide whether a new agent version is better than the current one. How do you set it up?

    What to look for: Builds an evaluation set with real cases and a threshold; measures rather than opines.

  2. 2. How many cases does a good eval need and how do you choose them?

    What to look for: Representativeness over quantity; starts with well-chosen cases and grows with production failures.

  3. 3. How do you separate tolerable errors from errors that can never happen?

    What to look for: Separates the average from critical cases and blocks those even if the average is good.

  4. 4. How do you evaluate something subjective, like the quality of generated text?

    What to look for: Uses rubrics, multiple evaluators or LLM-as-judge with judgement, not just "sounds good to me".

Scorecard

  • Eval designBuilds sets with real cases, expected output and a pass threshold.
  • RepresentativenessCovers the real variety of the process; grows with production failures.
  • CriticalitySeparates tolerable errors from ones that can’t happen and blocks those.
  • Evaluating the subjectiveApplies rubrics or judges with judgement for non-binary quality.

Red flags

  • Evaluates "by eye" reading a few outputs.
  • A toy eval with four cases that don’t represent production.
  • Looks only at the average and not the cases that can’t fail.

Hiring or assessing this profile?

Generate the full interview script and scorecard, tailored to the seniority you’re hiring for.

Open the interview scorecard