Skip to main content
Level: Advanced Updated: September 2026

LLM Evaluations — Evals

The step that separates amateurs from pros. Without a way to measure whether the system works well, every change is a gamble. Here's how to build systematic measurement.

Why evals are the most important step

Imagine you changed a prompt to improve one answer. How do you know you didn't break ten others? In regular software development you have tests. In AI Engineering, the equivalent is evals — a collection of test cases that automatically measure the quality of the system's output.

Without evals you develop "by feel": you change something, manually check 2–3 examples, and hope. With evals you know exactly whether a change improved things, broke them, or had no effect — across dozens or hundreds of cases. This is what lets you improve a system with confidence instead of being afraid to touch it.

The industry's truth

Leading AI teams say: "whoever has good evals wins." Most of a serious AI Engineer's time goes into evals, not the prompt itself.

Types of evals

1. Rule-based / code

The fastest and cheapest. You check objective things in code: is the output valid JSON? Does it contain the required fields? Is the number in range? Great for structured outputs.

2. Reference-based

You have a known "correct answer" and compare against it. Suited to tasks with an unambiguous answer (classification, data extraction). Metrics: accuracy, precision/recall.

3. LLM-as-Judge

For open-ended tasks (writing quality, answer relevance) there's no single "correct answer." The solution: use another model as a judge that scores the output against criteria. Powerful and flexible — more on it below.

4. Human eval

The real gold, but expensive and slow. Humans rate a sample. You use it to calibrate the LLM-judge and make sure it agrees with humans.

Building an eval set — where to start

  1. Collect real cases. Take 20–50 real (or realistic) inputs the system is supposed to handle.
  2. Include edge cases. Not just the "normal case" — also empty input, mixed language, a manipulation attempt, an out-of-scope question.
  3. Define "what good looks like" for each case — an expected answer, or judging criteria.
  4. Start small. 20 good cases beat 500 bad ones. Expand over time, mainly from failures you saw in production.

Store the eval set as a file (JSON/CSV) in git, like code. It's a valuable asset that grows over time.

LLM-as-Judge — code example

The idea: a judge model receives the input, the output, and criteria, and returns a score + rationale. It's important to ask for a numeric score + explanation and use temperature 0.

JUDGE_PROMPT = """You are a judge of a support bot's answer quality.
Rate the answer from 1 to 5 by:
- relevance to the question
- accuracy (no wrong information)
- professional tone
Return JSON only: {"score": 1-5, "reason": "..."}

Question: {question}
Bot answer: {answer}"""

def judge(question, answer, client):
    prompt = JUDGE_PROMPT.format(question=question, answer=answer)
    resp = client.chat.completions.create(
        model=JUDGE_MODEL, temperature=0,   # pin this; a judge that changes invalidates your history
        response_format={"type": "json_object"},
        messages=[{"role": "user", "content": prompt}],
    )
    return json.loads(resp.choices[0].message.content)

# run over the whole eval set and average
scores = [judge(c["q"], run_system(c["q"]), client)["score"] for c in eval_set]
print("avg score:", sum(scores) / len(scores))

Critical tip: calibrate the judge against humans on a sample. If it agrees with humans ~85%+ of the time, you can trust it for most cases.

Regression testing & CI

The real power: run the evals automatically on every change. Wire them into CI (e.g. GitHub Actions), and if the average score drops below a threshold — the build fails. That way a change that improves one thing and breaks another is caught immediately.

# pseudo: eval gate in CI
avg = run_evals(eval_set)
THRESHOLD = 4.2
assert avg >= THRESHOLD, f"quality dropped: {avg} < {THRESHOLD}"
print(f"Evals passed: {avg}")

This turns "I think it's better" into "the numbers prove it's better." See also Observability for production measurement (online evals) on real traffic.

The average is the wrong number to look at

The natural way to report an eval run is a mean score, and it hides the thing you most need to see.

Consider a change that breaks ten cases and improves ten others. The average does not move. You ship it, and you have traded ten working behaviours for ten different ones without noticing — possibly trading the cases your users actually hit for the ones that were easy to write.

So the useful output of an eval run is not a number. It is a list of which cases changed state: what passed and now fails, what failed and now passes. That list is short, readable, and it is where the decision lives.

The mean is still worth tracking as a trend line over months. It is a poor basis for deciding whether today's change ships.

Your set should look like your traffic

Eval sets are usually written by the person building the system, which means they contain the cases that person thought of — articulate, well-formed, and about the feature they were working on.

Real traffic is not like that. It contains typos, one-word questions, requests for things the product does not do, people pasting an entire email, and follow-ups that make no sense without the previous turn.

The fix is to sample rather than invent. Pull a hundred real inputs from logs, categorise them roughly, and build the set to match those proportions. If a fifth of real traffic is out-of-scope questions, then a fifth of your set should be too — otherwise you are measuring a system on a distribution it will never see.

Three categories almost every set is missing:

The judge has its own failure modes

Using a model to grade a model is practical and it is not neutral measurement. Known biases worth designing around:

None of this makes the approach unusable. It makes it a measurement instrument that needs calibrating, like any other.

Designing a judge that agrees with you

The single largest improvement available: stop asking for a score out of five. A five-point scale asks the model to make a fine-grained aesthetic judgement, and small rubric changes move it around. Binary questions are far more stable.

Replace "rate this answer 1–5" with several yes/no checks, each asked separately:

Each is answerable with high agreement between two careful humans, which is the test of whether it is answerable by a judge. Count how many pass. That gives you a score built from defensible parts, and when it drops you can see which part dropped.

Two more design points. Ask one criterion per call — a judge asked five things at once attends to the first and reasons its way to a consistent verdict. And give the judge only what it needs: the answer and the source, not the whole conversation, which gives it room to rationalise.

Calibrating, concretely

"Check the judge agrees with humans" is right and usually left as an aspiration. The procedure is small:

  1. Take fifty outputs from a real run, spanning good and bad.
  2. Have a person label them against the same criteria the judge uses. Ideally two people, so you can see how much humans agree with each other — if they only agree eighty per cent of the time, no judge will do better and your criteria need sharpening.
  3. Compare, and look at the disagreements individually rather than at the percentage. The pattern in them tells you what to fix.
  4. Adjust the rubric, not the labels, and re-run.
  5. Repeat when you change the judge model, the rubric, or the system in a way that changes the output shape.

The disagreements are the valuable part. They are usually cases where your criteria were ambiguous — which means your team disagreed about what good looks like, and the judge was only the first to notice.

Cost, speed and what to run when

A full eval run is a batch of model calls, which is real money and several minutes. That tension is what stops teams running them, so plan the tiers deliberately:

Run the cases concurrently — they are independent, and a serial run is what makes evals feel slow enough to skip. And cache results per (case, system version) so re-running after an unrelated change costs nothing.

Making the CI gate survivable

A gate that fails for reasons nobody caused gets disabled within a month, so the threshold design matters as much as the set.

Eval sets go stale

An eval set is code and it decays like code, in ways that are easy to miss because it keeps producing a number.

A quarterly review — read the set, delete what is obsolete, add what production taught you — keeps it honest. It is an hour, and skipping it is how a green build stops meaning anything.

The first afternoon, if you have nothing

Most teams know they should have evals and do not, because the topic sounds like infrastructure. It is not. The first useful version is a file and a loop, and it takes an afternoon.

  1. Open a spreadsheet or a JSON file. Two columns: the input, and what a good answer looks like — in prose, as a note to yourself, not as an exact string to match.
  2. Fill in twenty rows from real requests if you have logs, from your own honest guesses if you do not. Include three that should be declined and two that are adversarial.
  3. Write the loop that runs each input through your system and prints the input and the output side by side.
  4. Read all twenty yourself and mark pass or fail. This is the part that feels too manual to be real engineering, and it is where you learn more about your system than any dashboard will tell you.
  5. Now make a change — a prompt edit you were going to make anyway — and re-run. Read the twenty again. You have just done a controlled experiment, which is more than most teams shipping AI features ever do.
  6. Only then automate the grading, on the criteria you found yourself applying while reading.

The order matters. Teams that start by building a judging framework end up with infrastructure measuring criteria nobody validated. Teams that start by reading twenty outputs discover what their criteria actually are — and that knowledge is what makes the automated version worth anything.

What evals cannot tell you

Worth being clear, because a passing eval suite creates confidence that should be bounded.

Evals measure a fixed set of inputs. They say nothing about the inputs you have not thought of, and real traffic drifts — users discover new phrasings, a marketing campaign brings a different audience, someone links to you from an unexpected place. A system can pass every test and degrade in production because the distribution moved underneath it.

That is the gap observability fills, and the two are a loop rather than alternatives: production tells you what is actually being asked, the interesting failures become eval cases, and the set stays representative because it is fed from reality.

The practical implication: do not treat a green eval run as permission to stop watching. Treat it as permission to ship, which is a smaller and more useful claim.

Common mistakes

Next step

Got evals? Now you can add safety and monitor in production with confidence.