LLM Evaluations — Evals
The step that separates amateurs from pros. Without a way to measure whether the system works well, every change is a gamble. Here's how to build systematic measurement.
Why evals are the most important step
Imagine you changed a prompt to improve one answer. How do you know you didn't break ten others? In regular software development you have tests. In AI Engineering, the equivalent is evals — a collection of test cases that automatically measure the quality of the system's output.
Without evals you develop "by feel": you change something, manually check 2–3 examples, and hope. With evals you know exactly whether a change improved things, broke them, or had no effect — across dozens or hundreds of cases. This is what lets you improve a system with confidence instead of being afraid to touch it.
Leading AI teams say: "whoever has good evals wins." Most of a serious AI Engineer's time goes into evals, not the prompt itself.
Types of evals
1. Rule-based / code
The fastest and cheapest. You check objective things in code: is the output valid JSON? Does it contain the required fields? Is the number in range? Great for structured outputs.
2. Reference-based
You have a known "correct answer" and compare against it. Suited to tasks with an unambiguous answer (classification, data extraction). Metrics: accuracy, precision/recall.
3. LLM-as-Judge
For open-ended tasks (writing quality, answer relevance) there's no single "correct answer." The solution: use another model as a judge that scores the output against criteria. Powerful and flexible — more on it below.
4. Human eval
The real gold, but expensive and slow. Humans rate a sample. You use it to calibrate the LLM-judge and make sure it agrees with humans.
Building an eval set — where to start
- Collect real cases. Take 20–50 real (or realistic) inputs the system is supposed to handle.
- Include edge cases. Not just the "normal case" — also empty input, mixed language, a manipulation attempt, an out-of-scope question.
- Define "what good looks like" for each case — an expected answer, or judging criteria.
- Start small. 20 good cases beat 500 bad ones. Expand over time, mainly from failures you saw in production.
Store the eval set as a file (JSON/CSV) in git, like code. It's a valuable asset that grows over time.
LLM-as-Judge — code example
The idea: a judge model receives the input, the output, and criteria, and returns a score + rationale. It's important to ask for a numeric score + explanation and use temperature 0.
JUDGE_PROMPT = """You are a judge of a support bot's answer quality.
Rate the answer from 1 to 5 by:
- relevance to the question
- accuracy (no wrong information)
- professional tone
Return JSON only: {"score": 1-5, "reason": "..."}
Question: {question}
Bot answer: {answer}"""
def judge(question, answer, client):
prompt = JUDGE_PROMPT.format(question=question, answer=answer)
resp = client.chat.completions.create(
model=JUDGE_MODEL, temperature=0, # pin this; a judge that changes invalidates your history
response_format={"type": "json_object"},
messages=[{"role": "user", "content": prompt}],
)
return json.loads(resp.choices[0].message.content)
# run over the whole eval set and average
scores = [judge(c["q"], run_system(c["q"]), client)["score"] for c in eval_set]
print("avg score:", sum(scores) / len(scores))
Critical tip: calibrate the judge against humans on a sample. If it agrees with humans ~85%+ of the time, you can trust it for most cases.
Regression testing & CI
The real power: run the evals automatically on every change. Wire them into CI (e.g. GitHub Actions), and if the average score drops below a threshold — the build fails. That way a change that improves one thing and breaks another is caught immediately.
# pseudo: eval gate in CI
avg = run_evals(eval_set)
THRESHOLD = 4.2
assert avg >= THRESHOLD, f"quality dropped: {avg} < {THRESHOLD}"
print(f"Evals passed: {avg}")
This turns "I think it's better" into "the numbers prove it's better." See also Observability for production measurement (online evals) on real traffic.
The average is the wrong number to look at
The natural way to report an eval run is a mean score, and it hides the thing you most need to see.
Consider a change that breaks ten cases and improves ten others. The average does not move. You ship it, and you have traded ten working behaviours for ten different ones without noticing — possibly trading the cases your users actually hit for the ones that were easy to write.
So the useful output of an eval run is not a number. It is a list of which cases changed state: what passed and now fails, what failed and now passes. That list is short, readable, and it is where the decision lives.
- Store per-case results, not just the aggregate, and diff them between runs.
- Gate on regressions, not only on the mean: no previously-passing case may now fail, or if one does, a human has to say why that is acceptable.
- Keep the distribution visible. A set where most cases score four and three score one is a different system from one where everything scores three, and both average the same.
The mean is still worth tracking as a trend line over months. It is a poor basis for deciding whether today's change ships.
Your set should look like your traffic
Eval sets are usually written by the person building the system, which means they contain the cases that person thought of — articulate, well-formed, and about the feature they were working on.
Real traffic is not like that. It contains typos, one-word questions, requests for things the product does not do, people pasting an entire email, and follow-ups that make no sense without the previous turn.
The fix is to sample rather than invent. Pull a hundred real inputs from logs, categorise them roughly, and build the set to match those proportions. If a fifth of real traffic is out-of-scope questions, then a fifth of your set should be too — otherwise you are measuring a system on a distribution it will never see.
Three categories almost every set is missing:
- Unanswerable questions, where the correct behaviour is to decline. Without these you never test the abstention path, and you will not notice when a change makes the system answer everything.
- Adversarial inputs — attempts to extract the prompt, to get it off-topic, to make it say something it should not.
- The boring middle. Sets skew towards interesting edge cases; most traffic is dull and the system must handle it well.
The judge has its own failure modes
Using a model to grade a model is practical and it is not neutral measurement. Known biases worth designing around:
- Length bias. Longer answers tend to score higher, largely independent of quality. If your change made outputs more verbose, expect the judge to approve.
- Position bias in pairwise comparisons — whichever answer is shown first has an advantage. Run each comparison both ways round and discard disagreements.
- Self-preference. A judge tends to favour output from its own model family. Using a different provider for the judge than for the system is a cheap mitigation.
- Rubric sensitivity. Small wording changes in the judging prompt shift scores meaningfully, which is why the judge prompt is versioned code and not something to tweak casually.
- Confident scoring of things it cannot know. Ask it to judge factual accuracy without giving it the source and it will produce a number anyway.
None of this makes the approach unusable. It makes it a measurement instrument that needs calibrating, like any other.
Designing a judge that agrees with you
The single largest improvement available: stop asking for a score out of five. A five-point scale asks the model to make a fine-grained aesthetic judgement, and small rubric changes move it around. Binary questions are far more stable.
Replace "rate this answer 1–5" with several yes/no checks, each asked separately:
- Does the answer address the question that was asked?
- Does every factual claim appear in the provided source?
- Does it avoid stating anything the source contradicts?
- Does it decline appropriately if the source does not contain the answer?
- Is the tone consistent with the guidelines?
Each is answerable with high agreement between two careful humans, which is the test of whether it is answerable by a judge. Count how many pass. That gives you a score built from defensible parts, and when it drops you can see which part dropped.
Two more design points. Ask one criterion per call — a judge asked five things at once attends to the first and reasons its way to a consistent verdict. And give the judge only what it needs: the answer and the source, not the whole conversation, which gives it room to rationalise.
Calibrating, concretely
"Check the judge agrees with humans" is right and usually left as an aspiration. The procedure is small:
- Take fifty outputs from a real run, spanning good and bad.
- Have a person label them against the same criteria the judge uses. Ideally two people, so you can see how much humans agree with each other — if they only agree eighty per cent of the time, no judge will do better and your criteria need sharpening.
- Compare, and look at the disagreements individually rather than at the percentage. The pattern in them tells you what to fix.
- Adjust the rubric, not the labels, and re-run.
- Repeat when you change the judge model, the rubric, or the system in a way that changes the output shape.
The disagreements are the valuable part. They are usually cases where your criteria were ambiguous — which means your team disagreed about what good looks like, and the judge was only the first to notice.
Cost, speed and what to run when
A full eval run is a batch of model calls, which is real money and several minutes. That tension is what stops teams running them, so plan the tiers deliberately:
- On every commit: the rule-based checks. Free, instant, and they catch schema and format breakage — which is a large share of real regressions.
- On every pull request: a fast subset, perhaps a fifth of the set, chosen to cover the categories rather than at random.
- Nightly and before release: the full set, including the judged cases.
- On demand: the full set plus a larger sample, when changing model or making a structural change.
Run the cases concurrently — they are independent, and a serial run is what makes evals feel slow enough to skip. And cache results per (case, system version) so re-running after an unrelated change costs nothing.
Making the CI gate survivable
A gate that fails for reasons nobody caused gets disabled within a month, so the threshold design matters as much as the set.
- Gate with a margin. Provider-side variation moves scores slightly. A threshold set at exactly yesterday's number turns normal noise into a red build.
- Measure the noise first. Run the same set three times without changing anything and see how much the score moves. That range is your margin, and it is a number rather than a guess.
- Pin the judge model and its prompt. If either changes, every historical score becomes incomparable and your trend line is meaningless.
- Make the failure output readable. "Score dropped to 4.1" sends nobody anywhere. "These three cases regressed, here is the input and both outputs" is actionable.
- Allow a documented override. Sometimes a regression is an accepted trade. A human approving it in the pull request is fine; silently lowering the threshold is not.
Eval sets go stale
An eval set is code and it decays like code, in ways that are easy to miss because it keeps producing a number.
- Cases that no longer reflect the product. A feature changed and forty cases now assert the old behaviour. They fail, someone "fixes" them by loosening the assertion, and the set gets weaker.
- Overfitting. After months of tuning against the same hundred cases, you have a system that does well on those hundred. Hold back a portion you never tune against, and rotate in fresh cases from production.
- Leakage. Public benchmark questions end up in training data over time, which is part of why public leaderboards flatter models. Your own private set does not have this problem — one more reason it beats a benchmark.
- Stale expected answers, where the correct response depends on data that has since changed.
A quarterly review — read the set, delete what is obsolete, add what production taught you — keeps it honest. It is an hour, and skipping it is how a green build stops meaning anything.
The first afternoon, if you have nothing
Most teams know they should have evals and do not, because the topic sounds like infrastructure. It is not. The first useful version is a file and a loop, and it takes an afternoon.
- Open a spreadsheet or a JSON file. Two columns: the input, and what a good answer looks like — in prose, as a note to yourself, not as an exact string to match.
- Fill in twenty rows from real requests if you have logs, from your own honest guesses if you do not. Include three that should be declined and two that are adversarial.
- Write the loop that runs each input through your system and prints the input and the output side by side.
- Read all twenty yourself and mark pass or fail. This is the part that feels too manual to be real engineering, and it is where you learn more about your system than any dashboard will tell you.
- Now make a change — a prompt edit you were going to make anyway — and re-run. Read the twenty again. You have just done a controlled experiment, which is more than most teams shipping AI features ever do.
- Only then automate the grading, on the criteria you found yourself applying while reading.
The order matters. Teams that start by building a judging framework end up with infrastructure measuring criteria nobody validated. Teams that start by reading twenty outputs discover what their criteria actually are — and that knowledge is what makes the automated version worth anything.
What evals cannot tell you
Worth being clear, because a passing eval suite creates confidence that should be bounded.
Evals measure a fixed set of inputs. They say nothing about the inputs you have not thought of, and real traffic drifts — users discover new phrasings, a marketing campaign brings a different audience, someone links to you from an unexpected place. A system can pass every test and degrade in production because the distribution moved underneath it.
That is the gap observability fills, and the two are a loop rather than alternatives: production tells you what is actually being asked, the interesting failures become eval cases, and the set stays representative because it is fed from reality.
The practical implication: do not treat a green eval run as permission to stop watching. Treat it as permission to ship, which is a smaller and more useful claim.
Common mistakes
- No eval set at all. The big mistake. Even 20 cases change everything.
- Only "easy cases." If the evals don't include edge cases, they give false confidence.
- An uncalibrated judge. An LLM-judge without a check against humans may give misleading scores.
- High judge temperature. Causes inconsistent scores. Always 0.
- A single metric. One score hides problems. Measure several dimensions (accuracy, tone, safety) separately.
- Not updating. Every production failure should become a new case in the eval set.
Next step
Got evals? Now you can add safety and monitor in production with confidence.