2 system design questions lean on this idea. Each walks through the full answer.
A team ships a support chatbot backed by a large language model. It works well. One afternoon an engineer wants the tone to feel crisper, so they add a single line to the prompt: "Be concise and confident." The demo looks great. They merge it.
Three weeks later, complaints pile up. The bot is telling customers things that are flatly wrong, and saying them with total confidence. Refund windows that do not exist. Features the product does not have. Nobody changed the model. Nobody changed the data. Someone changed four words in a prompt, and quality fell off a cliff, silently, because "concise and confident" quietly told the model to stop hedging and stop citing its sources.
Here is the part that should bother you. This change was tested. The engineer ran it a few times by hand and it looked fine. That is the trap with language models. A handful of manual spot checks tells you almost nothing, because the output space is enormous and the failures hide in the cases you did not happen to try. The bad answers were never the ones the engineer typed during the demo. They were the odd refund question, the ambiguous account state, the customer who asked two things at once, and none of those were in the five prompts anyone tried before merging.
In traditional software, a bug throws an error. In an LLM system, a bug is a fluent, confident, wrong sentence that looks exactly like a correct one.
That asymmetry is the whole reason LLM evaluation is its own discipline. A crashing service pages you at 3am. A regressed language model just keeps answering, politely and wrongly, and the only alarm is a slow rise in support tickets weeks later. By the time a human notices, thousands of users have already been told something false. Evaluation exists to close that gap between the change and the consequence, from weeks down to minutes.
This lesson is about the discipline that catches that bug before it ships, and about the instruments that measure quality once it has. It is the least glamorous part of LLM ops and the part that separates a toy demo from a system you can trust with real users.
Grading a normal model is easy. A fraud classifier says fraud or not fraud, you have the true label, and accuracy is one line of code. The answer is right or it is wrong, and every metric you have ever used assumes exactly that: a prediction and a ground truth you can line up beside it.
Now grade this: "Explain our refund policy to an upset customer." There is no single correct string. A short answer can be perfect. A long answer can be perfect. Two completely different sentences can both be excellent, and a third that shares most of the same words can be terrible because it invented a policy. The thing you actually care about, whether this is a good answer, is not something you can check with an equals sign. There is no label to compare against, because the space of good answers is uncountable.
So the whole field of LLM evaluation is a set of tricks for approximating human judgment at a scale and cost where you cannot ask a human every time. Those tricks form a ladder, from cheap and objective at the bottom to expensive and human at the top. The bottom rungs are string and math checks that are perfectly precise but only apply when a right answer exists. The top rung is a person, which is the ground truth every other rung is trying to approximate.

The rule is simple: push as much grading as possible to the cheap bottom rungs, and spend your judge and human budget only on the open-ended quality that the cheap layers cannot reach. A production eval setup uses all five rungs at once, not one. You validate the JSON structure with a rule, catch an outright refusal with a regex, check a known numeric field with exact match, and only then hand the genuinely open-ended quality up to a judge, keeping a human sample to keep the judge honest.
The reason to be so careful about which rung you use is that they do not cost the same, and they do not buy the same thing. Plot every grader by what it costs per grade against how well it agrees with a human, and the shape is a curve with a sharp knee, not a straight line.

Offline evaluation runs before you ship, against a fixed set of test cases you control. That fixed set is called a golden set (or an eval set), and it is the single most important artifact in the whole system. Everything else in this lesson is machinery that operates on it.
A golden set is a curated collection of inputs paired with what good behavior looks like. Fifty to a few thousand examples covering the common cases, the tricky edge cases, and every past failure that once burned you. That last category matters more than the first. Every real production incident, the refund bug from the first slide included, should end its life as a permanent golden case, so the same failure can never ship twice. You store the set like code, usually as a JSONL file in git or a Hugging Face dataset, and you version it. When someone changes the prompt, it gets graded against the exact same cases as last time, so any comparison between two versions is fair rather than an accident of which inputs someone happened to try.
The graders that run against the golden set split into two families, and which family you can even use is decided by a single question: does this task have a known correct answer?

Reference-based metrics compare the output to a known correct answer: exact match, F1, BLEU, ROUGE, embedding similarity. They are cheap and objective but brittle. ROUGE will punish a perfect answer just for using different words than your reference, because it counts surface overlap, not meaning. Reference-free checks have no gold answer and instead test the output against a rule or a judge: JSON schema validity, toxicity and PII detection, on-topic checks, and the you will meet next. Reference-free is where most real LLM grading lives, because most real LLM tasks are open-ended and a near-identical answer can be wrong while two very different ones are both excellent.
Put the pieces together and the offline harness is not heavy machinery. It is one golden set feeding a runner, the runner calling the exact prompt and model you are grading, a stack of graders scoring every output, and one gate that decides ship or block.

When there is no reference answer and a rule cannot capture "is this good," you reach for the trick that makes modern LLM eval scale: use a strong model to grade the outputs of the model under test. This is , and it is the rung at the knee of the cost curve, the one that buys most of the human agreement for cents.
You hand a capable model (often GPT-4o or Claude) the question, the candidate answer, and a written rubric, and you ask it to score. Done well, it agrees with human graders often enough to replace them on most cases, at cents per grade instead of dollars and minutes. The rubric is what makes it repeatable: a vague instruction to rate quality gives you noise, while a written list of criteria (is it grounded in the provided facts, does it stay on tone, does it answer the whole question) gives you something a low score can be traced back to.

The catch is that judges have real, measurable biases, and if you ignore them your scores are noise dressed up as numbers:
You control these with a few habits. Force the judge to reason before it scores, because a score-first grade is measurably worse. Ask for a small discrete scale (1 to 5) with a written justification, so a low score is auditable. Pin the judge model version, because a silent upgrade changes your ruler underneath you. Traced as a sequence, a single grade is a short exchange, and the ordering of two of its steps is the whole trick.
If your system retrieves documents and then answers from them, a single quality score is worse than useless, because when an answer is wrong you cannot tell whether retrieval failed or the model failed. Those are two completely different bugs with two completely different fixes, and one number blends them into a shrug. evaluation fixes this by grading the pipeline in parts. The Ragas framework popularized three reference-free metrics that each tap a different stage.

Read them together and they point straight at the fix:
| Metric | Grades | Low score means |
|---|---|---|
| Context relevance | The retriever | Retrieval is pulling noise. Fix chunking, , top-k, or add a reranker. |
| Faithfulness | Model | The model is hallucinating on top of the context. It is asserting claims the documents never made. |
| Answer relevance | On-topic-ness | The model is evasive or wandering. Usually a prompt fix, not a retrieval fix. |
The dangerous quadrant is high context relevance with low faithfulness: the retriever did its job, handed the model the right documents, and the model still made things up. That is the case that produces confident, fluent, wrong answers, and it is exactly what a single accuracy number hides from you. A blended score of, say, 70 percent could be a retrieval problem or a grounding problem, and you would burn a week guessing. Grading these three separately turns that guess into a diagnosis: a low context-relevance number sends you to the embeddings and the reranker, a low faithfulness number sends you to the prompt and the model, and you stop debugging by superstition.
None of this is heavy machinery. A real eval harness is a loop over your golden set, a call to the system under test, a stack of scorers, and an assert on the aggregate. Here is the shape of one in Python, mixing an automatic check, an LLM judge, and a baseline gate.
import json
# 1. The golden set: versioned test cases in git or a HF dataset.
golden = [json.loads(line) for line in open("golden_set.jsonl")]
def score_case(case, system_under_test, judge):
output = system_under_test(case["question"]) # the prompt + model you are grading
# Cheap deterministic checks first (reference-free, free, instant).
valid_json = is_valid_json(output)
on_topic = not is_refusal(output)
# Reference-based check, only when a gold answer exists.
exact = None
if case.get("expected"):
exact = normalize(output) == normalize(case["expected"])
# LLM-as-judge for open-ended quality. Reason first, then score 1-5.
verdict = judge.grade(
question=case["question"],
answer=output,
rubric=case["rubric"], # written criteria, per task
reference=case.get("expected"), # optional
) # -> { "score": 4, "reason": "grounded, on tone" }
return {
"id": case["id"],
"valid_json": valid_json,
"on_topic": on_topic,
"exact": exact,
"judge_score": verdict["score"],
"judge_reason": verdict["reason"],
}
results = [score_case(c, my_system, my_judge) for c in golden]
avg_judge = sum(r["judge_score"] for r in results) / len(results)
# 2. The gate: this runs in CI on every prompt change.
BASELINE = 4.1 # last shipped version's score
assert all(r["valid_json"] for r in results), "some outputs are not valid JSON"
assert avg_judge >= BASELINE, f"quality regressed: {avg_judge:.2f} < {BASELINE}"
print(f"passed. avg judge score {avg_judge:.2f} (baseline {BASELINE})")
The three deterministic helpers (is_valid_json, is_refusal, normalize) are a few lines each, and system_under_test and judge are thin wrappers around your own prompt-plus-model and your judge call. That is the whole idea. It runs on your laptop in seconds and it runs in CI on every pull request. The assert at the bottom is what turns a nice report into a gate that blocks a bad change. Frameworks like Promptfoo, DeepEval, and Ragas give you the scorers, the reporting, and the CI wiring out of the box, but under the hood they are this loop, and understanding the loop is what lets you trust the framework instead of cargo-culting it.
Offline eval answers "is this change safe to ship?" It cannot answer "did this change actually help real users?" Only production can answer that. So you need both, and they are not interchangeable. Offline runs against a fixed golden set before the ship boundary; online runs against live traffic after it.

Online evaluation runs on live traffic. The signals you collect:
The habit that ties offline and online together, and the single highest-impact practice in LLM ops, is regression testing every prompt change. Treat the prompt like code. Every edit is a pull request, and CI runs the eval suite automatically, exactly like a unit test, blocking the merge if any protected case regresses.

This is not theory. Every serious LLM product runs some version of this loop, and the strongest teams have turned it into a standing process rather than a one-time audit.
Anthropic and OpenAI both lean heavily on human preference data and pairwise comparison to evaluate and align their models, which is where the whole "A versus B, which is better" pattern comes from. The reason they reach for comparison rather than absolute scoring is the same reason your judge should: people are far more consistent saying "this one is better" than "this one is a 7," and that consistency turns noisy individual votes into a stable ranking.

Chatbot Arena, run by LMSYS, evaluates frontier models purely through human pairwise votes at massive scale, and its leaderboard has become the number people actually trust, precisely because it is grounded in human preference rather than a gameable benchmark. The same A versus B pattern scales beyond evaluation into the preference data used to align models in the first place, which is why this one idea shows up everywhere from your CI gate to model training.
LangChain ships LangSmith specifically so teams can build golden datasets, run evaluators (including ) against every change, and track scores over time. The reason a whole product exists for this is that hand-rolled eval scripts do not survive contact with a growing team. The Ragas framework turned evaluation into a standard vocabulary. When engineers say faithfulness, context relevance, and answer relevance today, they are usually speaking Ragas, and those metrics are how RAG systems get debugged in the wild.
Put every piece from this lesson together and evaluation stops being a check you run once and becomes a loop the whole team runs continuously. You curate a golden set, change the prompt or model, grade it offline, gate it, ship a canary, watch the live signal, and feed every new production failure back into the golden set. The loop tightens the ruler over time instead of letting it rot.

4 questions - Score 80% to pass
Why is exact match a poor metric for most LLM tasks?
In a RAG system, an answer scores high on context relevance but low on faithfulness. What does that tell you?
What is the single highest-impact practice for preventing a one-line prompt edit from silently breaking production?
Why do teams prefer pairwise preference (is A or B better) over absolute 1-to-5 scores when evaluating with humans or an LLM judge?
The engineering goal is to sit at the knee, an LLM judge audited against humans, and to climb higher only where a wrong grade is genuinely expensive. A high-stakes medical or legal answer earns a human in the loop. A JSON formatting check does not. Match the grader's cost to the cost of being wrong about that particular answer, and let the cheap rungs carry the volume.
The gate is the part that turns a nice report into a control. Every hard check must pass and the average judge score must clear the last shipped baseline, and this exact gate is what runs in continuous integration on every prompt change. Without the gate you have a dashboard nobody reads. With it, a regression cannot merge.

Forcing the reasoning to come before the number is not decoration. A judge that commits to a score and then rationalizes it is both less accurate and impossible to audit, while a judge that argues each criterion first and scores last leaves a paper trail you can read when it disagrees with you.
The last habit is the one teams skip and regret: calibration. A judge score is worthless until you have measured it against the thing it approximates. Every so often, have humans grade the same sample and plot the judge score against the human score. Points on the diagonal are agreement; points far off it are the judge disagreeing with people, and those cases are gold because they show you exactly where the rubric or the judge is failing.

Report that agreement as a single tracked metric, not a one-time sanity check. If it slips, either the judge model changed under you or your rubric no longer matches what humans value, and both are things you want to catch before the judge quietly starts grading against the wrong standard. And where you can, prefer pairwise preference (is A or B better) over absolute scores, since humans and judges are both far more consistent at comparing two answers than at assigning an absolute number to one.
The candidate run cratered because the confident phrasing traded grounded facts for invented ones, and the gate refused the merge the moment the average dropped below the baseline. That is the entire mechanism, and it is small enough to hold in your head.
Watch this gate catch the exact bug from the first slide. The "be concise and confident" edit looked harmless and passed manual spot checks, but graded against the full golden set it cratered faithfulness, and the gate blocked it in minutes instead of letting it burn users for weeks. The four words that reached production for three weeks would have died in a failed pull request check, with a report naming the exact cases that broke.
One last thing to guard against: eval gaming. When a metric becomes a target, people optimize the metric instead of the goal. Teams overfit prompts to their golden set until the score is beautiful and production is unchanged. You defend against it by keeping a portion of the golden set held out and rotated, by refreshing cases from real production failures, and by always confirming an offline win with an online test before you believe it.
The through-line across all of them is the same lesson from the start of this track. The model is the engine. Evaluation is the instrument panel and the brakes, the thing that tells you whether the engine is still running right and stops you before you drive it off a cliff. A team without real LLM eval is not shipping carefully. They are shipping blind and finding out from their users.