None of these metrics can be computed from the system's output alone. Each test question needs a labeled record: the question, the gold chunk ids, and a reference answer written by a person from the source documents. Each run then adds the retrieved ids and the claim verdicts. The human work is done once, and every later run reuses it. This labeled set is often called a golden dataset.
Take questions from real production logs, not ones the team invents. Invented questions tend to be easy and use the documents' own words. The lesson suggests clustering logged questions and sampling across clusters. It also suggests adding some questions with no answer in the corpus, so refusals stay measurable. A set of tens to low hundreds of real, labeled questions is usually enough to locate failures. Freeze it, version it and run it on every change.
One trap: gold chunk ids point at one particular chunking. If you re-cut the documents, as covered under chunking strategies, map the labels again. Otherwise recall keeps printing numbers that mean nothing.
Recall@k and context precision are plain arithmetic once labels exist, so they can run on every commit. Faithfulness and relevance usually need an LLM judge, so check the judge against people. In the lesson's sample, judge and human agreed on 88 of 100 claim verdicts. The 12 disagreements show where the judge or its instructions fail.
The RAG Evaluation lesson has both runnable Python examples, the full decision tree and the 2x2 with its fixes. The Evaluating RAG and Agents lesson extends the same idea to agents.