Think of a student getting ready for an exam. For weeks she works through the same five practice papers. By the end she scores 95 out of 100 on every one of them. Is she ready? Partly. Some of that 95 is real skill. Some of it is that she has seen these exact questions many times, and she remembers where the tricky ones are. The real exam will have new questions. Her score on it will probably be a little lower, and nobody can know how much lower until she sits it.
A good teacher knows this. She writes the real exam separately, and she does not let the student see it while the student practises. The practice papers are for learning. The unseen exam is for finding out what the student can really do.

Every lesson in this chapter so far has measured prompts on the same 40 customer messages from lesson 1. Those 40 messages are my practice papers. I wrote the prompts with them in mind, I read their mistakes, and I chose which prompt to call "the best" by looking at their scores. So this lesson does what the teacher does: it scores three of the chapter's prompts on 40 new messages that no prompt had ever been scored on, and compares the two results.

Dev set. The cases you look at while you write, compare and choose a prompt. "Dev" is short for development. In this lesson the dev set is lesson 1's 40 messages, which I also call the old messages.
New set. Cases you keep aside and do not look at while you work. You score the prompt on them once, after it is final. People also call this a held-out set or a test set. Here it is 40 new customer messages.
Optimistic. A score is optimistic when it is higher than the score you will get on cases you have not seen yet. A score on the dev set tends to be optimistic, because the choices you made were made while looking at it.
Paired. Two prompts scored on the same cases. Because every message was answered by both prompts, you can compare them message by message: which messages one prompt fixed and which it broke. That is what the sign test from lesson 7 needs (it counts messages fixed against messages broken).
Unpaired. Two scores on different cases. The old set and the new set have no message in common, so you cannot match them up one by one. You can only compare the two totals, and that needs a different test.
As in lesson 1, an answer is usable when the whole reply, in lower case and without spaces at either end, is exactly the correct category word.

The lab uses three prompts that earlier lessons already measured, with not one word changed:
Each prompt was scored on two sets of 40 messages. The old set is lesson 1's 40, the dev set that every lesson in this chapter has used. The new set is 40 customer messages I wrote on 2026-09-28, 10 per category, each labelled with the category I judged right, and worded differently from lesson 1's messages and lesson 4's examples. The file that holds them records that they were written before any prompt was scored on them. Before the run I also reworded three of them to make them clearer; no score had been seen at that point. The lab checks only that none of the new messages is an exact copy of an old message or of an example; a few are close paraphrases of old ones, such as the toaster and the blender.
Every prompt went with all 80 messages to qwen2.5:3b and llama3.2:3b through Ollama, at temperature 0 (the model always takes its most likely next token), with room for up to 30 written tokens (the pieces of text a model reads and writes). Three prompts, 80 messages and two models make 480 calls. I designed the lab and wrote it down in the lab file's own description before it ran, and scoring is lesson 1's, unchanged.

Here is the whole result. On qwen2.5:3b, prompt C went from 36 usable answers on the old messages to 34 on the new ones, prompt D stayed at 40, and C+examples went from 39 to 36. On llama3.2:3b, prompt C went from 20 to 17, prompt D from 36 to 32, and C+examples from 33 to 31.
So of the six pairs of a model and a prompt, not one scored higher on the new messages. Five scored lower, by 2 to 4 answers, and one stayed exactly the same. Added up over the three prompts, qwen gave 115 usable answers of 120 on the old messages and 110 on the new ones, and llama gave 89 and 80.

The old scores are not new numbers. They are the same results earlier lessons reported, and the lab checked this: all 240 replies on the old messages (three prompts, 40 messages, two models) are identical, letter for letter, to the replies the same prompts gave in lessons 1, 4 and 9. So the old column is exactly what this chapter has been telling you, and the new column is what you would have seen on messages it had never met. Every reply in the lab also finished well inside the token limit: none was cut.

Prompt D deserves a closer look, because it is the prompt this chapter would recommend. Its context lines say where the edges between the categories are, and on qwen it was perfect in lesson 1.
On qwen2.5:3b, prompt D held completely: 40 of 40 on the old messages and 40 of 40 on the new ones. That is a good result, and it is worth saying plainly. The context lines did not only fit the 40 messages they were first tested on; on this model they also worked on 40 messages written later.
On llama3.2:3b, prompt D scored 36 on the old messages and 32 on the new ones. If you had tested prompt D only on the old 40, you would have told your team to expect about 36 right answers in 40 from llama. On new messages it gave 32. If a drop like that were real, 4 in 40 is 10 in 100, which you would notice in a product. The next slide shows why 40 messages cannot tell us whether it is real.
So the same prompt, on the same two sets of messages, held on one model and dropped on the other. With 40 messages we cannot say the two models really differ here (qwen was at 40 of 40 on both sets, with no room to fall), but it is one more reason to test on new cases with the model you will actually use.

Is llama's drop from 36 to 32 more than luck? This question needs more care than the comparisons in lessons 7 and 8.
In those lessons, two prompts were always scored on the same 40 messages. That made the comparison paired: for each message you could see whether the second prompt fixed it or broke it, and the sign test looked only at the messages that changed. Here, the old set and the new set are different messages. There is no message that appears in both, so there is nothing to pair. "Fixed" and "broken" have no meaning between two sets.
What you can compare is the two totals: 36 right of 40 in one group, 32 right of 40 in another. A standard test for this is Fisher's test. Its answer is a p value: the chance of a result this uneven if there were no real difference; below 0.05 is the usual line for "probably not luck". It asks: if both sets really had the same chance of a right answer, how often would two groups of 40 split as unevenly as 36 and 32, just by which messages happened to land in each group? For llama's prompt D the answer is p = 0.348, about one time in three. That is far above the usual line of 0.05, so this one drop cannot be told apart from luck.
The same test on the other five pairs gives p between 0.359 and 1.000. None of the six drops, taken alone, could be told apart from luck. What does carry some weight is the direction: none of the six went up. But be careful with that too. The six results share the same 80 messages, so they are not six separate pieces of evidence. If the new set happens to be a little harder than the old one, all six will move together. With 40 messages a set, a single held-out test can show you a warning sign, not prove a difference.

Within one set, the comparison is paired again, because all three prompts answered the same 40 new messages. So inside the new set you can use the sign test exactly as lesson 7 did.
On qwen, going from C to D on the new messages fixed 6 and broke none: p = 0.031, below 0.05. On the old messages the same change fixed 4 and broke none, p = 0.125, which could not be told from luck. So the new set gave clearer evidence than the old one that D is better than C on qwen, simply because C made more mistakes there for D to fix. One caution: this is one of twelve comparisons I ran after seeing the data, and among twelve, one p value near 0.03 can turn up by luck, so read it as supporting evidence, not proof. On llama, adding context or examples to C fixed 15 and 14 messages on the new set and broke none: p below 0.001, the same clear result as on the old set.
There is a second, smaller finding here. On both models and on both sets, the three prompts came out in the same order: D first, C+examples second, C last. Where the sign test could separate two prompts (C against D and against C+examples on llama, and C against D on qwen's new set), both sets pointed the same way. Where it could not, the order may be luck. So in this lab the old messages were a reasonable guide to which prompt was better, and a weaker guide to how high the score would be. You want both answers, so you need both sets.
One comparison did not transfer cleanly. On qwen's new messages, going from D to C+examples broke 4 and fixed none (p = 0.125); on the old messages it broke only 1. On llama, D against C+examples was 4 fixed and 5 broken on the new set, p = 1.000: the two prompts cannot be told apart there.

A score says how many messages went wrong, not which. So I read every wrong answer on the new set, and then counted them by the category that was right. Across all six runs (three prompts, two models), there were 36 wrong answers on the old messages and 50 on the new ones, out of 240 answers each.
The extra mistakes were not spread evenly. Account messages went wrong 26 times on the new set against 19 on the old one, and delivery messages 16 times against 7. Billing went slightly the other way, 8 against 10. No returns message was ever wrong, on either set, in any run: "returns" was often the wrong answer the models gave, but it was never the right answer they missed.
The mistakes also went in the same directions that earlier lessons found on the old messages. Qwen's mistakes with prompt C were five account messages answered "billing" and one delivery message answered "returns"; with C+examples, four account messages answered "billing". Llama's 23 mistakes with prompt C all answered "returns", the same habit lesson 4 found. So most of the new mistakes went in directions already seen on the old messages. Three small kinds were new: qwen with C answered one delivery message "returns", llama with C+examples answered two account messages "delivery", and llama with D answered one billing message "account".


Reading the account messages side by side, I noticed one simple difference. Five of the ten old account messages contain the word "account" itself: "How do I change the email address on my account?", "Please delete my account and all my data", and three more. Only one of the ten new ones does: "I want to add a second person to my account."
When a message contains the name of the category, a model has an easy clue. When it only describes a problem ("The app logs me out every few minutes"), the model has to know that logging in belongs to account, and prompts C and C+examples never say so. Prompt D does say it ("signing in, passwords, profile details, privacy and emails from us"), and on qwen, prompt D got every one of these messages right.
One possible reason the new account messages were harder is that fewer of them name their category. But I want to be clear about how I found this. I did not predict it; I noticed it after reading the results, and a pattern found that way is weak evidence. Ten messages per set is also very few. To test the idea properly, you would write many account messages in two versions, one with the word "account" and one without, before running anything, and compare them in pairs.
The practical point does not depend on the reason. When I wrote the old account messages, half of them named the category, perhaps because I had the names in my head. Real customers more often describe what happened to them. A test set written by the person who named the categories can be easier than the messages real customers send.

There is one more thing you should know before you trust prompt D's perfect score on qwen's new messages. I wrote prompt D's context lines. I also wrote the old messages, the new messages and every label. So the same idea of where the edges between the categories are went into the prompt and into both test sets.
You can see it in the new messages. "I keep getting text messages from you, please stop them" is labelled account, and the context line for account includes "emails from us", which is very close. "How do I download all the data you hold about me?" is labelled account, and that line includes "privacy". "Two of the plates in the set arrived broken" is labelled delivery, and the delivery line includes "arrived damaged". I did not write the new messages to match the context lines, but I could not forget the lines while writing them either.
This means the new set tests something narrower than it seems. It tests whether the prompts work on new wording of the same ideas. It does not test whether they work on a different person's idea of the categories. If a colleague had labelled these messages, they might have put the broken plates under returns, and prompt D's context line would then be pulling answers the wrong way. The fairest new set is one written and labelled by someone other than the person who wrote the prompt, or better still, real messages from real customers, labelled by the people who will act on them.

Why should a score on the cases you worked with be too high? None of this chapter's three prompts was changed to fix a particular old message; lesson 1 wrote its context lines before its first run. The optimism comes from choices that are easy to miss.
The first is selection. Whenever you try several prompts on the same cases and keep the best, you keep the one that happened to do best on those particular cases. Lesson 8 showed that seven wordings meaning the same thing spread from 32 to 39 on qwen. If you had tried all seven and kept the 39, part of that 39 would be real quality and part would be luck on those 40 messages. On new messages, the luck does not come with it. This chapter called prompt D its best prompt by looking at its scores on the old messages.
The second is shared habits. The person who writes the prompt and the test cases brings the same words and the same ideas to both, as the last two slides showed. The test cases end up closer to the prompt than real inputs would be.
The third is plain chance. Any 40 messages are a small sample. A new set of 40 might be a little easier or a little harder by luck. That is why one held-out set gives you a better estimate, not a certain one.
The cure for all three is the same order of work: do all your writing, comparing and choosing on the dev set; freeze the prompt; then score it once on the new set and report that number. If you read the new mistakes and change the prompt, the new set has become a dev set, and you need fresh cases to test the changed prompt.
This script sends four old messages and four new ones to qwen2.5:3b, with prompts C and D, and prints each answer and the usable count for each set.

Before you run this lab. It uses qwen2.5:3b, running in Ollama on your own computer. If you have not set that up yet, the lab setup guide shows how to install Ollama, download the model and check that everything works, on macOS, Windows or Linux. You can use a different model instead: the guide shows the one line to change, and your numbers will differ from the ones in this lesson.
"""Two prompts, scored on messages they were written around and on messages they never saw.
This one uses qwen2.5:3b. Pull it first (see the lab setup guide):
ollama pull qwen2.5:3b
python holdout_demo.py
"""
import json
import urllib.request
INSTRUCTION = "Put this customer message into one of these categories: billing, delivery, returns, account."
FORMAT = "Answer with the category word only, in lower case, and nothing else."
CONTEXT = ("billing: charges, payments, invoices, prices and refunds of money.\n"
"delivery: shipping, couriers, tracking and parcels that have not arrived or arrived damaged.\n"
"returns: sending an item back, exchanges and the return process.\n"
"account: signing in, passwords, profile details, privacy and emails from us.")
PROMPTS = {"C": [INSTRUCTION, FORMAT], "D": [INSTRUCTION, FORMAT, CONTEXT]}
DEV = [("I forgot my password and the reset email never comes.", "account"), # from lesson 1's 40
("I need a refund of the shipping charge you added by mistake.", "billing"),
("The courier left the box at the wrong address.", "delivery"),
("Can I exchange the blue shirt for a red one?", "returns")]
NEW = [("The app logs me out every few minutes.", "account"), # from the 40 new ones
("Two of the plates in the set arrived broken.", "delivery"),
("I keep getting text messages from you, please stop them.", "account"),
("Can I swap these boots for the same pair in brown?", "returns")]
def ask(text):
body = {"model": "qwen2.5:3b", "stream": False, "messages": [{"role": "user", "content": text}],
"options": {"temperature": 0, "num_predict": 30}}
req = urllib.request.Request("http://localhost:11434/api/chat", data=json.dumps(body).encode(),
headers={"Content-Type": "application/json"})
return json.loads(urllib.request.urlopen(req).read())["message"]["content"].strip()
def score(name, cases):
usable = 0
for msg, gold in cases:
reply = ask("\n\n".join(PROMPTS[name] + [f"Message: {msg}\nCategory:"]))
usable += reply.lower() == gold
print(f" {name} {msg[:44]:<45} {gold:<9} got {reply[:12]!r}")
return usable
for name in PROMPTS:
dev, new = score(name, DEV), score(name, NEW)
print(f"prompt {name}: usable {dev} of {len(DEV)} on dev, {new} of {len(NEW)} on new\n")

The measurement is scripts/labs/prompting/chapter_batch.py, mode holdout. It is one of three labs for this chapter that I designed together and wrote down, in the file's own description, before any of them ran. It imports lesson 1's 40 messages, its prompt parts and its scoring, lesson 4's example set, and the 40 new messages from holdout_messages.py, sends every prompt with every message, and stores every reply in one results file per model. It never reports a time, because other labs were sharing Ollama while it ran.
The report is a separate file, scripts/labs/prompting/holdout_report.py. It reads the two results files and never calls a model. It prints the old and new scores with the unpaired test between them, the paired sign tests between prompts inside each set, the mistakes on the new set, the wrong answers by category, the count of account messages that contain the word "account", and two checks: that no new message copies an old one or an example, and that the old replies repeat the earlier lessons word for word.
This box has no model. It holds the lab's real answers for all three prompts on both sets and both models, one letter per message for the category each reply named. It prints each prompt's old and new scores with the unpaired Fisher test, and two paired comparisons inside the new set.
It prints the same numbers as the lab: every change from old to new is 0 or a drop, and every Fisher p is 0.348 or more. The two paired comparisons show the two cases this lesson describes: qwen's C to D on the new set, 6 fixed and 0 broken, p = 0.031; and llama's D to C+examples, 4 fixed and 5 broken, p = 1.000, a tie. Try sign_test("qwen2.5:3b", "old", "C", "D") to see the same change on the old messages, or change a letter in a new-set line and watch the gap move.
The prompt parts. INSTRUCTION, FORMAT and CONTEXT are lesson 1's three strings, copied exactly. PROMPTS builds prompt C from the first two and prompt D from all three, so the only difference between the two prompts is the context lines.
Two lists, kept apart. DEV holds four of lesson 1's messages and NEW holds four of the new ones, each with the category I labelled it with. Keeping them in two separate lists is the whole idea of this lesson in two variables. In your own work, NEW would live in its own file that you do not open while you write the prompt.
The call. ask sends one prompt to qwen2.5:3b through Ollama's chat endpoint at temperature 0, with room for 30 tokens, the same limit as the lab, and returns the reply text without spaces at either end.
The score. score sends every case in a list with one prompt, prints each answer, and counts a reply as usable only when, in lower case, it equals the right category. That is lesson 1's usable check. The last loop scores each prompt on both lists and prints the two counts side by side, so the gap between old and new is the first thing you see.

Split before you start. When you collect and label test cases, put some of them aside at once, in a separate file. For example, keep a third of them for the new set. Do not read the new set while you write the prompt, and never pick examples from it.
Do all your choosing on the dev set. Try wordings, context lines and examples, and compare them with the sign test on the dev cases, as lessons 7 and 8 showed. The dev set is where you are allowed to look, change and look again.
Score the new set once, and report that number. When the prompt is final, score it on the new set. That is the number to tell your team. The dev score describes your working process, not what users will get.
If the new score is much lower, read the new mistakes. They may tell you what kind of case your dev set is missing, as the account messages may have here, or that the new set is simply harder. Add cases like them to the dev set, fix the prompt, and then write fresh new cases, because the old new set has now been looked at.
Add real cases over time. Hand-written cases are a start. Real inputs from real users are better: they are messier, and nobody wrote them with your categories in mind. Keep adding real, labelled cases to both sets.
Re-test after every prompt or model change. Lesson 8 showed that the same two wordings came top on both models, but below the top the order changed: the terse wording was middling on qwen and by far the worst on llama. Lesson 3 showed that moving the rules into the system message barely changed qwen under a polite push (38 against 40) but took llama from 22 to 5. Any change to the prompt or the model is a new prompt, and it needs a new score on the held-out cases.

It matters most when you report a number. If anyone will make a decision from your prompt's score, such as whether to ship it, give them the score on cases you did not tune on. The dev score here was higher on five of six model and prompt pairs, though no single drop could be told apart from luck.
It matters when you have tried many versions. The more prompts you compare on the same cases, the more the winner's score is flattered by luck on those cases. After a long search, as in lesson 8, a held-out score is the honest one.
It matters when the test cases and the prompt come from the same person. A colleague's cases, or real ones, will find the edges you did not think of. My own new set shared my idea of the categories with my context lines.
It matters less for choosing between two prompts. In this lab, the clear differences between prompts pointed the same way on the old and the new messages, on both models. If you only need to know which of two prompts is better, a good dev set with a paired test often tells you. You still need the new set to know how good the winner is.
It is not a reason to stop improving. A held-out set is not a place to hide mistakes; its job is to show them. When it shows a new kind of mistake, learn from it, move those cases into the dev set, and write new held-out cases.

This lab ran two small models, once each, at temperature 0, on one task. The new set is 40 messages that I wrote and labelled, the same person who wrote the old messages and prompt D's context lines. With 40 messages per set, no single drop from old to new could be told apart from luck by Fisher's test, and the six results share their messages, so their common direction is a warning sign rather than a proof.
One new set also cannot separate two explanations. The old scores may be optimistic because the chapter was built around those messages, or the new messages may simply be a little harder, for example because fewer account messages name their category. Both lead to the same advice, but this lab cannot say how much of the gap each one explains. A second new set, written by someone else, or real customer messages, would be the next test. What carries over is the method: a set of cases you never look at while writing, scored once, and compared with care.

Take a prompt from your own work that has a set of test cases. If you have been improving it against those cases, they are your dev set. Now write or collect 20 to 40 new cases, ideally real ones, and ask a colleague to label them without looking at your prompt. Put them in a separate file. Score the current prompt on them once, and compare the new score with your dev score. If the gap is large, read the new mistakes and ask what kind of case your dev set was missing. Then keep the habit: every time you change the prompt or the model, score the new version on cases it has never been tuned on, and report that number.

4 questions - Score 80% to pass
The chapter's prompts scored 36, 40 and 39 on qwen2.5:3b on lesson 1's 40 messages, and 34, 40 and 36 on 40 new ones. What does the score on the old messages tell you?
Llama's prompt D scored 36 of 40 on the old set and 32 of 40 on the new set. Why can you not use the sign test to compare these two scores?
Inside the new set, qwen's prompt C to prompt D fixed 6 messages and broke none, p = 0.031. What kind of comparison is this?
You read the mistakes on your held-out set and change the prompt to fix them. What should you do next?
Six new messages were wrong in 4 of the 6 runs, and none was wrong in more. Five of the six are account messages: the app that logs the customer out, the email that is "already registered", the unwanted text messages, the two-factor code going to an old phone, and the settings page that will not load. The sixth is "Two of the plates in the set arrived broken", a delivery message that four runs called "returns".
What do they share? Each of the five account messages describes a symptom, something that went wrong in the app or on the phone, and none of them names the part of the business it belongs to. The broken plates are a different kind of hard case: damage on arrival sits between delivery (the parcel was damaged) and returns (the customer may want to send it back). My label says delivery, and lesson 1's context line for delivery says "arrived damaged", but a reasonable person could argue for returns. I am describing what I see in six messages, not proving a cause.
This is a real run in VS Code's terminal.

Prompt C got 2 of the 4 old messages and 1 of the 4 new ones. Prompt D got all eight. I picked these eight messages because they show the changes, so do not read the small counts as a measurement; the lab's 80 messages are the measurement. What the script does show is the habit this lesson is about: two lists, kept apart, and every prompt scored on both.
On my laptop, all 16 answers matched the lab's stored replies for the same message and prompt. I checked this by comparing them with the results file, not by eye. The script leaves out the lab's seed (a number that fixes any random choices) and 4,096-token context (the most text the model reads at once) and uses Ollama's defaults; for these one-word replies that made no difference. On your computer, a different Ollama version or chip can change which of two almost equally likely tokens is chosen, so a small difference is possible.
I want you to know which parts came after I saw the data. The prompts, both sets of messages, their labels and the scoring were fixed before the run. The report was written afterwards, and so were its choices: which counts to print, the Fisher test, the paired tests, the mistake tables and the account-word count. The pattern on the account slide was found by reading the results, not predicted.

