Think about a school exam that decides who passes. A fair school does not let one teacher mark every paper alone. Two teachers mark the same papers, each in a separate room, and neither sees the other's marks. Afterwards someone puts the two sets of marks side by side. Where they are the same, the mark stands. Where they differ, someone reads the paper again.
Most of the time, the two teachers agree. When they do not, the reason is often not that one of them was careless. Very often the marking guide did not say what to do with that kind of answer. One teacher gave the point and the other did not, and both could defend their choice. The useful thing to do then is to add one sentence to the marking guide, so that the next hundred papers are marked the same way.

A set of examples for training a model is exactly like that pile of exam papers. Each example needs a right answer, and a person, or here a model, has to decide what the right answer is. If the answers are careless or unclear, a model trained on them learns the carelessness, and a test scored against them measures nothing useful. This lesson is about building the examples for the rest of the chapter, before any model is trained on them. I wrote the rules, had the messages written and marked twice, read every place the marks differed, fixed one rule, and checked the test messages had not leaked into the training messages.

Training set. The examples a model learns from. Each one is a customer message together with the answer we want, called its label.
Validation set. A smaller group of examples that the model never learns from. While training runs, it reports how wrong the model is on these too, so you can see whether it is learning something general or only memorising the training examples.
Test set. Examples used only once the model is finished, to score it. They must never be used to build anything, not the model and not the prompt, or the score stops meaning anything.
Annotator. Whoever gives each example its label. In industry this is often a person paid to read and label messages. The word labeller means the same thing.
Blind. An annotator is blind when they cannot see anyone else's labels while they work, like the two teachers in separate rooms.
Gold. A label we trust enough to train on and to score against. In this lesson, a field is gold only when two blind annotators gave it the same value.
Kappa. Short for Cohen's kappa, a number for how much two annotators agree beyond what chance would give. The agreement slide explains how it is worked out.
Leakage. When a test example, or something almost the same, is also in the training set. A model can then score well on it by remembering, not by understanding.
A boundary rule is one extra sentence in the rules for the messages that sit between two categories. This lesson ends up adding one.
Lesson 1 of this chapter showed that sorting a message into one of four categories was too easy to teach anything about : a simple lookup of the most similar labelled message did as well as a trained model. So from here on the task is harder. The model reads one customer message and writes one ticket, a small piece of JSON (a standard text format for data, with named fields in curly brackets) with five fields.

Before a single message was written, I wrote the rules for the five fields in one short file, the house rules. The same file is the guide the annotators read, and later it is the prompt the prompted models read. Each rule is there because two careful people could otherwise give two different answers.
Category is one of the same four as before: billing, delivery, returns or account. It decides which team gets the ticket. Wants is one of seven actions the customer asks for: a refund, a replacement, a cancellation, a change of details, stopping our messages, information, or other. If a message asks for two things, the rule says to choose the one asked for first, because a team needs one next step, not a list.
Order number is copied as digits only, with no "#", no word "order" and no spaces, or null (empty) when there is none. A phone number, a price or a date is never an order number. The rule is strict because a program will look the number up: "#4471" and "4471" must come out the same.

The obvious first idea was to reuse what I already had. Lesson 1 had 521 training messages for the four-way sorting task, plus the chapter's 80 test messages. I had two annotators label all 601 as tickets, blind, with training and test messages mixed together and each message given a meaningless id, so no annotator could tell which was which.
They agreed very well. Category agreed on all 601, order number on all 601, urgent on 598, item on 597, wants on 588, and all five fields together on 581. That looked like a good start, until I counted what the fields actually held.

The messages had been written for sorting, so almost none of them had the details a ticket is for. Only 10 of the 601 had an order number, and none of the 80 test messages had one. Annotator A marked 3 as urgent (annotator B marked 6). Only 3 asked to cancel something. An item was named in 106.
That is a trap for testing. A model that never reads the message, and always answers "no order number" and "not urgent", would match annotator A on the order number for 591 of the 601 messages and on urgency for 598. It would look almost perfect on those two fields while being unable to do the one thing those fields exist for.

A test set has to be able to tell a good model from a lazy one. These messages could not, so I did not use them as the main ticket data. New messages had to be written where the fields are really used. The 80 test messages stay in the chapter as a second, smaller ticket test, because the prompting chapter's results are on them, but the main test comes next.
I needed messages with real order numbers, real product names, real deadlines and every kind of request. I also wanted the test set to be harder than the training set in one specific way: written by someone else, for different customers. If the same writer wrote both, the test messages would share that writer's habits, and a model could score well by learning the habits.
All of these messages were written by an AI model, for this course, and so were all the labels in this lesson. I say it here and again wherever it matters, because it limits what every number means.
The first writer wrote the training pool: 640 messages, exactly 160 per category, spread over all seven kinds of request. The second writer, working separately and never shown the first writer's messages, wrote the test set: 140 messages, 35 per category, for a different customer base. I asked for people typing on a phone, people writing in English as a second language, people who explain at length, and people who write formal letters. Each writer also gave, for each message, the ticket they meant it to have. I call that the writer's intent and kept it as a third opinion, never as the answer.

The two sets use the fields in similar amounts, which is what I wanted: 311 of the 640 training messages have an order number and 77 of the 140 test messages do; an item is named in 444 and 97; 156 and 35 are urgent. The requests are spread out too. In the training pool, 118 ask for information, 100 for a refund, 93 for a change, 85 for a replacement, 82 to stop messages, 82 are other, and 80 ask to cancel.
The way they are written differs, and on purpose. I measured it with three simple checks that anyone can repeat on the files. The test messages are longer: a median (the middle value when all the lengths are put in order) of 20 words against 17, and 12 of the 140 are over 30 words against only 2 of the 640. More of them start with a small letter, as phone typing often does: 30 of 140 against 98 of 640. And 8 of the 140 start with "Dear", like a letter, against 5 of the 640.
Next, every one of the 780 messages, the 640 and the 140 together, was labelled twice, by two annotators, A and B. Three things kept the labelling honest.

First, the messages were shuffled together and renamed. Each got a meaningless id from k000 to k779, and a separate key file, which neither annotator saw, maps each id back to the real one. So an annotator could not tell a training message from a test message and could not treat the test more carefully.
Second, neither annotator saw the writer's ticket. They saw only the text of each message and the house rules. Third, neither saw the other's labels, like the two teachers in separate rooms. Only after both were finished were the two sets of labels matched up through the key.
The annotators were two separate runs of an AI model of the same family as the writers. That matters for the next slide: two copies of the same kind of thinking tend to agree more than two different people would. So every agreement number in this lesson is likely higher than two people would reach, not a promise of it. The same goes for the later steps: the third annotator's settling of the splits and the relabelling agreement are votes from the same model family too.
The simplest measure is to count, for each field, the messages where A and B gave the same value. I counted each field separately, because one overall number would hide which rule is weak.

The order number and urgent agreed on all 780 messages. Those two rules are mechanical: copy the digits, look for a deadline or one of the listed words. Item agreed on 771 and wants on 770. Category, which I expected to be the easiest field, agreed on only 743: 37 messages where A and B chose different categories. All five fields together agreed on 724 of the 780.
Counting is not quite enough, though. Think back to the first try, where almost no message was urgent. Two annotators who both always wrote "false" would agree on nearly every message, without reading any of them. So plain agreement looks good on a field where one answer is very common, even when nobody is reading. Cohen's kappa fixes that in two steps. First, it works out how often the two would agree just by chance, from their habits. For each answer, take how often A gives it times how often B gives it, and add these up. Second, it asks: of the agreement that chance does not explain, how much did they actually reach?

The sketch shows it on the 20 rows of this lesson's small script. On those rows A and B agree on the category 16 times out of 20, which is 0.80. Chance agreement from their habits is 0.265. So kappa is (0.80 minus 0.265) divided by (1 minus 0.265), which is 0.728. A kappa of 1.0 means they agreed on everything; 0 means no better than chance. There is no official line for "good enough". A common rule of thumb comes from Landis and Koch, who called 0.81 to 1.00 "almost perfect" agreement.
On all 780 messages, kappa was 0.937 for category, 0.985 for wants, 0.987 for item, and exactly 1.000 for order number and urgent. Those are high, and the warning from the last slide applies: two annotators of the same model family agreeing is weaker evidence than two people agreeing. The more useful fact is not the size of the numbers but where the disagreements sit, and that is the next slide.
When two annotators disagree, the first thing to do is not to count, but to read. I read all 37 messages where the category split.

They were not scattered. 33 of the 37 were the same disagreement: annotator A said returns and annotator B said billing. The other four were one each of four different pairs. Reading the 33 showed why. Almost all were a customer who has sent something back and wants their money: "sent back the jeans last week order 51130 pls refund me", or "I'd like a refund on order 58660. The towels I returned were all unused." 29 of the 33 asked for a refund.
The house rules said returns is "sending an item back, exchanges and the return process", and billing is "charges, payments, invoices, prices and refunds of money". A refund for a returned item is both. Annotator A read it as the last step of a return. Annotator B read it as a refund of money. Both followed the rules as written; the rules did not decide.

Even the writers were not consistent. For the 33 messages, the writer had meant returns 27 times and billing 6 times. The last example above, about a kettle refund "approved last Tuesday", does not even say the kettle was sent back; that one is honestly unclear. When the person who wrote a message and two careful readers cannot agree, the problem is the rule, not the readers.
The other two fields had fewer splits: 10 on wants and 9 on item. They were the same kind of thing. For wants, most were messages that ask a question and then a request only if the answer is no, such as "Has my payment for the heater been received? If not, please cancel the order." A said cancel, B said information. For item, most were how short the noun should be: "yoga mat" or "mat", "bar stool" or "stool". Those are rule gaps too, but small ones, and I did not change the rules for them. I dropped those messages from training instead, as the final slide shows.
For the 33 refunds, I added one boundary rule to the house rules:
A refund for an item the customer has already sent back, or is sending back, is returns, because the refund is the last step of the return. Money that is not tied to sending an item back (a wrong or double charge, an invoice, a price, the refund of a cancelled order or a missing item) is billing.
Then a third annotator, a fresh run that had not seen any earlier labels, relabelled the category only, with the updated rules, blind. It did not see just the 37 split messages. They were mixed with 40 messages where A and B had agreed, called controls, and all 77 were given new meaningless ids. The controls answer a simple question: does the third annotator agree with A and B where they already agreed? If it did not, its answers on the splits would not be worth much.

On all 37 splits, the third annotator chose either A's category or B's, so every split was settled. It sided with A 33 times and with B 4 times. Of the 33 returns-or-billing splits, it chose returns for 32 and billing for 1, the kettle message that never says the kettle went back. The rule for gold, at that point, was: for a split message, the third annotator's category counts only if it matches A or B.
On the 40 controls, it matched A and B on 38. I read the two it did not match, and they turned out to be the most useful two messages in this lesson.

The two controls meant that the labels I had were a mix. Most categories came from A and B working without the boundary rule; 37 came from the third annotator working with a rule that had a defect in it. The only clean way out was to fix the rule and then label the category again for every message, not only for the ones that had caused trouble.

First, I tightened the rule. The words "or a missing item" came out, and one sentence went in: a parcel that never arrived or arrived damaged stays delivery, even when the customer asks for their money back. The rule now settles only the question it was written for, returns against billing, and leaves the delivery definition alone.
Then two more annotators, D and E, fresh runs that had seen none of the earlier labels, labelled the category of all 780 messages under the tightened rules. They worked blind, on the same meaningless ids, with only the message text and the rules. From now on, a category is gold only where D and E agree. The other four fields keep their gold from A and B, because the rule change did not touch them.

D and E agreed on the category of 773 of the 780 messages, a kappa of 0.988. 14 labels in the old gold changed when these two fresh annotators relabelled under the tightened guide: 13 agreed labels plus 1 settled split. The 13 were labels A and B had agreed on under the original guide, before any boundary rule existed. The 1 is te108, the salad bowl message, where A said billing and B delivery and the third annotator had settled it. 10 of the 14 are in the training pool and 4 in the test set. D and E also disagreed on 7 messages, 3 training and 4 test, and those now have no category gold.
The last check is the one from lesson 1: leakage. The first writer never saw the test messages, but both writers wrote about the same shop in plain English, so some training messages were bound to say almost the same thing as a test message. A model that learns one of those could get the matching test message right by remembering it.
I embedded every training message and every test message, the 140 new ones and the chapter's 80 older ones, with nomic-embed-text, the model from lesson 1. An embedding is a list of numbers for a text, and texts that mean similar things get similar lists. For each training message I stored its similarity to the closest of the 220 test messages. Similarity here is : 1.0 means two texts point the same way, and lower means less alike.

No training message scored 0.90 or more against any test message; the highest was 0.893. 14 scored 0.85 or more, 65 scored 0.80 or more, and 184 scored 0.75 or more. The middle value was 0.72.
I used the same rule as lesson 1, chosen there before any model was trained: drop every training message with a similarity of 0.80 or more to any test message. I did not choose a new line for this data after seeing these scores. That dropped 65 training messages. 50 of them were closest to one of the 140 new test messages, 8 to one of the chapter's newer 40, and 7 to one of its older 40. By category they were 25 account, 18 returns, 14 billing and 8 delivery.

Now everything comes together. A training message is kept only if it passed the leak check and all five of its fields are gold.

Of the 640 training messages, 65 were dropped by the leak check. 17 more had a field that is not gold: 3 where D and E split on the category, 8 where A and B split on wants and 6 on item. 2 messages were in both groups, so 560 were kept; all five fields are gold on 623 of the 640 before the leak check. From the 560, a fixed random shuffle (seed 1, so the same shuffle every time) took 15 per category for the validation set, 60 in all, and the other 500 became the training set. The training set is not perfectly balanced: 129 delivery, 127 returns, 123 billing and 121 account, because the leak check and the splits took more from some categories than others. (Before the relabelling, the same steps had given 503 training rows; that version is kept but no lesson uses it.)

The test set is all 140 messages from the second writer. Test messages are never dropped for a split, because dropping the hard ones would make the test easier. Instead, the scoring has one rule: a field counts only where it is gold. 131 of the 140 test messages have all five fields gold. The other 9 each have one field that is not: the category on 4, where D and E split, wants on 2 and item on 3, where A and B split. For those 9, the other four fields are scored and the unclear one is left out. So 691 of the 700 test fields are scored. A message counts as a whole ticket right only if all five fields are gold and all five match.
The chapter's 80 older test messages stay as a second, smaller test, with the labels from the first try. All five fields agreed on 78 of those 80.
This script measures agreement the way this lesson did, on a small real sample. It carries 20 of the 780 messages with the tickets annotators A and B gave them in the first round, and prints, for each field, how often they agree, the chance agreement, and Cohen's kappa. Then it lists every place they split. The 20 rows are not a random sample. They are a sample chosen to contain splits: I picked 4 category splits, 2 wants splits and 1 item split so that the output has something to show, and 13 rows where the annotators agreed on everything, from a one-off draw that is not saved in the repo.

Before you run this lab. This script needs only Python 3; it calls no model and installs nothing. The lab's leak check, which you do not need for this script, uses nomic-embed-text running in Ollama on your own computer. If you want to try that part and have not set Ollama up yet, the lab setup guide shows how to install Ollama, download a model and check that everything works, on macOS, Windows or Linux. For the leak check you need only the model: ollama pull nomic-embed-text.
"""How much do two labellers agree? Per-field agreement and Cohen's kappa, on 20 real rows.
The rows are 20 of the 780 ticket messages from lesson 2 of the fine-tuning chapter, with the tickets
two annotators gave them blind. Each ticket is (category, wants, order_number, item, urgent).
Nothing to install and no model:
python agreement_demo.py
"""
from collections import Counter
FIELDS = ("category", "wants", "order_number", "item", "urgent")
# (id, message, annotator A's ticket, annotator B's ticket)
ROWS = [
('tr193', 'Order 70832, the lamp I returned was received on the 1st. I need the refund before my card statement closes on Friday.',
('returns', 'refund', '70832', 'lamp', True),
('billing', 'refund', '70832', 'lamp', True)),
('te113', 'sent back the jeans last week order 51130 pls refund me',
('returns', 'refund', '51130', 'jeans', False),
('billing', 'refund', '51130', 'jeans', False)),
('tr531', 'Please delete my saved card details from your system.',
('billing', 'other', None, None, False),
('account', 'other', None, None, False)),
('tr020', 'Cancel order 3895 please. Shipping the rug to Spain turned out to be £40 on top.',
('delivery', 'cancel', '3895', 'rug', False),
('billing', 'cancel', '3895', 'rug', False)),
('tr584', 'Has my payment of £85.20 for the heater been received? I need to know asap. If not, please cancel the order.',
('billing', 'cancel', None, 'heater', True),
('billing', 'information', None, 'heater', True)),
('te047', 'Is my account linked to my old work email, and if so, can you move it to my personal one?',
('account', 'change', None, None, False),
('account', 'information', None, None, False)),
('tr618', 'Refund order 88210 immediately. I have been charged for a yoga mat I cancelled in March.',
('billing', 'refund', '88210', 'yoga mat', True),
('billing', 'refund', '88210', 'mat', True)),
('tr122', "I'm still getting delivery update texts for the trainers on order 63004, which arrived weeks ago. Stop them please.",
('account', 'stop', '63004', 'trainers', False),
('account', 'stop', '63004', 'trainers', False)),
('tr515', 'order 52019 came in 4 seperate boxes for 4 small items. what a waste of packaging',
('delivery', 'other', '52019', None, False),
('delivery', 'other', '52019', None, False)),
('tr461', 'I got a kettle from you once and now you emial me every day. Please stop.',
('account', 'stop', None, 'kettle', False),
('account', 'stop', None, 'kettle', False)),
('tr006', 'Hello, my display name shows my full name on my blender review. Please change it to initials only.',
('account', 'change', None, 'blender', False),
('account', 'change', None, 'blender', False)),
('tr269', "Please change the collection address for the desk I'm returning on 92241 to my work address.",
('returns', 'change', '92241', 'desk', False),
('returns', 'change', '92241', 'desk', False)),
('tr527', 'Faulty toaster. I want an exchange please.',
('returns', 'replacement', None, 'toaster', False),
('returns', 'replacement', None, 'toaster', False)),
('tr382', 'Hello, is two-factor login available on your site? I need it switched on today if so.',
('account', 'information', None, None, True),
('account', 'information', None, None, True)),
('tr550', "Please resend #47318 today, the first dress went missing and it's for a wedding on Sunday.",
('delivery', 'replacement', '47318', 'dress', True),
('delivery', 'replacement', '47318', 'dress', True)),
('tr503', 'Order no. 11825 came with a box of broken biscuits. Could you send another box?',
('delivery', 'replacement', '11825', 'biscuit', False),
('delivery', 'replacement', '11825', 'biscuit', False)),
('te073', 'Please send a new lampshade for order 15520, mine arrived dented, and change the delivery address for the new one to my work.',
('delivery', 'replacement', '15520', 'lampshade', False),
('delivery', 'replacement', '15520', 'lampshade', False)),
('tr529', 'Do you do returns collection for the armchair on order 39402 or do I have to take it to the post office?',
('returns', 'information', '39402', 'armchair', False),
('returns', 'information', '39402', 'armchair', False)),
('te015', 'Unsubscribe me from the newsletter and delete my account while you are at it.',
('account', 'stop', None, None, False),
('account', 'stop', None, None, False)),
('tr377', 'Stop calling me at work. It must stop today, my manager has noticed.',
('account', 'stop', None, None, True),
('account', 'stop', None, None, True)),
]
# ---- the arithmetic
def kappa(a, b):
"""Cohen's kappa. Plain agreement flatters a field where almost every answer is the same: if 19 of 20
messages have no order number, two labellers who both write null every time agree 95% of the time
without reading anything. Kappa asks how much better than CHANCE they did. Chance agreement is worked
out from each labeller's own habits: how often A gives each answer times how often B gives it, added
up over the answers. Then kappa = (agreement - chance) / (1 - chance): 1.0 means perfect, 0 means no
better than chance. When both give the same single answer to every row, chance is 1 and kappa is undefined."""
n = len(a)
agree = sum(x == y for x, y in zip(a, b)) / n
count_a, count_b = Counter(a), Counter(b)
chance = sum(count_a[v] * count_b[v] for v in set(a) | set(b)) / (n * n)
if chance == 1:
return agree, chance, None
return agree, chance, (agree - chance) / (1 - chance)
print(f"{'field':<14}{'agree':>9}{'chance':>9}{'kappa':>8}")
for k, field in enumerate(FIELDS):
a = [str(row[2][k]) for row in ROWS] # str() so that None and "None" count as one answer
b = [str(row[3][k]) for row in ROWS]
agree, chance, kap = kappa(a, b)
same = sum(x == y for x, y in zip(a, b))
print(f"{field:<14}{same:>5} / {len(ROWS)}{chance:>9.3f}{'n/a' if kap is None else f'{kap:.3f}':>8}")
print("\nwhere they split:")
for rid, text, a, b in ROWS:
for k, field in enumerate(FIELDS):
if a[k] != b[k]:
print(f" {rid} {field}: A {a[k]}, B {b[k]} ({text[:60]})")

The report lives in scripts/labs/finetune/dataset_report.py. It reads only stored files: the writers' messages with their intended tickets, the two annotators' labels and the key that maps their meaningless ids back to real ones, the third annotator's labels and its key, the relabelling by D and E, the first and the current gold files, and the first and the current split. The default report uses only Python's standard library and calls no model. Every table in this lesson is printed by it, and a json mode writes the same numbers to results/ds-report.json, which is what the figures read.
It also checks the files against each other before it prints anything, and stops with an error if any check fails. It confirms that:
This box has no model. It holds 192 of the 780 messages: every message where A and B split on any field, every message the third annotator saw, every training message the leak check dropped, every category the relabelling changed or split on, and a random few that went straight through to training, validation or the test set. For each one you can see the writer's intended ticket, annotator A's, annotator B's, the third annotator's category where there is one, the category D and E gave it, and where the message ended up and why.
As it is, the box shows three messages. The first, tr193, is a returned lamp and a refund: A said returns and B said billing, the third annotator and later D and E all said returns, and then the message was dropped anyway, because it was too close to a test message about another returned lamp, at 0.8854. The second is tr156, the return-label refund: A and B agreed on billing, the third annotator said returns, D and E both said returns, and it went to training as returns. The third is te125, a polite request to cancel a garden bench before it is dispatched: the writer, A and B all said delivery, and D and E both said billing, which is the cancellation change from the relabelling slide. At the end, count() shows that of the 192 in the box, 7 are messages where D and E split, 63 went to training, 13 to validation, 36 are test messages and 80 were dropped.
Try find('cancel') to list messages that mention a cancellation, then show() with any id it prints. Try show('te108'), the salad bowl message, where the writer, A and B gave three different categories, and show('tr229'), a cancellation where D and E split.
The data. ROWS is a list of 20 real rows. Each row holds the message id, the message text, annotator A's ticket and annotator B's ticket. A ticket is a tuple (a fixed list) of the five field values in order: category, wants, order number, item and urgent. None is Python's word for null.
Counting agreement. For each field, the script takes the value at the same position from every row, once for A and once for B. It turns each value into text with str(), so that None from A and None from B are counted as the same answer. Agreement is the number of rows where the two are equal, divided by the number of rows.
Chance agreement. Counter counts how often each annotator uses each answer. In the 20 rows, A says account 7 times and B says it 8 times, so two annotators guessing with those habits would both say account on about 7/20 times 8/20 of the rows. The script adds that up over every answer either of them used: for category, (7 x 8 + 5 x 3 + 5 x 4 + 3 x 5) / 400 = 106 / 400 = 0.265, the chance agreement it prints.
Kappa. Kappa is agreement minus chance, divided by 1 minus chance: the share of the room above chance that the annotators actually used. When both annotators give the same single answer on every row, chance is 1 and there is no room above it, so kappa cannot be computed; the script prints n/a instead of dividing by zero. That case does not happen in these 20 rows, but it would if you picked only messages with no order number.
The splits. The last loop prints every row and field where A and B differ, with the start of the message. On your own data, this list is the part to read first.
To use it on your own labels, replace ROWS with your own rows, and with your own field names. The same few lines work for one field or for twenty.

Write the rules first. Before anyone labels anything, write down what each field means and what to do at its edges: what counts as urgent, how to write a product name, what to do when a message asks for two things. Give a reason for each rule. Your labellers will read it, and later your prompt may be built from it.
Label every example twice, blind. Two labellers, each working alone, with training and test examples mixed and given meaningless ids. If you can only afford two labels on part of the data, do it on the test set first, because the test set is what every later decision rests on. Use people if the model will serve people. If you use a model to help, as I did, say so, and expect its agreement to be higher than two people would reach.
Measure agreement per field. Count agreement and compute kappa for each field separately. A single overall number hides the one weak rule.
Read the splits, fix the rule, and relabel everything it touches. Group the disagreements and read them. If many split the same way, the rule has a gap: add a sentence. Then check the new sentence against your other definitions, word by word, because a new rule can contradict an old one, as mine did. Finally, relabel the field for every example, not only the ones that split, with fresh labellers. A new rule can change answers on examples nobody flagged: here the full relabel changed 13 agreed labels and 1 settled split, and it also exposed a case the rules do not cover at all. Relabelling only the splits, with agreed controls mixed in, is still a useful cheap check, because the controls are what showed me the rule reached further.
Check for leaks. Embed your training and test examples and drop training examples that are almost the same as a test example. Read the closest pairs to choose the line, and keep the same line once chosen.
Freeze the test set, and write down what you changed. Fix the sets before you train or score anything, and keep a short record of every decision and when it was made.
It is worth it whenever a number will decide something. If you will choose between a prompt and a fine-tune, or between two models, by their scores, the scores are only as good as the labels. A few wrong labels in a test set of 140 can move a score by more than the difference you are trying to measure.
It is worth it for any field that needs judgement. Fields like category, wants and item depend on how people read a rule, and people read rules differently. The only way to find where is to label twice and compare. Here the two fields with near-mechanical rules, order number and urgent, agreed on every message, and the field I expected to be easy, category, had the most splits.
It is worth it when your test data comes from a different place than your training data. Real users rarely write like the people who wrote your examples. A test set written by someone else, for other users, tells you more than a random slice of your training data.
Two labels on everything may be more than you need. If the labels are simple and mechanical, like copying an order number, and two labellers agree on the first hundred, labelling the rest once is a reasonable saving. Keep double labelling for the test set and for the fields that split.
It does not replace real data. Everything in this lesson was written and labelled with an AI model's help. That is fine for teaching how to check a data set, but for a real product, messages from real customers and labels from people who know the business are worth far more than any amount of care with written messages.

Written by an AI model. Both writers were AI models, working for this course. The test writer was asked to write like different customers, and the simple checks show it did write differently, but real customers are messier than any written set, and they write about things no writer thought of.
Labelled by the same model family. The annotators A, B, the third one, D and E were all runs of AI models from the same family as the writers. Two copies of one way of thinking probably agree more than two people do. The agreement numbers here, and the kappas, are likely higher than two people would reach, and that includes the third annotator's votes and the relabelling's 0.988. A real team would use people and measure agreement in exactly the same way, and should expect lower numbers.
Rules changed after seeing the disagreements. The boundary rule was added after reading the 37 splits, and tightened after reading the third annotator's two misses. Both changes happened before any training, and the category was then relabelled for all 780. Two things remain. The rules have a hole: nothing says where a plain request to cancel an unshipped order belongs. A and B said delivery for 6 such messages, D and E said billing, and those 6 keep billing, which is a reading of an unclear rule, not a decision. And D and E still split on 7 messages, one of them a cancellation; those 7 have no category gold.
The test labels changed after prompted models were scored. Prompted models had already been run on the 140 test messages when the category was relabelled. The change was triggered by the two controls, not by any score. Still, by this lesson's own question, "was the test set fixed before any model was scored on it?", this test set does not fully pass, and you should know that.
140 test messages is small. One test message is less than 1% of the set, and a difference of one or two messages between two models will usually not be told apart from luck, as lesson 1 showed with the sign test. The chapter's 80 older messages add a second check, but they have far fewer order numbers and deadlines.

The next time someone hands you a labelled data set, or you are about to build one, ask the five questions in the list. Who wrote the examples, and did the test examples come from somewhere else? Was each example labelled twice, blind, and was agreement counted per field? What did the labellers disagree on, and were rules changed because of it, and when? Were near copies of test examples removed from training, and at what line? Was the test set fixed before any model was scored on it?
If you cannot answer those questions, you do not yet know what any score on that data means. And if the answer to the third is "yes, a rule changed", ask one more: was everything relabelled afterwards, or only the examples that caused the change? Run the small script from this lesson on a sample of your own labels: twenty rows labelled by two people is enough to find a weak rule.

The next lesson uses these frozen sets for the first time. It gives the house rules to prompted models of several sizes and scores them field by field on the 140 test messages, to find where a prompt alone starts to slip.
4 questions - Score 80% to pass
On the first try, the old sorting messages had an order number in only 10 of 601. Why was that a problem for testing a ticket model?
33 of the 37 category splits were annotator A saying returns and annotator B saying billing, almost all refunds for an item sent back. What was the right response?
Urgent agreed on all 780 messages, and most messages are not urgent. Why is Cohen's kappa more useful than plain agreement for a field like this?
Why were the 140 test messages written by a second writer, for a different kind of customer, instead of taken from the same pool as the training messages?
Item is the product as a short common noun: lower case, singular, with no colour, size, brand or number. "Two blue ceramic mugs" becomes "mug". "My new Sony headphones" becomes "headphones", which stays plural because some nouns are always plural in English. If a team wants to count complaints per product, "mug" and "mugs" and "blue mug" must all be the same thing. Urgent is true only when the customer states a deadline or uses one of a short list of words: urgent, asap, today, tonight, right now or immediately. An angry message is not urgent by itself. Without that line, every angry customer would jump the queue.
These rules were the first draft. The category rules turned out to have a gap, and then my fix for it turned out to have one of its own. Later slides show how the data found both.

Reading them side by side shows it better than counts. "The courier leave the package in the rain and the book is all wet now" is the English of someone who learned it later in life. "charged twice for order 55120 the blue wool scarf. pls refund the extra asap" is a phone. Those are the kinds of messages a model trained on tidier writing might miss, and that is what the test is for.
One more thing you should know about the training pool. In 46 of the 640 training messages, one common word was misspelled by a small script, for example "charged" became "chraged". The writer marked them and I recorded which ones. So some spelling mistakes in the training data were added mechanically, not typed by a person in a hurry. I have no such record for the test messages, so I cannot tell you how their mistakes were made.
Message tr156 asks for a refund of the cost of a return label. A and B, working without the boundary rule, said billing; the third annotator, with it, said returns. That is not the third annotator's mistake: the money is tied to sending an item back, and the new rule says that is returns. So the rule seemed to reach a label that A and B had agreed on. It did not only settle the 37 splits; it could change messages nobody had flagged.
Message tr416 showed a second, worse problem. It is an order that never arrived, and the customer wants their money back. A and B said delivery. The third annotator said billing, and it was following my words: the rule said "the refund of a cancelled order or a missing item is billing". But the delivery definition says delivery covers "parcels that have not arrived". My rule contradicted my own definition. The third annotator had simply obeyed the newer sentence. The next slide is what I did about both problems.
The 14 fall into four kinds. One caution before reading them: D and E are different runs from A and B. A change can come from the tightened rule, or simply from two new annotators reading the same message differently, and with this data I cannot separate the two.

Delivery to billing, 6 messages. All 6 ask to cancel an order that has not shipped yet, such as "if order 40127 hasn't left yet please cancel it immediately", and none asks for money back. A and B said delivery for all 6; D and E said billing for all 6. The rules do not cover this case. The boundary rule makes only "the refund of a cancelled order" billing, and nothing in the rules says where a plain request to cancel belongs. So billing here is how D and E read an unclear rule, not what the rules say. The clearest sign is tr229, "Please cancel the bookcase before it ships today", almost the same request as the test message te058 ("Please cancel my order before it ships today"): it is one of the 7 new splits, with D saying delivery and E billing. This is a hole the rules still have. I am leaving it in the data, and telling you, rather than relabelling a third time: the 6 messages keep D and E's billing, and tr229 has no category gold. A label is only as settled as the rule behind it.
Billing to returns, 4 messages. These fit the boundary rule: money tied to an item sent back, such as "my card got hit for £38 for the jeans ... but i returned them three weeks ago", and tr156, the return label. Billing to delivery, 3 messages. These fit the tightened sentence: an order "never dispatched", or a bowl that arrived cracked, where the customer wants a refund, now stays delivery. Delivery to account, 1 message: te052, a customer asking us to delete a photo the driver took of their front door. Neither rule touches this one. Privacy was in the account definition from the start, so this change is two sets of annotators reading the same message differently.
As for the two controls: tr156 is now returns, and tr416, the parcel that never arrived, is delivery again, which is what the definition said all along.
Changing a rule after seeing where annotators disagree is normal labelling work. It is how rules get good. The lesson here is about how far the change has to reach: when you change a rule, relabel everything the rule could touch, not only the messages that exposed the problem. Relabelling everything is also what showed me the cancellation hole: it appeared only because D and E read all 780.
When this happened matters too, so here it is plainly. No model had been trained on any of this data. Prompted models had already been run on the 140 test messages and on the chapter's 80 older ones, using the house rules of the time as their prompt; the reason for the change was the two controls above, not any model result. Their stored replies are being scored again against the new labels, and no result from them appears in this lesson.
The closest pairs are the same request reworded: a copy of personal data, a password reset. The third pair is interesting for this task. Both messages are about a returned lamp and a refund before Friday, but the order numbers differ, 70832 and 66390. For ticket extraction a model still has to copy the right digits from the message in front of it, so a near copy does not hand it the whole answer. It does hand it the category, what the customer wants, the item and the urgency, which is why the rule still applies.
The line is a line. The message about unsubscribing scored exactly 0.8000 and was dropped; a message about a crushed box and a broken lamp scored 0.7986 and was kept, with no one judging its meaning. The stored scores were checked: the lab found each training message's closest test message again, and every score matched the stored one to the fourth decimal place.
These sets are now frozen: fixed, and not to be changed because of anything a model does on them later. Every lesson from here on trains on the 500, watches its loss (a number for how wrong it is) on the 60, and is scored on the 140 and the 80.
This is a real run in VS Code's terminal.

Look at urgent: 20 of 20 agree, but chance agreement is already 0.580, because most messages are not urgent. Kappa is still 1.000, because they agreed on every row, including the urgent ones. Category agrees on only 16 of 20, because I put four splits in, and chance there is lower, 0.265, since four categories are used more evenly. Its kappa, 0.728, is the lowest here. On the full 780 messages category's kappa was 0.937; this sample was chosen to contain splits, so its numbers are lower than the whole set's and should not be quoted as the lesson's result.
I checked the script against the stored data in code, not by eye: the lab's demo mode reads the script's 20 rows, confirms every message text and every label is the same as in the stored annotation files, and computes agreement, chance and kappa for the same 20 ids from those files. All five fields matched the script's output exactly.
data/ticket4_changes.jsonTwo modes do more than read. The pairs mode, which also uses NumPy, embeds all 640 training messages and the 220 test messages again with nomic-embed-text (the only model call in the lab, and only for ), finds which test message each training message is closest to, and checks the score against the stored one: all 640 matched, and 65 were at 0.80 or more, as stored. The typos mode recorded once which training messages had the scripted misspelling, from the first writer's working files. A demo mode checks the student script, and a box mode writes the playground on the next slide.
What came after I saw the data: the choice to write new messages came after counting the empty fields in the first try. The boundary rule came after reading the 37 splits, and the third annotator after that. The tightened rule and the relabelling by D and E came after reading the third annotator's two misses on the controls. The 0.80 leak rule was carried over unchanged from lesson 1. No model was trained on any of this data before the sets were fixed.

FIELDSThe leak line catches copies, not cousins. Training messages just under 0.80 can still be close rewordings of a test message, as the kept lamp message at 0.7986 shows. That closeness helps any model that learns from them.