Fine Tuning

Building a Training Set You Can Trust: House Rules, Two Blind Labellers and a Leak Check

0 of 21 complete

0%

Contents

Back|Fine TuningBuilding a Training Set You Can Trust: House Rules, Two Blind Labellers and a Leak Check
1/21
73 min left
Prerequisites
Before You Fine-Tune: Try the Cheapest Baseline Firstrequired
Related Topics
Inter-Annotator Agreement: Before You Trust Your Own LabelsLLM Evaluation and Error AnalysisWhich Labels You Check AgainstLLM Evaluation and Error AnalysisThirty Rows, Ten Questions, One Leaky SplitLLM Evaluation and Error AnalysisWhat an Embedding Is: Meaning as a List of NumbersTokens and EmbeddingsSearching by Meaning: Real Semantic Search, in Six LanguagesTokens and Embeddings
1 of 21

Two Teachers, One Pile of Exam Papers

Think about a school exam that decides who passes. A fair school does not let one teacher mark every paper alone. Two teachers mark the same papers, each in a separate room, and neither sees the other's marks. Afterwards someone puts the two sets of marks side by side. Where they are the same, the mark stands. Where they differ, someone reads the paper again.

Most of the time, the two teachers agree. When they do not, the reason is often not that one of them was careless. Very often the marking guide did not say what to do with that kind of answer. One teacher gave the point and the other did not, and both could defend their choice. The useful thing to do then is to add one sentence to the marking guide, so that the next hundred papers are marked the same way.

An illustration of a room with one long wooden table. At the left end a man in glasses reads a sheet of paper and writes on a small note; at the right end a woman looks at her phone, with a blank sheet of paper in front of her; a plant sits in the middle between them. Headed one table, two people, each working alone, titled label it twice, and read where you differ. Beneath: 780 messages, two blind annotators. All five fields agreed on 724. Order number and urgent: 780 of 780. Category: 743.

A set of examples for training a model is exactly like that pile of exam papers. Each example needs a right answer, and a person, or here a model, has to decide what the right answer is. If the answers are careless or unclear, a model trained on them learns the carelessness, and a test scored against them measures nothing useful. This lesson is about building the examples for the rest of the chapter, before any model is trained on them. I wrote the rules, had the messages written and marked twice, read every place the marks differed, fixed one rule, and checked the test messages had not leaked into the training messages.

Eight Words for This Lesson

A hand-drawn list headed eight words for this lesson, titled the vocabulary of a labelled data set. Training set: the examples a model learns from. Validation set: examples kept out of training; training reports its loss (how wrong it is) on them. Test set: examples used only to score the finished model, never to build it. Annotator: whoever gives each example its right answer, its label. Blind: labelling without seeing anyone else's labels. Gold: a label both annotators agree on; only gold trains or scores. Kappa: agreement after taking away what chance alone would give. Leakage: a test example, or a near copy of one, inside the training set. Beneath: boundary rule, a rule for the cases that sit between two categories.

Training set. The examples a model learns from. Each one is a customer message together with the answer we want, called its label.

Validation set. A smaller group of examples that the model never learns from. While training runs, it reports how wrong the model is on these too, so you can see whether it is learning something general or only memorising the training examples.

Test set. Examples used only once the model is finished, to score it. They must never be used to build anything, not the model and not the prompt, or the score stops meaning anything.

Annotator. Whoever gives each example its label. In industry this is often a person paid to read and label messages. The word labeller means the same thing.

Blind. An annotator is blind when they cannot see anyone else's labels while they work, like the two teachers in separate rooms.

Gold. A label we trust enough to train on and to score against. In this lesson, a field is gold only when two blind annotators gave it the same value.

Kappa. Short for Cohen's kappa, a number for how much two annotators agree beyond what chance would give. The agreement slide explains how it is worked out.

Leakage. When a test example, or something almost the same, is also in the training set. A model can then score well on it by remembering, not by understanding.

A boundary rule is one extra sentence in the rules for the messages that sit between two categories. This lesson ends up adding one.

The Ticket and Its House Rules

Lesson 1 of this chapter showed that sorting a message into one of four categories was too easy to teach anything about : a simple lookup of the most similar labelled message did as well as a trained model. So from here on the task is harder. The model reads one customer message and writes one ticket, a small piece of JSON (a standard text format for data, with named fields in curly brackets) with five fields.

An editorial frame headed one customer message in, one ticket out, titled five fields, each with a house rule. A zone labelled the message holds: Please resend #47318 today, the first dress went missing and it's for a wedding on Sunday. A zone labelled the ticket holds five boxes. Category: delivery, the team that handles it. Wants: replacement, the first thing asked for. Order_number: 47318, digits only, no #. Item: dress, a short common noun. Urgent: true, the message says today. Beneath: training row tr550; both annotators gave this ticket.

Before a single message was written, I wrote the rules for the five fields in one short file, the house rules. The same file is the guide the annotators read, and later it is the prompt the prompted models read. Each rule is there because two careful people could otherwise give two different answers.

Category is one of the same four as before: billing, delivery, returns or account. It decides which team gets the ticket. Wants is one of seven actions the customer asks for: a refund, a replacement, a cancellation, a change of details, stopping our messages, information, or other. If a message asks for two things, the rule says to choose the one asked for first, because a team needs one next step, not a list.

Order number is copied as digits only, with no "#", no word "order" and no spaces, or null (empty) when there is none. A phone number, a price or a date is never an order number. The rule is strict because a program will look the number up: "#4471" and "4471" must come out the same.

A table headed the house rules, and why each one exists, titled a rule for every place two people could differ. Category: one of four; a boundary rule for refunds. It decides which team gets the ticket. Wants: one of seven; the first action asked for. A team needs one next step, not a list. Order_number: digits only; null if none; a phone number is never one. A program looks it up. Item: lower case, singular, no colour, size or brand. Counts per product must add up. Urgent: true only for a stated deadline or words like today or asap. Anger is not urgency. Beneath: without the rule, two careful people give two different right answers.

A First Try on the Old Messages

The obvious first idea was to reuse what I already had. Lesson 1 had 521 training messages for the four-way sorting task, plus the chapter's 80 test messages. I had two annotators label all 601 as tickets, blind, with training and test messages mixed together and each message given a meaningless id, so no annotator could tell which was which.

They agreed very well. Category agreed on all 601, order number on all 601, urgent on 598, item on 597, wants on 588, and all five fields together on 581. That looked like a good start, until I counted what the fields actually held.

A bar chart headed the first try: the 601 sorting messages, labelled as tickets, titled an order number in 10 of 601. Four bars, counting messages as annotator A filled each field, on a scale up to 600: order number, a sliver near 10; urgent true and wants cancel, slivers near 3; item named, a bar near 100. Beneath: order number 10, urgent 3, cancel 3, item 106. None of the 80 test messages had an order number.

The messages had been written for sorting, so almost none of them had the details a ticket is for. Only 10 of the 601 had an order number, and none of the 80 test messages had one. Annotator A marked 3 as urgent (annotator B marked 6). Only 3 asked to cancel something. An item was named in 106.

That is a trap for testing. A model that never reads the message, and always answers "no order number" and "not urgent", would match annotator A on the order number for 591 of the 601 messages and on urgency for 598. It would look almost perfect on those two fields while being unable to do the one thing those fields exist for.

Two panels headed a model that never reads the message, titled empty answers would score almost full marks. Always null order: 591 / 601, matches annotator A. Always urgent false: 598 / 601, matches annotator A. Beneath: so I wrote new messages where the fields are really used.

A test set has to be able to tell a good model from a lazy one. These messages could not, so I did not use them as the main ticket data. New messages had to be written where the fields are really used. The 80 test messages stay in the chapter as a second, smaller ticket test, because the prompting chapter's results are on them, but the main test comes next.

Two Writers, Two Kinds of Customer

I needed messages with real order numbers, real product names, real deadlines and every kind of request. I also wanted the test set to be harder than the training set in one specific way: written by someone else, for different customers. If the same writer wrote both, the test messages would share that writer's habits, and a model could score well by learning the habits.

All of these messages were written by an AI model, for this course, and so were all the labels in this lesson. I say it here and again wherever it matters, because it limits what every number means.

The first writer wrote the training pool: 640 messages, exactly 160 per category, spread over all seven kinds of request. The second writer, working separately and never shown the first writer's messages, wrote the test set: 140 messages, 35 per category, for a different customer base. I asked for people typing on a phone, people writing in English as a second language, people who explain at length, and people who write formal letters. Each writer also gave, for each message, the ticket they meant it to have. I call that the writer's intent and kept it as a third opinion, never as the answer.

A bar chart headed writer 1: 640 training messages; writer 2: 140 test messages, titled same fields, a different kind of customer. Paired bars show the share of each writer's messages, in percent, training writer first and test writer second: order number about 49 and 55; item named about 69 and 69; urgent about 24 and 25; lower case start about 15 and 21; 30+ words about 0 and 9; Dear about 1 and 6. Beneath: over 30 words, 2 of 640 against 12 of 140. Median length 17 words against 20.

The two sets use the fields in similar amounts, which is what I wanted: 311 of the 640 training messages have an order number and 77 of the 140 test messages do; an item is named in 444 and 97; 156 and 35 are urgent. The requests are spread out too. In the training pool, 118 ask for information, 100 for a refund, 93 for a change, 85 for a replacement, 82 to stop messages, 82 are other, and 80 ask to cancel.

The way they are written differs, and on purpose. I measured it with three simple checks that anyone can repeat on the files. The test messages are longer: a median (the middle value when all the lengths are put in order) of 20 words against 17, and 12 of the 140 are over 30 words against only 2 of the 640. More of them start with a small letter, as phone typing often does: 30 of 140 against 98 of 640. And 8 of the 140 start with "Dear", like a letter, against 5 of the 640.

Two Annotators Who Never Saw Each Other's Work

Next, every one of the 780 messages, the 640 and the 140 together, was labelled twice, by two annotators, A and B. Three things kept the labelling honest.

A sequence diagram with three columns: the key, annotator A, annotator B. Step one, the key sends annotator A 780 texts, ids k000 to k779. Step two, the key sends annotator B the same, separately. Step three, annotator A sends back 780 tickets. Step four, annotator B sends back 780 tickets. Step five, the key turns them back to real ids. Beneath: training and test messages were mixed, so no annotator knew which was which. Neither saw the writer's ticket or the other's labels.

First, the messages were shuffled together and renamed. Each got a meaningless id from k000 to k779, and a separate key file, which neither annotator saw, maps each id back to the real one. So an annotator could not tell a training message from a test message and could not treat the test more carefully.

Second, neither annotator saw the writer's ticket. They saw only the text of each message and the house rules. Third, neither saw the other's labels, like the two teachers in separate rooms. Only after both were finished were the two sets of labels matched up through the key.

The annotators were two separate runs of an AI model of the same family as the writers. That matters for the next slide: two copies of the same kind of thinking tend to agree more than two different people would. So every agreement number in this lesson is likely higher than two people would reach, not a promise of it. The same goes for the later steps: the third annotator's settling of the splits and the relabelling agreement are votes from the same model family too.

How Often They Agreed, Field by Field

The simplest measure is to count, for each field, the messages where A and B gave the same value. I counted each field separately, because one overall number would hide which rule is weak.

A table headed annotator A against annotator B, field by field, titled all five fields agreed on 724 of 780. Category: 610 of 640 training, 133 of 140 test; kappa 0.937. Wants: 632 of 640 training, 138 of 140 test; kappa 0.985. Order_number: 640 of 640 training, 140 of 140 test; kappa 1.000. Item: 634 of 640 training, 137 of 140 test; kappa 0.987. Urgent: 640 of 640 training, 140 of 140 test; kappa 1.000. All five: 596 of 640 training, 128 of 140 test. Beneath: category had the most splits, 37 of 780.

The order number and urgent agreed on all 780 messages. Those two rules are mechanical: copy the digits, look for a deadline or one of the listed words. Item agreed on 771 and wants on 770. Category, which I expected to be the easiest field, agreed on only 743: 37 messages where A and B chose different categories. All five fields together agreed on 724 of the 780.

Counting is not quite enough, though. Think back to the first try, where almost no message was urgent. Two annotators who both always wrote "false" would agree on nearly every message, without reading any of them. So plain agreement looks good on a field where one answer is very common, even when nobody is reading. Cohen's kappa fixes that in two steps. First, it works out how often the two would agree just by chance, from their habits. For each answer, take how often A gives it times how often B gives it, and add these up. Second, it asks: of the agreement that chance does not explain, how much did they actually reach?

A hand-drawn sketch of four boxes joined by arrows, headed sketched: kappa on the demo's 20 rows, category, titled agree 16 of 20, kappa 0.728. They agree on 16 of 20: 0.80. Chance, from each one's own habits: 0.265. Room above chance: 1 - 0.265. Kappa: (0.80 - 0.265) / (1 - 0.265) = 0.728. Beneath: 1.0 is perfect agreement; 0 is no better than chance.

The sketch shows it on the 20 rows of this lesson's small script. On those rows A and B agree on the category 16 times out of 20, which is 0.80. Chance agreement from their habits is 0.265. So kappa is (0.80 minus 0.265) divided by (1 minus 0.265), which is 0.728. A kappa of 1.0 means they agreed on everything; 0 means no better than chance. There is no official line for "good enough". A common rule of thumb comes from Landis and Koch, who called 0.81 to 1.00 "almost perfect" agreement.

On all 780 messages, kappa was 0.937 for category, 0.985 for wants, 0.987 for item, and exactly 1.000 for order number and urgent. Those are high, and the warning from the last slide applies: two annotators of the same model family agreeing is weaker evidence than two people agreeing. The more useful fact is not the size of the numbers but where the disagreements sit, and that is the next slide.

Reading the 37 Disagreements

When two annotators disagree, the first thing to do is not to count, but to read. I read all 37 messages where the category split.

Five isometric columns headed the 37 category splits, by pair, titled 33 were returns against billing. The first column is tall: A returns, B billing, 33. Four flat ones follow, each 1: A delivery, B returns; A billing, B account; A billing, B delivery; A delivery, B billing. Beneath: height is messages. 29 of the 33 ask for a refund; the writer meant returns for 27.

They were not scattered. 33 of the 37 were the same disagreement: annotator A said returns and annotator B said billing. The other four were one each of four different pairs. Reading the 33 showed why. Almost all were a customer who has sent something back and wants their money: "sent back the jeans last week order 51130 pls refund me", or "I'd like a refund on order 58660. The towels I returned were all unused." 29 of the 33 asked for a refund.

The house rules said returns is "sending an item back, exchanges and the return process", and billing is "charges, payments, invoices, prices and refunds of money". A refund for a returned item is both. Annotator A read it as the last step of a return. Annotator B read it as a refund of money. Both followed the rules as written; the rules did not decide.

A two-column page headed returns or billing? A refund for an item sent back, titled the guide did not say. Message, then A, B and the writer. te113: sent back the jeans last week order 51130 pls refund me: A returns, B billing; writer returns. tr407: I'd like a refund on order 58660. The towels I returned were all unused: A returns, B billing; writer returns. tr300: How long do refunds normally take to reach a credit card? I sent the tablet from order 39816 back on Monday: A returns, B billing; writer billing. tr433: When will my refund for teh kettle on order 20114 show up? It was approved last Tuesday: A returns, B billing; writer billing. Beneath: refund is money, billing. The item went back, returns.

Even the writers were not consistent. For the 33 messages, the writer had meant returns 27 times and billing 6 times. The last example above, about a kettle refund "approved last Tuesday", does not even say the kettle was sent back; that one is honestly unclear. When the person who wrote a message and two careful readers cannot agree, the problem is the rule, not the readers.

The other two fields had fewer splits: 10 on wants and 9 on item. They were the same kind of thing. For wants, most were messages that ask a question and then a request only if the answer is no, such as "Has my payment for the heater been received? If not, please cancel the order." A said cancel, B said information. For item, most were how short the noun should be: "yoga mat" or "mat", "bar stool" or "stool". Those are rule gaps too, but small ones, and I did not change the rules for them. I dropped those messages from training instead, as the final slide shows.

One Rule Added, and a Third Annotator

For the 33 refunds, I added one boundary rule to the house rules:

A refund for an item the customer has already sent back, or is sending back, is returns, because the refund is the last step of the return. Money that is not tied to sending an item back (a wrong or double charge, an invoice, a price, the refund of a cancelled order or a missing item) is billing.

Then a third annotator, a fresh run that had not seen any earlier labels, relabelled the category only, with the updated rules, blind. It did not see just the 37 split messages. They were mixed with 40 messages where A and B had agreed, called controls, and all 77 were given new meaningless ids. The controls answer a simple question: does the third annotator agree with A and B where they already agreed? If it did not, its answers on the splits would not be worth much.

Two panels headed one boundary rule added; a third annotator, blind, category only, titled every split resolved; 38 of 40 controls matched. The 37 splits: 37 / 37, sided with A 33, with B 4. 40 agreed controls: 38 / 40, matched A and B. Beneath: the 77 were mixed together, so the third annotator could not tell a split from a control.

On all 37 splits, the third annotator chose either A's category or B's, so every split was settled. It sided with A 33 times and with B 4 times. Of the 33 returns-or-billing splits, it chose returns for 32 and billing for 1, the kettle message that never says the kettle went back. The rule for gold, at that point, was: for a split message, the third annotator's category counts only if it matches A or B.

On the 40 controls, it matched A and B on 38. I read the two it did not match, and they turned out to be the most useful two messages in this lesson.

A two-column page headed where the third annotator went its own way, titled two controls that showed the rule reached further. tr156 (control): The return label for the toaster on 81104 cost me £4.50. Can I have that refunded?: A and B billing; third returns; writer returns; later D and E returns. tr416 (control): Order 28410 never arrived and I don't want a replacement sweater. Refund me: A and B delivery; third billing; writer delivery; later D and E delivery. tr433 (split): When will my refund for teh kettle on order 20114 show up? It was approved last Tuesday: A returns, B billing; third billing; writer billing; later D and E billing. te108 (split): Please refund the cracked salad bowl and cancel the other bowl in the same order, #61047: A billing, B delivery; third billing; writer returns; later D and E delivery. Beneath: the first rule seemed to reach labels A and B had agreed on. In that round, a split took the third label if it matched A or B.

Fixing the Rule Again, and Relabelling Everything

The two controls meant that the labels I had were a mix. Most categories came from A and B working without the boundary rule; 37 came from the third annotator working with a rule that had a defect in it. The only clean way out was to fix the rule and then label the category again for every message, not only for the ones that had caused trouble.

A table headed the rule had a second gap, titled its own words fought the delivery definition. Delivery: parcels that have not arrived or arrived damaged. First rule: the refund of a cancelled order or a missing item is billing. tr416: Order 28410 never arrived and I don't want a replacement sweater. Refund me. A and B: delivery. Third: billing. Tightened: a parcel that never arrived or arrived damaged stays delivery, even when the customer wants money back. Beneath: then two fresh annotators relabelled the category of all 780.

First, I tightened the rule. The words "or a missing item" came out, and one sentence went in: a parcel that never arrived or arrived damaged stays delivery, even when the customer asks for their money back. The rule now settles only the question it was written for, returns against billing, and leaves the delivery definition alone.

Then two more annotators, D and E, fresh runs that had seen none of the earlier labels, labelled the category of all 780 messages under the tightened rules. They worked blind, on the same meaningless ids, with only the message text and the rules. From now on, a category is gold only where D and E agree. The other four fields keep their gold from A and B, because the rule change did not touch them.

A bar chart headed D and E, blind, the tightened rule, all 780 messages, titled 13 agreed labels and 1 settled split changed. Four bars, counting messages by old category to new category: delivery to billing 6, billing to returns 4, billing to delivery 3, delivery to account 1. Beneath: D and E agreed on 773 of 780, kappa 0.988. 7 new splits. Changed: 10 training, 4 test. All 6 delivery to billing ask to cancel; none asks for money back.

D and E agreed on the category of 773 of the 780 messages, a kappa of 0.988. 14 labels in the old gold changed when these two fresh annotators relabelled under the tightened guide: 13 agreed labels plus 1 settled split. The 13 were labels A and B had agreed on under the original guide, before any boundary rule existed. The 1 is te108, the salad bowl message, where A said billing and B delivery and the third annotator had settled it. 10 of the 14 are in the training pool and 4 in the test set. D and E also disagreed on 7 messages, 3 training and 4 test, and those now have no category gold.

Keeping the Test Set Clean

The last check is the one from lesson 1: leakage. The first writer never saw the test messages, but both writers wrote about the same shop in plain English, so some training messages were bound to say almost the same thing as a test message. A model that learns one of those could get the matching test message right by remembering it.

I embedded every training message and every test message, the 140 new ones and the chapter's 80 older ones, with nomic-embed-text, the model from lesson 1. An embedding is a list of numbers for a text, and texts that mean similar things get similar lists. For each training message I stored its similarity to the closest of the 220 test messages. Similarity here is : 1.0 means two texts point the same way, and lower means less alike.

A bar chart headed each training message against its closest test message, of 220, titled 65 of 640 too close, dropped. Bars count training messages by their similarity to the closest test message; each bar covers 0.05, and its label is where it starts, from 0.50 to 0.85: the tallest bars are at 0.65 and 0.70, near 180 and 190; a dashed line just before the 0.80 bar is labelled 0.80, dropped from here; the 0.80 bar is near 50 and the 0.85 bar near 14. Beneath: median 0.72, highest 0.893. At 0.90 or more: 0; 0.85: 14; 0.80: 65.

No training message scored 0.90 or more against any test message; the highest was 0.893. 14 scored 0.85 or more, 65 scored 0.80 or more, and 184 scored 0.75 or more. The middle value was 0.72.

I used the same rule as lesson 1, chosen there before any model was trained: drop every training message with a similarity of 0.80 or more to any test message. I did not choose a new line for this data after seeing these scores. That dropped 65 training messages. 50 of them were closest to one of the 140 new test messages, 8 to one of the chapter's newer 40, and 7 to one of its older 40. By category they were 25 account, 18 returns, 14 billing and 8 delivery.

A two-column page headed training message, and its closest test message, titled close in meaning, not always in the ticket. 0.8927, dropped: I'd like a copy of all the personal data you hold about me; closest, Under data protection law I request a copy of all personal data you hold about me. Please confirm receipt of this request, ticket test. 0.8890, dropped: I think someone knows my password. Please reset it immediately; closest, Someone may have my password. Please reset it right now and log me out of all devices, ticket test. 0.8854, dropped: Order 70832, the lamp I returned was received on the 1st. I need the refund before my card statement closes on Friday; closest, I need the refund for order no. 66390 before Friday, my card bill is due. It was for the grey metal lamp you already collected, ticket test. 0.8000, dropped: unsubscribe me from everything. I get 4 emails a day from you and it's ridiculous; closest, I unsubscribed three times and your newsletter is STILL in my inbox every morning. Take me off the list, ticket test. 0.7986, kept: Order 30119 arrived with the box crushed and the lamp inside is broken. Can you send a new one?; closest, The box arrived crushed and the two tall glasses inside are broken. Could you send them again? Order #56011, ticket test. Beneath: score, nomic-embed-text similarity. Of the 65 dropped, 50 were closest to a ticket test message.

The Sets Every Later Lesson Uses

Now everything comes together. A training message is kept only if it passed the leak check and all five of its fields are gold.

A flowchart headed from written messages to the frozen sets, titled 500 to train, 60 to validate, 140 to test. Writer 1: 640 messages, which branches to drop 65, 0.80 or more to a test message, and drop 17, a field that is not gold; both lead to kept: 560, 2 were in both drops. Kept branches to valid: 60, 15 per category, and train: 500. Separately, writer 2: 140 messages leads to test: 140, 131 fully gold. Beneath: no model was trained before these sets were fixed.

Of the 640 training messages, 65 were dropped by the leak check. 17 more had a field that is not gold: 3 where D and E split on the category, 8 where A and B split on wants and 6 on item. 2 messages were in both groups, so 560 were kept; all five fields are gold on 623 of the 640 before the leak check. From the 560, a fixed random shuffle (seed 1, so the same shuffle every time) took 15 per category for the validation set, 60 in all, and the other 500 became the training set. The training set is not perfectly balanced: 129 delivery, 127 returns, 123 billing and 121 account, because the leak check and the splits took more from some categories than others. (Before the relabelling, the same steps had given 503 training rows; that version is kept but no lesson uses it.)

A table headed the sets every later lesson uses, titled frozen before any model trained on them. Train: 500, billing 123, delivery 129, returns 127, account 121. Valid: 60, 15 per category. Test: 140 from writer 2; 131 fully gold, 691 of 700 fields scored. Second test: the chapter's 80 sorting messages; all five fields agreed on 78 of 80. Beneath: a field that is not gold is left out of the score, never guessed.

The test set is all 140 messages from the second writer. Test messages are never dropped for a split, because dropping the hard ones would make the test easier. Instead, the scoring has one rule: a field counts only where it is gold. 131 of the 140 test messages have all five fields gold. The other 9 each have one field that is not: the category on 4, where D and E split, wants on 2 and item on 3, where A and B split. For those 9, the other four fields are scored and the unclear one is left out. So 691 of the 700 test fields are scored. A message counts as a whole ticket right only if all five fields are gold and all five match.

The chapter's 80 older test messages stay as a second, smaller test, with the labels from the first try. All five fields agreed on 78 of those 80.

Try It Yourself

This script measures agreement the way this lesson did, on a small real sample. It carries 20 of the 780 messages with the tickets annotators A and B gave them in the first round, and prints, for each field, how often they agree, the chance agreement, and Cohen's kappa. Then it lists every place they split. The 20 rows are not a random sample. They are a sample chosen to contain splits: I picked 4 category splits, 2 wants splits and 1 item split so that the output has something to show, and 13 rows where the annotators agreed on everything, from a one-off draw that is not saved in the repo.

A real screenshot of VS Code with agreement_demo.py open, showing the top of the file: the docstring and the start of the ROWS list, each row an id, a message and the two annotators' tickets. The rest of the file is in the code box on this slide. Beneath: copy it from the box on the slide.

Before you run this lab. This script needs only Python 3; it calls no model and installs nothing. The lab's leak check, which you do not need for this script, uses nomic-embed-text running in Ollama on your own computer. If you want to try that part and have not set Ollama up yet, the lab setup guide shows how to install Ollama, download a model and check that everything works, on macOS, Windows or Linux. For the leak check you need only the model: ollama pull nomic-embed-text.

"""How much do two labellers agree? Per-field agreement and Cohen's kappa, on 20 real rows.

The rows are 20 of the 780 ticket messages from lesson 2 of the fine-tuning chapter, with the tickets
two annotators gave them blind. Each ticket is (category, wants, order_number, item, urgent).
Nothing to install and no model:
    python agreement_demo.py
"""
from collections import Counter

FIELDS = ("category", "wants", "order_number", "item", "urgent")

# (id, message, annotator A's ticket, annotator B's ticket)
ROWS = [
    ('tr193', 'Order 70832, the lamp I returned was received on the 1st. I need the refund before my card statement closes on Friday.',
     ('returns', 'refund', '70832', 'lamp', True),
     ('billing', 'refund', '70832', 'lamp', True)),
    ('te113', 'sent back the jeans last week order 51130 pls refund me',
     ('returns', 'refund', '51130', 'jeans', False),
     ('billing', 'refund', '51130', 'jeans', False)),
    ('tr531', 'Please delete my saved card details from your system.',
     ('billing', 'other', None, None, False),
     ('account', 'other', None, None, False)),
    ('tr020', 'Cancel order 3895 please. Shipping the rug to Spain turned out to be £40 on top.',
     ('delivery', 'cancel', '3895', 'rug', False),
     ('billing', 'cancel', '3895', 'rug', False)),
    ('tr584', 'Has my payment of £85.20 for the heater been received? I need to know asap. If not, please cancel the order.',
     ('billing', 'cancel', None, 'heater', True),
     ('billing', 'information', None, 'heater', True)),
    ('te047', 'Is my account linked to my old work email, and if so, can you move it to my personal one?',
     ('account', 'change', None, None, False),
     ('account', 'information', None, None, False)),
    ('tr618', 'Refund order 88210 immediately. I have been charged for a yoga mat I cancelled in March.',
     ('billing', 'refund', '88210', 'yoga mat', True),
     ('billing', 'refund', '88210', 'mat', True)),
    ('tr122', "I'm still getting delivery update texts for the trainers on order 63004, which arrived weeks ago. Stop them please.",
     ('account', 'stop', '63004', 'trainers', False),
     ('account', 'stop', '63004', 'trainers', False)),
    ('tr515', 'order 52019 came in 4 seperate boxes for 4 small items. what a waste of packaging',
     ('delivery', 'other', '52019', None, False),
     ('delivery', 'other', '52019', None, False)),
    ('tr461', 'I got a kettle from you once and now you emial me every day. Please stop.',
     ('account', 'stop', None, 'kettle', False),
     ('account', 'stop', None, 'kettle', False)),
    ('tr006', 'Hello, my display name shows my full name on my blender review. Please change it to initials only.',
     ('account', 'change', None, 'blender', False),
     ('account', 'change', None, 'blender', False)),
    ('tr269', "Please change the collection address for the desk I'm returning on 92241 to my work address.",
     ('returns', 'change', '92241', 'desk', False),
     ('returns', 'change', '92241', 'desk', False)),
    ('tr527', 'Faulty toaster. I want an exchange please.',
     ('returns', 'replacement', None, 'toaster', False),
     ('returns', 'replacement', None, 'toaster', False)),
    ('tr382', 'Hello, is two-factor login available on your site? I need it switched on today if so.',
     ('account', 'information', None, None, True),
     ('account', 'information', None, None, True)),
    ('tr550', "Please resend #47318 today, the first dress went missing and it's for a wedding on Sunday.",
     ('delivery', 'replacement', '47318', 'dress', True),
     ('delivery', 'replacement', '47318', 'dress', True)),
    ('tr503', 'Order no. 11825 came with a box of broken biscuits. Could you send another box?',
     ('delivery', 'replacement', '11825', 'biscuit', False),
     ('delivery', 'replacement', '11825', 'biscuit', False)),
    ('te073', 'Please send a new lampshade for order 15520, mine arrived dented, and change the delivery address for the new one to my work.',
     ('delivery', 'replacement', '15520', 'lampshade', False),
     ('delivery', 'replacement', '15520', 'lampshade', False)),
    ('tr529', 'Do you do returns collection for the armchair on order 39402 or do I have to take it to the post office?',
     ('returns', 'information', '39402', 'armchair', False),
     ('returns', 'information', '39402', 'armchair', False)),
    ('te015', 'Unsubscribe me from the newsletter and delete my account while you are at it.',
     ('account', 'stop', None, None, False),
     ('account', 'stop', None, None, False)),
    ('tr377', 'Stop calling me at work. It must stop today, my manager has noticed.',
     ('account', 'stop', None, None, True),
     ('account', 'stop', None, None, True)),
]

# ---- the arithmetic


def kappa(a, b):
    """Cohen's kappa. Plain agreement flatters a field where almost every answer is the same: if 19 of 20
    messages have no order number, two labellers who both write null every time agree 95% of the time
    without reading anything. Kappa asks how much better than CHANCE they did. Chance agreement is worked
    out from each labeller's own habits: how often A gives each answer times how often B gives it, added
    up over the answers. Then kappa = (agreement - chance) / (1 - chance): 1.0 means perfect, 0 means no
    better than chance. When both give the same single answer to every row, chance is 1 and kappa is undefined."""
    n = len(a)
    agree = sum(x == y for x, y in zip(a, b)) / n
    count_a, count_b = Counter(a), Counter(b)
    chance = sum(count_a[v] * count_b[v] for v in set(a) | set(b)) / (n * n)
    if chance == 1:
        return agree, chance, None
    return agree, chance, (agree - chance) / (1 - chance)


print(f"{'field':<14}{'agree':>9}{'chance':>9}{'kappa':>8}")
for k, field in enumerate(FIELDS):
    a = [str(row[2][k]) for row in ROWS]          # str() so that None and "None" count as one answer
    b = [str(row[3][k]) for row in ROWS]
    agree, chance, kap = kappa(a, b)
    same = sum(x == y for x, y in zip(a, b))
    print(f"{field:<14}{same:>5} / {len(ROWS)}{chance:>9.3f}{'n/a' if kap is None else f'{kap:.3f}':>8}")

print("\nwhere they split:")
for rid, text, a, b in ROWS:
    for k, field in enumerate(FIELDS):
        if a[k] != b[k]:
            print(f"  {rid} {field}: A {a[k]}, B {b[k]}  ({text[:60]})")

The Lab Report

A real terminal recording of python dataset_report.py, in seven numbered sections. 1, the first attempt on 601 sorting messages: an order number in only 10; always answering empty matches annotator A on 591 order numbers and 598 urgent flags. 2, the two writers side by side, train 640 and test 140, with each field's counts, lengths, and 46 scripted misspellings in training. 3, A against B on all 780: category 743, kappa 0.937; wants 770; order number 780; item 771; urgent 780; all five 724; 37 category splits, 33 of them returns against billing. 4, the third annotator: 37 of 37 splits settled; 38 of 40 controls matched, missing tr156 and tr416. 5, the relabel by D and E: 773 of 780, kappa 0.988, 7 splits; 14 labels changed: delivery to billing 6, billing to returns 4, billing to delivery 3, delivery to account 1. 6, the leak check: 0, 14, 65 and 184 at 0.90, 0.85, 0.80 and 0.75; highest 0.8927; every stored score reproduced. 7, the final sets: 560 kept, train 500 (503 before), valid 60, test 131 of 140 fully gold (135 before), 691 of 700 fields scored. Beneath: the lab's own report. It calls no model.

The report lives in scripts/labs/finetune/dataset_report.py. It reads only stored files: the writers' messages with their intended tickets, the two annotators' labels and the key that maps their meaningless ids back to real ones, the third annotator's labels and its key, the relabelling by D and E, the first and the current gold files, and the first and the current split. The default report uses only Python's standard library and calls no model. Every table in this lesson is printed by it, and a json mode writes the same numbers to results/ds-report.json, which is what the figures read.

It also checks the files against each other before it prints anything, and stops with an error if any check fails. It confirms that:

  • the merged label file holds exactly what each annotator wrote;
  • the third annotator saw exactly the 37 split messages and the 40 controls;
  • the current gold is exactly D and E's agreed category, with every other field unchanged;
  • the 14 changed labels are the ones stored in ;

Pick a Message and Follow It

This box has no model. It holds 192 of the 780 messages: every message where A and B split on any field, every message the third annotator saw, every training message the leak check dropped, every category the relabelling changed or split on, and a random few that went straight through to training, validation or the test set. For each one you can see the writer's intended ticket, annotator A's, annotator B's, the third annotator's category where there is one, the category D and E gave it, and where the message ended up and why.

As it is, the box shows three messages. The first, tr193, is a returned lamp and a refund: A said returns and B said billing, the third annotator and later D and E all said returns, and then the message was dropped anyway, because it was too close to a test message about another returned lamp, at 0.8854. The second is tr156, the return-label refund: A and B agreed on billing, the third annotator said returns, D and E both said returns, and it went to training as returns. The third is te125, a polite request to cancel a garden bench before it is dispatched: the writer, A and B all said delivery, and D and E both said billing, which is the cancellation change from the relabelling slide. At the end, count() shows that of the 192 in the box, 7 are messages where D and E split, 63 went to training, 13 to validation, 36 are test messages and 80 were dropped.

Try find('cancel') to list messages that mention a cancellation, then show() with any id it prints. Try show('te108'), the salad bowl message, where the writer, A and B gave three different categories, and show('tr229'), a cancellation where D and E split.

The Code, Part by Part

The data. ROWS is a list of 20 real rows. Each row holds the message id, the message text, annotator A's ticket and annotator B's ticket. A ticket is a tuple (a fixed list) of the five field values in order: category, wants, order number, item and urgent. None is Python's word for null.

Counting agreement. For each field, the script takes the value at the same position from every row, once for A and once for B. It turns each value into text with str(), so that None from A and None from B are counted as the same answer. Agreement is the number of rows where the two are equal, divided by the number of rows.

Chance agreement. Counter counts how often each annotator uses each answer. In the 20 rows, A says account 7 times and B says it 8 times, so two annotators guessing with those habits would both say account on about 7/20 times 8/20 of the rows. The script adds that up over every answer either of them used: for category, (7 x 8 + 5 x 3 + 5 x 4 + 3 x 5) / 400 = 106 / 400 = 0.265, the chance agreement it prints.

Kappa. Kappa is agreement minus chance, divided by 1 minus chance: the share of the room above chance that the annotators actually used. When both annotators give the same single answer on every row, chance is 1 and there is no room above it, so kappa cannot be computed; the script prints n/a instead of dividing by zero. That case does not happen in these 20 rows, but it would if you picked only messages with no order number.

The splits. The last loop prints every row and field where A and B differ, with the start of the message. On your own data, this list is the part to read first.

To use it on your own labels, replace ROWS with your own rows, and with your own field names. The same few lines work for one field or for twenty.

How to Build a Training Set You Can Trust

A hand-sketched column of six boxes joined by arrows, headed building your own set, titled six steps, in this order. 1, write the rules first. 2, label twice, blind. 3, count agreement per field. 4, fix the rule; relabel all it touches. 5, drop near copies of tests. 6, freeze the test set. Beneath: change rules before training, and say what you changed.

Write the rules first. Before anyone labels anything, write down what each field means and what to do at its edges: what counts as urgent, how to write a product name, what to do when a message asks for two things. Give a reason for each rule. Your labellers will read it, and later your prompt may be built from it.

Label every example twice, blind. Two labellers, each working alone, with training and test examples mixed and given meaningless ids. If you can only afford two labels on part of the data, do it on the test set first, because the test set is what every later decision rests on. Use people if the model will serve people. If you use a model to help, as I did, say so, and expect its agreement to be higher than two people would reach.

Measure agreement per field. Count agreement and compute kappa for each field separately. A single overall number hides the one weak rule.

Read the splits, fix the rule, and relabel everything it touches. Group the disagreements and read them. If many split the same way, the rule has a gap: add a sentence. Then check the new sentence against your other definitions, word by word, because a new rule can contradict an old one, as mine did. Finally, relabel the field for every example, not only the ones that split, with fresh labellers. A new rule can change answers on examples nobody flagged: here the full relabel changed 13 agreed labels and 1 settled split, and it also exposed a case the rules do not cover at all. Relabelling only the splits, with agreed controls mixed in, is still a useful cheap check, because the controls are what showed me the rule reached further.

Check for leaks. Embed your training and test examples and drop training examples that are almost the same as a test example. Read the closest pairs to choose the line, and keep the same line once chosen.

Freeze the test set, and write down what you changed. Fix the sets before you train or score anything, and keep a short record of every decision and when it was made.

When This Much Care Is Worth It, and When It Is Not

It is worth it whenever a number will decide something. If you will choose between a prompt and a fine-tune, or between two models, by their scores, the scores are only as good as the labels. A few wrong labels in a test set of 140 can move a score by more than the difference you are trying to measure.

It is worth it for any field that needs judgement. Fields like category, wants and item depend on how people read a rule, and people read rules differently. The only way to find where is to label twice and compare. Here the two fields with near-mechanical rules, order number and urgent, agreed on every message, and the field I expected to be easy, category, had the most splits.

It is worth it when your test data comes from a different place than your training data. Real users rarely write like the people who wrote your examples. A test set written by someone else, for other users, tells you more than a random slice of your training data.

Two labels on everything may be more than you need. If the labels are simple and mechanical, like copying an order number, and two labellers agree on the first hundred, labelling the rest once is a reasonable saving. Keep double labelling for the test set and for the fields that split.

It does not replace real data. Everything in this lesson was written and labelled with an AI model's help. That is fine for teaching how to check a data set, but for a real product, messages from real customers and labels from people who know the business are worth far more than any amount of care with written messages.

What This Data Set Is Not

A two-column page headed read before you trust these labels, titled what this data set is, and what it is not. It is: written by an AI model, for the course; labelled by the same AI model family; 140 test messages; rules fixed before any training. It is not: not real customer messages; not people, agreement likely higher than theirs; not large, each test message is under 1%; not fixed from the start, one rule, tightened once. A fifth row: it is test labels final after the relabel; it is not frozen before scoring, prompts ran first.

Written by an AI model. Both writers were AI models, working for this course. The test writer was asked to write like different customers, and the simple checks show it did write differently, but real customers are messier than any written set, and they write about things no writer thought of.

Labelled by the same model family. The annotators A, B, the third one, D and E were all runs of AI models from the same family as the writers. Two copies of one way of thinking probably agree more than two people do. The agreement numbers here, and the kappas, are likely higher than two people would reach, and that includes the third annotator's votes and the relabelling's 0.988. A real team would use people and measure agreement in exactly the same way, and should expect lower numbers.

Rules changed after seeing the disagreements. The boundary rule was added after reading the 37 splits, and tightened after reading the third annotator's two misses. Both changes happened before any training, and the category was then relabelled for all 780. Two things remain. The rules have a hole: nothing says where a plain request to cancel an unshipped order belongs. A and B said delivery for 6 such messages, D and E said billing, and those 6 keep billing, which is a reading of an unclear rule, not a decision. And D and E still split on 7 messages, one of them a cancellation; those 7 have no category gold.

The test labels changed after prompted models were scored. Prompted models had already been run on the 140 test messages when the category was relabelled. The change was triggered by the two controls, not by any score. Still, by this lesson's own question, "was the test set fixed before any model was scored on it?", this test set does not fully pass, and you should know that.

140 test messages is small. One test message is less than 1% of the set, and a difference of one or two messages between two models will usually not be told apart from luck, as lesson 1 showed with the sign test. The chapter's 80 older messages add a second check, but they have far fewer order numbers and deadlines.

What to Do Next

A hand-drawn list headed for a data set someone hands you, titled five questions before you trust it. Who wrote it?: real customers, staff, or a model; the test set from a different source. Labelled twice?: by two labellers, blind, with agreement counted per field. What split?: read the disagreements; were rules changed, and when? Leaks?: near copies of test cases removed, with the line stated. Frozen?: the test set fixed before models were scored on it, or every later change stated. Beneath: a set that cannot answer these is a guess, not a measurement.

The next time someone hands you a labelled data set, or you are about to build one, ask the five questions in the list. Who wrote the examples, and did the test examples come from somewhere else? Was each example labelled twice, blind, and was agreement counted per field? What did the labellers disagree on, and were rules changed because of it, and when? Were near copies of test examples removed from training, and at what line? Was the test set fixed before any model was scored on it?

If you cannot answer those questions, you do not yet know what any score on that data means. And if the answer to the third is "yes, a rule changed", ask one more: was everything relabelled afterwards, or only the examples that caused the change? Run the small script from this lesson on a sample of your own labels: twenty rows labelled by two people is enough to find a weak rule.

A closing card headed to keep, titled change a rule, relabel all it touches. In large type: 14 of 780. Beneath: labels in the old gold changed when two fresh annotators relabelled every message: 13 agreed labels and 1 settled split. Then: fix the rule, then relabel everything it could touch.

The next lesson uses these frozen sets for the first time. It gives the house rules to prompted models of several sizes and scores them field by field on the 140 test messages, to find where a prompt alone starts to slip.

Knowledge Check

Knowledge Check

4 questions - Score 80% to pass

Q1

On the first try, the old sorting messages had an order number in only 10 of 601. Why was that a problem for testing a ticket model?

Q2

33 of the 37 category splits were annotator A saying returns and annotator B saying billing, almost all refunds for an item sent back. What was the right response?

Q3

Urgent agreed on all 780 messages, and most messages are not urgent. Why is Cohen's kappa more useful than plain agreement for a field like this?

Q4

Why were the 140 test messages written by a second writer, for a different kind of customer, instead of taken from the same pool as the training messages?

Item is the product as a short common noun: lower case, singular, with no colour, size, brand or number. "Two blue ceramic mugs" becomes "mug". "My new Sony headphones" becomes "headphones", which stays plural because some nouns are always plural in English. If a team wants to count complaints per product, "mug" and "mugs" and "blue mug" must all be the same thing. Urgent is true only when the customer states a deadline or uses one of a short list of words: urgent, asap, today, tonight, right now or immediately. An angry message is not urgent by itself. Without that line, every angry customer would jump the queue.

These rules were the first draft. The category rules turned out to have a gap, and then my fix for it turned out to have one of its own. Later slides show how the data found both.

A two-column page headed real messages from each writer, titled the test writer wrote for other customers. Training writer: do i have to pay return postage on a bike; Dear sir or madam, please cancel my annual membership and confirm no further payments will be taken; Your site took payment for a phone charger on order 85561 in dollars and my bank added a conversion fee. I'd like a refund of the fee as it was your error. Test writer: The courier leave the package in the rain and the book is all wet now. Please can you send me new one? Order number is 71502; charged twice for order 55120 the blue wool scarf. pls refund the extra asap; Dear Sir or Madam, I write regarding order number 70244. I was billed for three ceramic plates, yet only two were supplied. I would be grateful if you could refund the difference at your earliest convenience. Yours faithfully. Beneath: training, short and tidy, mostly. Test: second-language English, phone typing, formal letters.

Reading them side by side shows it better than counts. "The courier leave the package in the rain and the book is all wet now" is the English of someone who learned it later in life. "charged twice for order 55120 the blue wool scarf. pls refund the extra asap" is a phone. Those are the kinds of messages a model trained on tidier writing might miss, and that is what the test is for.

One more thing you should know about the training pool. In 46 of the 640 training messages, one common word was misspelled by a small script, for example "charged" became "chraged". The writer marked them and I recorded which ones. So some spelling mistakes in the training data were added mechanically, not typed by a person in a hurry. I have no such record for the test messages, so I cannot tell you how their mistakes were made.

Message tr156 asks for a refund of the cost of a return label. A and B, working without the boundary rule, said billing; the third annotator, with it, said returns. That is not the third annotator's mistake: the money is tied to sending an item back, and the new rule says that is returns. So the rule seemed to reach a label that A and B had agreed on. It did not only settle the 37 splits; it could change messages nobody had flagged.

Message tr416 showed a second, worse problem. It is an order that never arrived, and the customer wants their money back. A and B said delivery. The third annotator said billing, and it was following my words: the rule said "the refund of a cancelled order or a missing item is billing". But the delivery definition says delivery covers "parcels that have not arrived". My rule contradicted my own definition. The third annotator had simply obeyed the newer sentence. The next slide is what I did about both problems.

The 14 fall into four kinds. One caution before reading them: D and E are different runs from A and B. A change can come from the tightened rule, or simply from two new annotators reading the same message differently, and with this data I cannot separate the two.

A two-column page headed real messages whose agreed category changed, titled changed by D and E, and not all by the rule. Message, then before and after. te125: Good morning. Owing to a change in my circumstances I must ask you to cancel order number 62015 for the garden bench before it is dispatched, and to ... : delivery, now billing (wants cancel). tr639: Hi, if order 40127 hasn't left yet please cancel it immediately and keep it from shipping: delivery, now billing (wants cancel). tr602: hi, my card got hit for £38 for the jeans on order 61930 but i returned them three weeks ago. can i have that money back please: billing, now returns (wants refund). tr102: My order number is 60317. The cordless drill was charged but your email says it was never dispatched. I'd like a refund rather than waiting: billing, now delivery (wants refund). te052: The delivery driver took a photo of my front door with my child in it. Please delete that photo: delivery, now account (wants other). Beneath: no rule covers a plain request to cancel. Billing there is how D and E read it; A and B said delivery.

Delivery to billing, 6 messages. All 6 ask to cancel an order that has not shipped yet, such as "if order 40127 hasn't left yet please cancel it immediately", and none asks for money back. A and B said delivery for all 6; D and E said billing for all 6. The rules do not cover this case. The boundary rule makes only "the refund of a cancelled order" billing, and nothing in the rules says where a plain request to cancel belongs. So billing here is how D and E read an unclear rule, not what the rules say. The clearest sign is tr229, "Please cancel the bookcase before it ships today", almost the same request as the test message te058 ("Please cancel my order before it ships today"): it is one of the 7 new splits, with D saying delivery and E billing. This is a hole the rules still have. I am leaving it in the data, and telling you, rather than relabelling a third time: the 6 messages keep D and E's billing, and tr229 has no category gold. A label is only as settled as the rule behind it.

Billing to returns, 4 messages. These fit the boundary rule: money tied to an item sent back, such as "my card got hit for £38 for the jeans ... but i returned them three weeks ago", and tr156, the return label. Billing to delivery, 3 messages. These fit the tightened sentence: an order "never dispatched", or a bowl that arrived cracked, where the customer wants a refund, now stays delivery. Delivery to account, 1 message: te052, a customer asking us to delete a photo the driver took of their front door. Neither rule touches this one. Privacy was in the account definition from the start, so this change is two sets of annotators reading the same message differently.

As for the two controls: tr156 is now returns, and tr416, the parcel that never arrived, is delivery again, which is what the definition said all along.

Changing a rule after seeing where annotators disagree is normal labelling work. It is how rules get good. The lesson here is about how far the change has to reach: when you change a rule, relabel everything the rule could touch, not only the messages that exposed the problem. Relabelling everything is also what showed me the cancellation hole: it appeared only because D and E read all 780.

When this happened matters too, so here it is plainly. No model had been trained on any of this data. Prompted models had already been run on the 140 test messages and on the chapter's 80 older ones, using the house rules of the time as their prompt; the reason for the change was the two controls above, not any model result. Their stored replies are being scored again against the new labels, and no result from them appears in this lesson.

The closest pairs are the same request reworded: a copy of personal data, a password reset. The third pair is interesting for this task. Both messages are about a returned lamp and a refund before Friday, but the order numbers differ, 70832 and 66390. For ticket extraction a model still has to copy the right digits from the message in front of it, so a near copy does not hand it the whole answer. It does hand it the category, what the customer wants, the item and the urgency, which is why the rule still applies.

The line is a line. The message about unsubscribing scored exactly 0.8000 and was dropped; a message about a crushed box and a broken lamp scored 0.7986 and was kept, with no one judging its meaning. The stored scores were checked: the lab found each training message's closest test message again, and every score matched the stored one to the fourth decimal place.

These sets are now frozen: fixed, and not to be changed because of anything a model does on them later. Every lesson from here on trains on the 500, watches its loss (a number for how wrong it is) on the 60, and is scored on the 140 and the 80.

This is a real run in VS Code's terminal.

A real screenshot of VS Code's terminal after running python agreement_demo.py. A table of field, agree, chance and kappa. Category: 16 / 20, chance 0.265, kappa 0.728. Wants: 18 / 20, chance 0.152, kappa 0.882. Order_number: 20 / 20, chance 0.230, kappa 1.000. Item: 19 / 20, chance 0.122, kappa 0.943. Urgent: 20 / 20, chance 0.580, kappa 1.000. Then where they split: tr193, te113, tr531 and tr020 on category; tr584 and te047 on wants; tr618 on item, yoga mat against mat.

Look at urgent: 20 of 20 agree, but chance agreement is already 0.580, because most messages are not urgent. Kappa is still 1.000, because they agreed on every row, including the urgent ones. Category agrees on only 16 of 20, because I put four splits in, and chance there is lower, 0.265, since four categories are used more evenly. Its kappa, 0.728, is the lowest here. On the full 780 messages category's kappa was 0.937; this sample was chosen to contain splits, so its numbers are lower than the whole set's and should not be quoted as the lesson's result.

I checked the script against the stored data in code, not by eye: the lab's demo mode reads the script's 20 rows, confirms every message text and every label is the same as in the stored annotation files, and computes agreement, chance and kappa for the same 20 ids from those files. All five fields matched the script's output exactly.

data/ticket4_changes.json
  • the training and validation files hold exactly the kept messages with their gold tickets.
  • Two modes do more than read. The pairs mode, which also uses NumPy, embeds all 640 training messages and the 220 test messages again with nomic-embed-text (the only model call in the lab, and only for ), finds which test message each training message is closest to, and checks the score against the stored one: all 640 matched, and 65 were at 0.80 or more, as stored. The typos mode recorded once which training messages had the scripted misspelling, from the first writer's working files. A demo mode checks the student script, and a box mode writes the playground on the next slide.

    What came after I saw the data: the choice to write new messages came after counting the empty fields in the first try. The boundary rule came after reading the 37 splits, and the third annotator after that. The tightened rule and the relabelling by D and E came after reading the third annotator's two misses on the controls. The 0.80 leak rule was carried over unchanged from lesson 1. No model was trained on any of this data before the sets were fixed.

    Three brand cards headed the tools, with their logos, titled what ran where. Python: the report, the demo, the box. NumPy: the similarity table of the leak check. Ollama: nomic-embed-text, the leak check only.

    FIELDS

    The leak line catches copies, not cousins. Training messages just under 0.80 can still be close rewordings of a test message, as the kept lamp message at 0.7986 shows. That closeness helps any model that learns from them.