Fine Tuning

A Fine-Tune on New Cases: Messages Shaped Unlike Its Training, and a Gap No Model Can Fix

0 of 25 complete

0%

Contents

Back|Fine TuningA Fine-Tune on New Cases: Messages Shaped Unlike Its Training, and a Gap No Model Can Fix
1/25
72 min left
Prerequisites
Serving a Fine-Tune: Moving It Into Ollama, Shrinking It, and Testing It Againrequired
Related Topics
Testing a Prompt on New Cases: Why the Score You Tuned On May Be OptimisticPrompting as EngineeringSmall Rewordings: The Same Instruction, Seven WaysPrompting as EngineeringInstructions Around a Long Document: Where to Put the RulesPrompting as EngineeringPrompts as Code: Templates That a Customer Cannot BreakPrompting as EngineeringPutting a Prompt Together: Which Parts Still Earn Their PlacePrompting as Engineering
1 of 25

A Referee in a New League

Think of a young football referee. She learned the job in one local league. The matches there were tidy: the same kind of teams, the same weather, the same small crowd, players who spoke her language. After a season she made almost every call correctly, and her coach was proud of her.

Then she is sent to referee a cup match far from home. The pitch is wet, the crowd is loud, two players shout at her in a language she does not know, and three things happen at once near the goal. Some of her calls are now harder. That is expected, and it tells her coach something useful: how good she is away from the matches she practised on.

Then something happens that the rule book simply does not cover. She checks the book, carefully, and finds nothing. A second referee, watching the same moment, makes the opposite call. Neither of them is careless. The book is missing a rule. No amount of practice fixes that. Only the league can fix it, by writing the rule down.

An illustration of a woman at a desk reading a thick open book through a magnifying glass while she writes in a notebook, with a laptop and a pot of pens beside her and books stacked on the desk. Headed a case the rule book does not cover, titled practised on one kind of message, tested on six new kinds. Beneath: on 53 new messages, whole tickets right: fine-tuned 0.5B 22, fine-tuned 1.5B 21, prompted 7B 24, prompted 14B 25. On cancel requests the two annotators agreed on the category 0 times in 15.

This lesson does the same to the two small models I trained in lesson 4. So far I have tested them on messages written in the same style as their training messages. Now I send them messages of six new kinds, and I compare them with two larger models that were only given the written rules. Both halves of the referee story happened: every model scored lower on the new messages, and one kind of message exposed a missing rule.

Nine Words for This Lesson

A hand-drawn list headed nine words for this lesson, titled testing a fine-tune on messages unlike its training. New cases: messages shaped unlike the ones a model was trained or tuned on. Kind: one shape of message: long emails, two requests, other languages, decoy numbers, rare products, cancels. Annotator: a person, or here an AI model, who writes the right ticket for a message. Gold field: a field both annotators filled the same way; only gold fields are scored. Fully gold: all five fields gold; only these messages can score a whole ticket. Leak line: cosine similarity 0.80 to a training message; at or over it, a new message is dropped. Rule gap: a case the house rules do not decide, so careful people split. Paired: two set-ups on the same messages, compared message by message. Unpaired: two scores on different messages; only the totals can be compared. Beneath: a whole ticket: all five fields gold, and all five right.

New cases are messages shaped unlike the ones a model was trained on, or unlike the ones you tested it on while you built it. People also call this a distribution shift: the kind of input changes, even though the task stays the same. A kind is one such shape. I chose six: long emails, two requests in one message, other languages (English mixed with Hindi, Bengali or Tagalog, or written wholly in Spanish or German), messages with decoy numbers (numbers that are not the order number), rare products with brands and sizes, and plain requests to cancel.

As in every lesson of this chapter, the task is a ticket: five fields (category, wants, order number, item and urgent) filled in by the house rules of lesson 2. An annotator writes the right ticket for a message. A field is gold when two annotators, working apart, filled it in the same way, and only gold fields are scored. A message is fully gold when all five fields are gold, and only those messages can score a whole ticket: all five fields right.

The leak line is the chapter's rule against testing on a copy of the training data: a new message whose closest training message has a (a score for how alike two texts are in meaning, usually between 0 and 1 here, where 1 means the same) of 0.80 or more is dropped. A rule gap is a case the house rules do not decide. Paired and unpaired comparisons are the two kinds from the prompting chapter's lesson 11: paired means the same messages, compared one by one; unpaired means different messages, where only totals can be compared.

The Question, and How the Messages Were Made

Lesson 4 ended with a result that surprised me. On the chapter's 140 test messages, the fine-tuned 0.5B model got 93 whole tickets right of the 131 fully gold ones, the same as the prompted 14B, a model almost thirty times its size, and the fine-tuned 1.5B got 102. The fine-tunes had only a one-line prompt. The 14B had the full rules and a JSON schema.

But those 140 test messages were written like the training messages: short, one request each, plain English. The training messages run from 7 to 32 words, with a median of 17. A fair worry is that the fine-tunes learned the shape of those messages, not the job. Real customers do not write to a template. Some send a long email with their life story; some ask two things at once; some write half the message in their own language; some paste a tracking code where you expected an order number. A model that only works on tidy messages will fail quietly on exactly the messages that need the most care, and the ordinary test set will not warn you, because it is tidy too. So the question of this lesson is: on messages shaped unlike its training data, can the fine-tune still not be told apart from a prompted 14B, or does it fall behind? And a second question came out of the data: what can no model fix?

A flowchart headed how the new messages were made and scored, titled written blind, labelled blind, checked for leaks. A writer who saw only the house rules leads to 90 messages, 15 of each kind. That leads three ways: to annotator H, blind; to annotator I, blind; and to nearest training message. Both annotators lead to gold: where H and I agree. Nearest training message leads to 7 at 0.80 or over: dropped. Gold and the dropped box both lead to 83 scored, four set-ups. Beneath: the design, the kinds and the 14 sign tests: recorded in CHAPTER-PLAN.md and newcases.py before any message was written. 53 of the 83 scored messages are fully gold.

I wrote the design down in the description at the top of the lab file newcases.py: the six kinds, the four set-ups, the scoring and the 14 sign tests I would run. The record that this came before any message was written is the dated entry in the chapter's plan file, CHAPTER-PLAN.md (2026-09-28, the new-cases lab "designed before any message was written"). Then a writer who saw only the house rules and the list of kinds wrote 90 messages, 15 of each kind. The writer never saw the training set or the test set. The messages were written by an AI model and labelled by two AI models (annotators H and I), as with all of this chapter's data; they are not real customers.

The Six Kinds

Each kind tests something the training data did not show. Here is one real message of each.

An editorial page headed the six kinds, one real message each, titled shaped unlike the training messages. Six labelled zones. Long, an email of 120 to 250 words: Hello there, I hope you are all well. I ordered a garden bench from you about three weeks ago as a surprise for my mum's 70th birthday. The ... (140 words). Two, two requests in one message: Where is my order 60419? It was meant to come Monday. If it's not coming this week, cancel it and refund me. Mixed, other languages: English mixed with Hindi, Bengali or Tagalog, or written wholly in Spanish or German. This one is all Spanish: Buenos días, el paquete con la manta llegó mojado y la manta tiene manchas. Quiero una nueva por favor. Pedido 83012. Decoy, other numbers beside the order number: The 1.7 litre kettle, item code KT-2210, ordered 19/09 for £24.00, hasn't shipped. Tracking number TRK88214573 doesn't work. Where is it? Rare, unusual products, with brands and sizes: The solar-powered copper-finish bird bath fountain with LED ring I bought was charged at £89 but listed at £69. Refund the difference. Cancel, a plain cancel, not yet shipped: Please cancel order 41822, I no longer need it. Thank you. Beneath: the 500 training messages: 7 to 32 words, median 17, one request, plain English.

Long messages are emails with a greeting, a story, sometimes a quoted earlier reply and a signature with a phone number. The request is buried in the middle. Two messages ask for two things; the rules say to choose the action asked for first. Mixed messages are in other languages: 9 mix English with Hindi or Bengali (written in Latin letters) or Tagalog, and 6 are written wholly in Spanish or German. The design asked for English mixed with another language; the writer wrote 6 with no English at all, and I kept them, because customers do that too. Decoy messages hold other numbers next to the order number, or instead of it: a phone number, a price, a date, a tracking code, an invoice number. Rare messages are about unusual products, with brands, sizes and adjectives that the item rule says to strip. Cancel messages ask to cancel an order that has not shipped, with no refund asked.

An isometric drawing of seven blocks side by side, headed median words per message, heights to scale, titled training: 17 words, the long emails: 142. A short block for training, 17; a tall block for long, 142; short blocks for two, 22; mixed, 21; decoy, 22; rare, 20; cancel, 16. Beneath: long emails ran 130 to 227 words; no training message was longer than 32. The other kinds are about as long as the training messages; they differ in shape.

The Four Set-Ups

The lab uses the same four set-ups as the chapter's test runs, with nothing changed, so the two sets of results can be put side by side.

A sequence diagram with four columns: the lab, fine-tunes (mlx), 7B and 14B (Ollama), the grader. Step 1, the lab sends the fine-tunes one line plus the message. Step 2, the fine-tunes send the lab a reply. Step 3, the lab sends the 7B and 14B the rules plus the schema plus the message. Step 4, they send the lab a reply. Step 5, the lab sends the grader the replies plus the gold. Step 6, the grader scores the gold fields. Headed how each new message was answered and scored, titled the same four set-ups as the chapter's test. Beneath: fine-tunes: the one-line prompt they were trained with, greedy, at most 80 tokens. Prompted: temperature 0, seed 1, the full rules and the JSON schema.

The two fine-tunes are lesson 4's adapters on Qwen2.5 0.5B and 1.5B, run in mlx, the tool that trained them. They get only the one line they were trained with: "Turn the customer message into a support ticket as one line of JSON." A LoRA adapter is the small add-on lesson 4 trained beside a frozen base model. They choose the most likely next token (a piece of a word, the unit a model reads and writes) every time, which is called greedy decoding, and may write up to 80 tokens. Why only one line? Because that is how they were trained, in lesson 4: the house rules were never shown to them as text. Everything they know about the rules, they learned from 500 labelled examples. That is exactly what this lesson tests: whether rules learned from examples carry over to messages unlike the examples.

The two prompted models are qwen2.5:7b and qwen2.5:14b in Ollama. They get the full house rules as the system message and the chapter's JSON schema (lesson 3: a description of the allowed shape of a reply, which Ollama enforces as the model writes). They run at temperature 0, which means always take the most likely next token, with seed 1, a fixed starting number for any random choice, so a rerun gives the same reply.

The grader is the chapter's own. It reads each reply as JSON and compares each field with the gold value, after putting text in lower case. Every reply from all four set-ups was valid JSON holding the five fields, so no score here comes from a reply the grader could not read.

Seven Messages Were Too Close to Training

Before scoring, the lab found each new message's closest training message, using the same model as the chapter's other leak checks (nomic-embed-text). An embedding turns a text into a list of numbers so that texts with similar meaning get similar numbers, and compares two such lists.

A two-column page headed the 7 new messages at 0.80 or over, and their nearest training message, titled too close to training: dropped before scoring. New message, then nearest training message. Decoy, 0.817: Paid £45 for 2 pillows, the site now says £32 for 2. Can I get the ...; The pillows on #26093 were two for £30 in the basket but I was charged £40. ... Decoy, 0.813: please call me on +44 7700 900610 about the rug I sent back on the ...; My return of the rug on order 43517 was accepted but you only refunded the ... Cancel, 0.805: Hello, I need to cancel order 58430 today please, it must not ship.; Hi, if order 40127 hasn't left yet please cancel it immediately and keep it ... Long, 0.804: Hi team, So the table lamp finally arrived yesterday after a week of ...; Order 30119 arrived with the box crushed and the lamp inside is broken. Can you ... Two, 0.803: I need a refund for the toaster from order 55012, it stopped working ...; Hello, I'd like to return the toaster from order 38850, it burns one side of ... Long, 0.802: Dear Customer Services, I am writing about a payment for order 71560, ...; I was charged £34.99 twice for the kettle on order #4471. Please refund the ... Cancel, 0.800: We no longer need the high chair we ordered, order 49105. Please ...; Cancel the payment for the high chair on order #72114 now and charge my other ... Beneath: dropped: decoy 2, cancel 2, long 2, two 1. The closest message kept: cancel_01, 0.791.

Seven of the 90 were at or over the leak line of 0.80: two decoy, two cancel, two long and one two-request message. Reading the pairs, most are the same situation in new words: a price difference on pillows, a refund for a returned rug, a lamp that arrived broken. The writer never saw the training data, so these are not copies; a shop's customers simply ask similar things. The rule drops them anyway, because a close match could let a fine-tune score by memory rather than skill. Dropping them costs a little test data; keeping them could make a fine-tune look better than it is on exactly the question this lesson asks. When in doubt, a test set should lean towards being too strict.

That leaves 83 scored messages. The closest of them to any training message was a plain cancel, "Please cancel order 41822, I no longer need it", at 0.791, just under the line. The mixed messages were the furthest from training, with a median best match of 0.574: an embedding of a sentence written wholly in Spanish is far from any English one.

The Annotators Agreed, Except on One Kind

Before looking at any model, look at the people, or here the AI annotators, who wrote the answers. Their agreement decides what can be scored at all. A model can only be marked right or wrong against an answer that someone settled, and a field the annotators never settled cannot be marked at all, however good or bad the model's reply.

Two panels headed annotators H and I, blind, all five fields agreed, titled every kind but one mostly agreed. The other five kinds: 57 of 75; category agreed on 69 of 75. Cancel: 0 of 15; category: H delivery 15, I billing 15. Beneath: on cancel, H and I agreed on wants, order number and urgent for all 15, and on item for 13.

On the five kinds other than cancel, H and I agreed on all five fields for 57 of 75 messages. Most of their splits were small: whether the item is "table lamp" or "lamp", "frying pan" or "pan". The two-request messages were the least settled of these, fully agreed on 8 of 15, because deciding which request comes first is itself a judgement.

On cancel, the picture is completely different. H said delivery for all 15 cancel messages. I said billing for all 15. They agreed on the category of none. They were not careless: on the other four fields of the same messages they agreed almost everywhere. They read an unclear rule two ways, and each kept to their own reading. That is the useful part. When annotators split here and there, a message is usually just hard to read. When they split the same way on every message of a kind, they are following two different rules, and the written rules allow both. Here it was 15 of 15 one way and 15 of 15 the other.

The same split appeared outside the cancel kind: 5 messages of other kinds also split between delivery and billing, and every one of them involves cancelling something, such as "please cancel order 40277 and also how do I change the password". So of the 83 scored messages, 53 are fully gold, and none of the 13 scored cancel messages is among them. Cancel has no whole ticket to score, though its other four fields can still be scored.

Four Set-Ups, the Same 53 Messages

Here is the headline, on the 53 fully gold new messages, which all four set-ups answered.

A bar chart headed whole tickets right, of the 53 fully gold new messages, titled 22, 21, 24 and 25: close together. Four bars on a scale up to 60, with a dashed line at all 53: FT 0.5B about 22, FT 1.5B about 21, prompted 7B about 24, prompted 14B about 25. Beneath: the same 53 messages for all four. Against the 14B: FT 1.5B +5 -9, p 0.4240; FT 0.5B +6 -9, p 0.6072.

The fine-tuned 0.5B got 22 whole tickets right, the fine-tuned 1.5B 21, the prompted 7B 24 and the prompted 14B 25. The four are close together, and all four got fewer than half right.

Because all four set-ups answered the same messages, this comparison is paired, and the sign test from earlier lessons applies. It looks only at the messages where two set-ups disagree, one right and one wrong, and asks how likely a split that uneven would be if only luck were at work. That chance is the p value. Against the 14B, the fine-tuned 1.5B was right alone on 5 messages and wrong alone on 9: p 0.4240. The 0.5B: 6 against 9, p 0.6072. Neither can be told apart from luck.

It helps to see what those numbers mean in messages. Of the 53, the fine-tuned 1.5B and the 14B were both right on 16 and both wrong on 23. They disagreed on only 14 messages, and the split of those 14 was 5 to 9. Suppose the two were equally good. Then a split of 14 messages would often be uneven just by chance: 5 to 9, or something more lopsided, would happen about four times in ten. That is what p 0.4240 means. So a gap of four tickets on 53 messages says very little.

A table headed the 14 sign tests, declared before the run, titled none below the line of 0.00357. Fourteen rows, one per kind and pair, each fine-tune to the 14B. Long: +1 -3 and +2 -3. Two: +1 -1 and +0 -1. Mixed: +0 -1 and +0 -2. Decoy: +0 -2 and +1 -2. Rare: +3 -2 and +3 -1. Cancel: +0 -0 twice, none fully gold. All: FT 1.5B +5 -9, p 0.4240; FT 0.5B +6 -9, p 0.6072. Every other p is 0.5000, 0.6250 or 1.0000. Beneath: + right only for the fine-tune; - right only for the 14B. Smallest p: 0.4240.

I declared 14 sign tests in the design (the record is the CHAPTER-PLAN.md entry above): each fine-tune against the 14B, per kind and over all messages. With 14 tests, the Bonferroni line (the usual 0.05 divided by the number of tests, so that running many tests does not make a lucky result easy to find) is 0.00357. The smallest p of the 14 is 0.4240, nowhere near it. Two of the 14 turned out empty once the labels came in, because cancel has no fully gold message; I keep them in the table because they were declared.

Lower Everywhere, on Different Messages

Every set-up got a smaller share right on the new messages than on the chapter's 140 test messages. This comparison needs care, because the two sets are different messages.

A bar chart headed share of whole tickets right, on two different sets of messages, titled every set-up scored lower on the new messages. Four pairs of bars on a scale of percent, the chapter's 140 then the new messages: FT 0.5B about 71 and 42; FT 1.5B about 78 and 40; prompted 7B about 61 and 45; prompted 14B about 71 and 47. Beneath: FT 0.5B 93 of 131 and 22 of 53; FT 1.5B 102 of 131 and 21 of 53; 7B 80 of 131 and 24 of 53; 14B 93 of 131 and 25 of 53. Different messages, writers and annotators: not a drop on the same messages.

On the chapter's 140 test messages, the fine-tuned 0.5B got 93 of 131 fully gold messages right, 71%. On the new messages it got 22 of 53, 42%. For the fine-tuned 1.5B it was 102 of 131 (78%) and 21 of 53 (40%); for the prompted 7B, 80 of 131 (61%) and 24 of 53 (45%); for the prompted 14B, 93 of 131 (71%) and 25 of 53 (47%).

It is tempting to say "the 0.5B dropped from 93 to 22". That sentence is wrong. No message is in both sets. The new messages were written by a different writer, labelled by different annotators, and chosen to be hard. The two numbers describe two different exams, not the same exam taken twice. What you can say is that every set-up scored lower on this new set than on the 140.

The comparison is unpaired, so the sign test does not apply. The prompting chapter's lesson 11 used Fisher's test for this: it asks how often two groups of this size would split this unevenly if both had the same chance of a right answer. I ran it after seeing the data, so read it as a description. For the two fine-tunes it gives p 0.0004 and below 0.0001; for the 14B, 0.0037; for the 7B, 0.0704. So three of the four scored lower by more than luck would usually give, and the 7B's lower score cannot be told apart from luck. The fine-tunes had further to fall, because they started higher on the 140. But whether they fell further than the 14B is a question about two different sets, and on the same 53 new messages, the paired test above could not separate them.

The two sets also differ in what they contain. The 131 include cancel requests whose category the annotators happened to agree on, so they were scored: te057 and te059 as delivery, te058 and te125 as billing. The 53 new ones include no plain order cancel at all, because the new annotators split on every one.

Kind by Kind

The totals hide how each kind went. The fully gold messages per kind are few, from 7 to 13, so read these counts as a picture, not a measurement.

A bar chart headed whole tickets right, by kind, titled no kind where the fine-tunes fell apart. Five groups of four bars, fine-tuned 0.5B, fine-tuned 1.5B, prompted 7B, prompted 14B, on a scale up to 14: long (12) about 8, 7, 8, 9; two (7) about 1, 2, 3, 2; mixed (11) about 2, 3, 3, 4; decoy (10) about 5, 4, 6, 6; rare (13) about 6, 5, 4, 4. Beneath: long 8/7/8/9; two 1/2/3/2; mixed 2/3/3/4; decoy 5/4/6/6; rare 6/5/4/4. Cancel: no fully gold message, so no whole ticket to score.

The long emails went best for every set-up: 8, 7, 8 and 9 of 12. That surprised me most. The fine-tunes had never seen a message longer than 32 words, and here they read emails of up to 227 words and found the request in the middle. The two-request messages went worst: 1, 2, 3 and 2 of 7. The mixed kind was hard for everyone too: 2, 3, 3 and 4 of 11. On decoys the prompted models did a little better (6 and 6 against 5 and 4), and on rare products the fine-tunes did (6 and 5 against 4 and 4). None of these per-kind differences can be told apart from luck.

Across the 53, all four set-ups got the same 10 messages right, and all four got 17 wrong. That second number matters: on a third of the messages, no set-up at all wrote the whole ticket. The 17 are not spread evenly. Seven are mixed messages, four are rare products, three are two-request messages, two are decoys and one is a long email. On mixed messages, 7 of the 11 fully gold ones defeated every set-up, which says the kind itself, or its gold, is hard, not that one kind of model is weak. The next slides read the replies to find out why.

A two-column page headed field by field, right where that field is gold, titled item and wants: the weakest fields on the new messages. Fine-tuned 1.5B, new, then 140; prompted 14B, new, then 140. Category: 56 of 64 (88%); 90%. 57 of 64 (89%); 88%. Wants: 68 of 81 (84%); 99%. 67 of 81 (83%); 89%. Order number: 78 of 83 (94%); 99%. 80 of 83 (96%); 100%. Item: 60 of 73 (82%); 91%. 57 of 73 (78%); 91%. Urgent: 74 of 82 (90%); 96%. 76 of 82 (93%); 97%. Beneath: different messages in the two sets: read the share, not a change. FT 0.5B wants 88%, item 77%.

Field by field, item and wants were the weakest fields on the new messages, for the fine-tunes and the prompted 14B alike. On the 140, the fine-tuned 1.5B had wants right 99% of the time; here, 84%. The 14B had item right 91% of the time on the 140 and 78% here. Again, these are two different sets, so these shares describe the sets; they are not a measured change on the same messages.

Decoy Numbers: A Code That Looks Like an Order

The rules say: "A phone number, price or date is never an order number." Decoy messages test that sentence, and also numbers it does not mention.

A page headed decoy numbers: what each set-up wrote as the order number, titled the fine-tunes took a tracking or item code as the order. Five rows, message then replies. Tracking code JD014472896GB has said 'out for delivery' ...: right: null. FT 0.5B JD014472896; FT 1.5B JD014472896GB; 7B null; 14B null. Your courier left a card saying parcel no. 5521 7730 0419 ...: right: null. FT 0.5B 5521; FT 1.5B 5521 7730 0419; 7B 552177300419; 14B null. The 1.7 litre kettle, item code KT-2210, ordered 19/09 ...: right: null. FT 0.5B KT-2210; FT 1.5B KT-2210; 7B 19/09; 14B null. Invoice INV-2026-0481 shows 3 lamps but I ordered 2, ...: right: null. All four INV-2026-0481. Deactivate my account please. Last login 2024, member ...: right: null. FT 0.5B 440192; FT 1.5B 440192; 7B null; 14B null. Beneath: a number written where the right answer is null, of the 83: FT 0.5B 5, FT 1.5B 5, 7B 3, 14B 1. The rules say a phone number, price or date is never an order number.

Where the right order number was null, the fine-tunes wrote a number anyway 5 times each, the 7B 3 times and the 14B once. Reading them, the fine-tunes took a tracking code, a parcel number, an item code, an invoice number and a customer number. The 7B once took the date "19/09". The 14B was caught only by the invoice number, which all four took.

Here is one possible reason, and it is a guess. In the training data, 280 of the 500 messages hold a digit, and 238 of those have an order number. So in the examples, a number in the message was usually the order number, and the fine-tunes may have learned "copy the number". The rule sentence names phone numbers, prices and dates, but not tracking codes, invoices or customer numbers. The 14B, reading the full rules, seems to have taken the idea behind the sentence; the fine-tunes never read that sentence at all, only examples of it. None of them ever took a phone number, a price or a date as the order, apart from the 7B's one date.

The decoy kind was the only one where the prompted models were ahead on the order number field: 12 of 13 for the 14B against 8 of 13 for each fine-tune.

What would I change? The rule itself is the first place to look. One more sentence, "a tracking code, parcel number, invoice number or customer number is never an order number", would cost nothing for the prompted models: they would read it on their next reply. For the fine-tunes, a sentence does nothing, because they never read the rules. They would need training examples that hold decoy numbers, labelled with a null order number, and a new training run. This pattern, a rule fix that is cheap for a prompt and costly for a fine-tune, comes back on the cancel slide.

Other Languages: The Word From the Other Language

The item rule asks for "the product the message is about, as a short common noun". For a message written in Spanish or German, the annotators wrote the English noun: "blanket" for "manta", "kettle" for "Wasserkocher".

A two-column page headed the mixed kind: the item, where the right answer is an English noun, titled the word from the other language, kept. Message and right item, then what the four set-ups wrote. Buenos días, el paquete con la manta llegó mojado y ...: blanket; manta, manta, manta, manta. Hallo, wann kommt meine Lieferung? Ich habe den ...: kettle; kocher, wasserkok, wasserkocher, wasserkocher. Ich möchte die Kissenbezüge umtauschen, die Farbe ...: cushion cover; kissen, kissenbezug, pillowcase, kissenbezug. Guten Tag, ich wurde für die Bettwäsche zweimal ...: bedding; bath, bedsheet, bettwäsche, bedsheet. Beneath: item wrong on mixed, of 13 gold: FT 0.5B 6, FT 1.5B 5, 7B 7, 14B 5. In order: FT 0.5B, FT 1.5B, 7B, 14B.

All four set-ups wrote "manta" for the Spanish blanket. For the German kettle the prompted models wrote "wasserkocher" and the fine-tunes wrote pieces of it, "kocher" and "wasserkok". For "Bettwäsche" the answers were "bath", "bedsheet", "bettwäsche" and "bedsheet". The item field on mixed messages was wrong 6, 5, 7 and 5 times of 13 gold.

This is a mistake shared by all four, not a fine-tune problem. And one could argue about the gold itself: the rules never say directly that the item must be in English. The annotators agreed on English, so English is gold here. This is a smaller version of the cancel gap: the model is asked to follow a rule that is not written down. The fix is again a sentence: "write the item as an English noun, whatever language the message is in". Until the rules say it, a model that keeps "manta" has written a short common noun, as the item rule asks, and has broken only a rule nobody wrote.

What the mixed messages did well: the order number. The fine-tunes read "Pedido 83012" and "Bestellung 27459" as order numbers every time on these messages, and the order number field was right for all four set-ups on all 15 mixed messages.

Two Requests: First, or Second?

The rules say: "If the message asks for two actions, choose the one it asks for FIRST." A two-request message tests whether the model follows that sentence.

Three panels headed two requests in one message: the rule says take the first, titled the prompted models more often took the later request. Fine-tunes: 1 and 1; first: 11 and 11; neither: 2 and 2. Prompted 7B: 4; first: 9; neither: 1. Prompted 14B: 5; first: 8; neither: 1. Beneath: big number: replies whose wants was the later request's action, of 14. e.g. "Where is my order 60419? It was meant to come Monday. If it's not coming this week, ...": 7B cancel, 14B refund; first request: information.

I read the 15 two-request messages, from their text only, and wrote down the first and the second action each one asks for, as wants values; the file is data/newcases/two_actions.json, with the rule I used. I did this after the results were known. Then the report counts, for the 14 scored ones, whether each reply's wants was the first action, the second (or a later one), or neither. The fine-tunes took the first action 11 times each and the later one once each. The prompted 7B took the first 9 times and a later one 4 times; the 14B the first 8 times and a later one 5 times. "Where is my order 60419? ... If it's not coming this week, cancel it and refund me": the first request is information, and the 7B wrote cancel, the 14B refund.

The one message all four sent to the later action was "Please send me the return label for the duvet. Once that's done, can you let me know how long the refund usually takes?": all four wrote information. The 14B's later choices were not always the more forceful action: it chose information over stop ("Stop texting me about reviews. I also want to know why order 71114 still hasn't shipped") and over the label. Its replies suggest a habit of taking the later request, but that is a pattern I read after the results, on 14 messages and one run, and no reason for it is supported by this data. The later request also pulled the 14B's order number: on two messages where the order number belongs to the second request, it wrote null. If your customers often ask two things at once, the rule "choose the one asked for first" is worth testing on its own: write twenty such messages, mark the first request before any model runs, and count.

Two replies from the fine-tuned 1.5B were a new kind of broken. On one two-request message it added a sixth key, other, and on another it wrote extra keys after the five, repeating urgent three times. Both were still valid JSON, so the grader read them, taking the last value of a repeated key. The prompted models, held to the schema, never did this here. It is the same weakness lesson 9 found: without a schema, a fine-tune can wander out of the shape.

Rare Products and Long Emails

The item rule says: no colour, size, brand, quantity or adjective. The rare products test it with names like "the 25cm round rattan banneton with linen liner".

A page headed rare products: the item the rules want, and what came back, titled brands and sizes: what came back. Four rows. The solar-powered copper-finish bird bath fountain with LED ... Right: bird bath. FT 0.5B bird bath fountain; FT 1.5B bird bath fountain; 7B bird bath fountain; 14B bird bath fountain. My sourdough proving basket (the 25cm round rattan banneton ... Right: basket. FT 0.5B sourdough proving basket; FT 1.5B basket; 7B banneton; 14B proving basket. My galvanised steel 240 litre rotating composter with twin ... Right: composter. FT 0.5B composter; FT 1.5B galvanised steel; 7B composter; 14B composter. ordered a Dyson V15 Detect Absolute cordless vac in ... Right: vacuum. FT 0.5B vac; FT 1.5B detect; 7B vacuum; 14B vac. Beneath: item wrong on rare, of 15: FT 0.5B 7, FT 1.5B 6, 7B 3, 14B 4. Kept the right noun inside a longer name, all 83: FT 0.5B 2, FT 1.5B 3, 7B 4, 14B 4.

The prompted 7B and 14B had the item right on 12 and 11 of the 15 rare messages; the fine-tunes on 8 and 9. The mistakes differ in kind. All four wrote "bird bath fountain" where the annotators wanted "bird bath"; that one is arguable, since a fountain is part of what the product is. The fine-tunes' own mistakes are stranger: the 1.5B wrote "galvanised steel" for the composter and "detect" (a word from the model name "V15 Detect") for the vacuum. These look like a model grabbing a word from the product's description instead of the product itself, not like a model that kept too much.

The prompted models were more likely to keep the right noun inside a longer name, 4 times each over all 83 messages, against 2 and 3 for the fine-tunes. On whole tickets, the rare kind was where the fine-tunes did best against the 14B: 6 and 5 against 4 of 13.

A page headed the long emails: where each set-up missed, titled up to 227 words, and mostly right. long_01 (140 w): 14B: item. long_02 (144 w): FT 0.5B: category. long_08 (130 w): FT 0.5B: wants, item; FT 1.5B: wants, item. long_09 (227 w): FT 1.5B: wants; 7B: wants, urgent; 14B: urgent. long_11 (140 w): FT 1.5B: urgent; 7B: item. long_12 (142 w): 7B: category. long_13 (139 w): 7B: wants; 14B: wants. long_14 (138 w): FT 0.5B: wants; FT 1.5B: wants. long_15 (138 w): FT 0.5B: urgent; FT 1.5B: urgent; 7B: category; 14B: item, urgent. Beneath: whole tickets of 12: FT 0.5B 8, FT 1.5B 7, 7B 8, 14B 9.

The long emails deserve a second look, because they are the kind I expected the fine-tunes to fail. They did not. Their misses were few and mostly on wants. On a password reset email the fine-tunes wrote "other" where the answer is "change". That also happened on a short password message in Bengali and English and on a two-request one. One possible reason is in the training data. I found the training messages about signing in by a keyword search (password, passcode, reset, log in, login, log-in, sign in, signin, sign-in, locked out, two-step, verification code): 12 messages, labelled change 4 times, other 5 times and information 3 times. The fine-tunes may have learned a split habit from split examples.

Cancel: A Gap in the Rules

Now the kind that no model could score on. Here is the whole story of one message.

A hand-drawn sketch headed sketched: a plain cancel request, titled two annotators, two answers, every time. At the top a box: "Please cancel order 41822, I no longer need it." Two arrows lead down to two boxes: annotator H: delivery, 15 of 15; annotator I: billing, 15 of 15. Below them a box: the four set-ups wrote billing: 11, 12, 11, 12 of 13. Beneath: neither is wrong: the rules name no category for a cancel with no refund. So no cancel message has a gold category, and none can score a whole ticket.

"Please cancel order 41822, I no longer need it." Which team gets this ticket? The rules give four categories. Delivery covers "shipping, couriers, tracking and parcels that have not arrived". Billing covers "charges, payments, invoices, prices and refunds of money", and the boundary rule adds that "the refund of a cancelled order" is billing. But this customer asks for no refund. The order has not shipped, so it is not a delivery problem either. Nothing in the rules decides it.

The annotation files hold labels, not reasons, so here is my guess at the two readings. One reading is delivery (the order is stopped before it ships, which is the warehouse's and the courier's business); the other is billing (stopping an order stops a payment). Annotator H chose delivery and annotator I billing, on all 15 cancel messages without exception. Lesson 2 found exactly this hole when two other annotators relabelled the training data, and I left it in on purpose. Here it is, measured: 0 of 15.

The four set-ups did not split. The fine-tuned 0.5B wrote billing on 11 of the 13 scored cancel messages, the 1.5B on 12, the 7B on 11, the 14B on 12. On the other four fields they did well: wants was right on all 13 for all four, and so was the order number.

An editorial page in three zones, headed where the models' billing came from, one possible reason, titled the training data may already have chosen. One training message: "Hi, if order 40127 hasn't left yet please cancel it immediately and keep it from shipping." Lesson 2: A and B said delivery; D and E, relabelling, said billing. It trained as billing. It is the nearest training message to 5 of the 15 cancel messages. Training messages that ask to cancel: billing 33, returns 16, delivery 10; every training row whose ticket says wants: cancel, by the category it trained with; the returns ones cancel a return. Nearest training message, for the 15 cancels: billing 9, returns 4, delivery 2; the prompted models had no training data, only the rules, and also wrote billing. Beneath: a fine-tune may repeat the reading its labels took; it cannot settle a rule the labels never settled.

What No Model Can Fix

A model can learn the rules, from examples or from reading them. It cannot write a rule that is missing. When two careful annotators split 15 to 15, a model that picks one side is right by one reading and wrong by the other. No amount of training data, model size or prompt wording changes that, because the disagreement is not in the model. It is in the rules.

Two panels headed when the rule is decided: what changes, titled one line for a prompt; relabel and retrain for a fine-tune. Prompted model: edit the rules; add one line to the guide it reads; the next reply can follow it. Fine-tune: relabel, retrain; relabel every training row the new rule touches, then train again and test again. Beneath: either way, the annotators must first agree on the new rule, or the test cannot score it.

The fix is a decision, not a model. The shop decides: a plain cancel goes to, say, delivery, because the warehouse team stops the parcel. Someone writes that down as one sentence in the rules. Then the two routes differ in cost. For the prompted model, the new sentence goes into the guide it reads, and the next reply can follow it. For the fine-tune, the rule lives in its training labels, so every training row the rule touches must be relabelled, and the model trained again and tested again. Here the prompt is cheaper to change, and that is one real cost of that this chapter had not measured before. It is easy to miss, because on the day you train a fine-tune its rules look settled. Rules change for ordinary reasons: a new team is formed, a product line is added, a lawyer asks for privacy requests to go somewhere else. Each change is one sentence for a prompt, and a relabel-and-retrain cycle for a fine-tune, with a fresh test at the end. If you expect your rules to change often, count that cost before you choose.

Either way, the test set needs the rule too. Until the annotators agree, the cancel messages cannot score a whole ticket, for any model.

A page headed the same gold field wrong in all four set-ups, titled 16 fields no set-up got right. Invoice INV-2026-0481 shows 3 lamps but I ordered 2, total ...: order number: right null; wrote inv-2026-0481. The solar-powered copper-finish bird bath fountain with LED ...: item: right bird bath; wrote bird bath fountain. Buenos días, el paquete con la manta llegó mojado y la manta ...: item: right blanket; wrote manta. Please send me the return label for the duvet. Once that's ...: wants: right other; wrote information. Is the 12-piece Villeroy & Boch Manufacture Rock stoneware ...: category: right billing; wrote delivery or returns or account. Ami je garden chair order korechilam (order 19984) seta bhanga ...: category: right delivery; wrote returns or billing. Beneath: by field: wants 4, category 5, item 6, order number 1. A message every set-up gets wrong is worth reading: the rule, or the label, may be the problem.

Try It Yourself

This script sends five of the new messages, one per kind apart from long, to a fine-tuned model and a prompted model in Ollama, each set up the way the lab ran it: the fine-tune gets the one line it was trained with and no schema; the prompted model gets the full rules and the schema. It prints each reply and which agreed fields it got wrong. It uses only Python's standard library. (A comment in the script says the five messages are one per kind; it means one per kind other than long.)

A real screenshot of VS Code with newcases_demo.py open, showing the docstring, the one-line prompt SHORT and the start of the house rules RULES. The rest of the file is in the code box on this slide. Beneath: copy it from the box on the slide.

Before you run this lab. It needs Ollama running on your own computer, with qwen2.5:14b (ollama pull qwen2.5:14b, about 9 GB) and the fine-tuned 1.5B you imported in lesson 9 as ft-q15-f16 (its 16-bit import: the model merged with its adapter and loaded into Ollama at 16 bits per weight). If you have not set Ollama up yet, the lab setup guide shows how to install it and check that everything works, on macOS, Windows or Linux. A smaller machine can use qwen2.5:7b and ft-q05-f16 instead: python newcases_demo.py ft-q05-f16 qwen2.5:7b. I do not give a running time, because another lab was using the same machine while mine ran.

If you use Windows or Linux: the fine-tunes were trained and scored in mlx, which is built mainly for Apple silicon, and everything here ran on a Mac. Lesson 9 shows how to move a fine-tune into Ollama, which runs everywhere; to train one without mlx, Hugging Face's library on a free Colab GPU is the usual route. I have not run that path for this chapter, and nothing here claims the Mac's numbers hold elsewhere.

"""Five new kinds of message, sent to a fine-tuned model and to a prompted one, each set up the way the lab ran it.

This is the lab of lesson 10 of the fine-tuning chapter, made small. It needs Ollama running on your computer,
the fine-tuned 1.5B you imported in lesson 9 (ft-q15-f16) and qwen2.5:14b (ollama pull qwen2.5:14b).
It uses only Python's standard library:
    python newcases_demo.py                          # ft-q15-f16 and qwen2.5:14b
    python newcases_demo.py ft-q05-f16 qwen2.5:7b    # any models; a name starting "ft-" gets the fine-tune's set-up
"""
import json
import sys
import urllib.request

FIELDS = ["category", "wants", "order_number", "item", "urgent"]

# The fine-tunes: the one line they were trained with, no schema.
SHORT = "Turn the customer message into a support ticket as one line of JSON."

# The prompted models: the house rules, word for word as in TICKET-GUIDE.md, plus the JSON schema.
RULES = """Turn one customer message into one ticket, a JSON object with exactly these five fields.

## category (one of four, same definitions as the prompting chapter)
- billing: charges, payments, invoices, prices and refunds of money.
- delivery: shipping, couriers, tracking and parcels that have not arrived or arrived damaged.
- returns: sending an item back, exchanges and the return process.
- account: signing in, passwords, profile details, privacy and emails from us.

Boundary rule: a refund for an item the customer
has already sent back, or is sending back, is **returns**, because the refund is the last step of the return.
Money that is not tied to sending an item back (a wrong or double charge, an invoice, a price, the refund of a
cancelled order) is **billing**. A parcel that never arrived or arrived damaged stays **delivery**, even when the
customer asks for their money back.

## wants (one of seven): what the customer is asking us to DO
- refund: give money back.
- replacement: send the same item again, or swap it for another (an exchange).
- cancel: stop an order, subscription or payment before it happens or continues.
- change: edit details we hold (address, email, name, phone, password, payment card, delivery slot).
- stop: stop messages from us (emails, texts, calls, newsletters).
- information: they ask a question or for a status, and ask for no other action.
- other: none of the above (a complaint with no request, praise, deleting data, something else).
If the message asks for two actions, choose the one it asks for FIRST.

## order_number (string or null)
The order number exactly as digits, without "#", "order" or spaces ("order #4471" -> "4471").
null when no order number is written. A phone number, price or date is never an order number.

## item (string or null)
The product the message is about, as a short common noun: lower case, singular, no colour, size, brand,
quantity or adjective ("two blue ceramic mugs" -> "mug", "my new Sony headphones" -> "headphones",
"a pair of trainers" -> "trainers"). Keep nouns that are always plural in English (headphones, trainers,
jeans, scissors). null when no product is named ("my order", "the parcel", "it" are not products).

## urgent (true or false)
true only when the customer states a time limit or deadline, or says urgent, asap, today, tonight,
right now, or immediately. Anger alone is not urgency.

## Output
One line of JSON, keys in this order:
{"category": "...", "wants": "...", "order_number": "..." or null, "item": "..." or null, "urgent": true/false}"""

SCHEMA = {"type": "object", "properties": {
    "category": {"type": "string", "enum": ["billing", "delivery", "returns", "account"]},
    "wants": {"type": "string", "enum": ["refund", "replacement", "cancel", "change", "stop", "information", "other"]},
    "order_number": {"type": ["string", "null"]}, "item": {"type": ["string", "null"]}, "urgent": {"type": "boolean"}},
    "required": FIELDS}

# Five of the 90 new messages, one per kind, with the ticket both annotators agreed on ("?" = they did not agree).
MESSAGES = {
    "nc_two_15": ("Where is my order 60419? It was meant to come Monday. If it's not coming this week, cancel it and refund me.",
                  ["delivery", "information", "60419", None, True]),
    "nc_mixed_07": ("Buenos días, el paquete con la manta llegó mojado y la manta tiene manchas. Quiero una nueva por favor. Pedido 83012.",
                    ["delivery", "replacement", "83012", "blanket", False]),
    "nc_decoy_02": ("Tracking code JD014472896GB has said 'out for delivery' since Tuesday. Postcode is BS7 8QA. When will my parcel actually get here?",
                    ["delivery", "information", None, None, False]),
    "nc_rare_08": ("The solar-powered copper-finish bird bath fountain with LED ring I bought was charged at £89 but listed at £69. Refund the difference.",
                   ["billing", "refund", None, "bird bath", False]),
    "nc_cancel_01": ("Please cancel order 41822, I no longer need it. Thank you.",
                     ["?", "cancel", "41822", None, False]),
}


def ask(model, message):
    fine_tune = model.startswith("ft-")
    body = {"model": model, "stream": False,
            "messages": [{"role": "system", "content": SHORT if fine_tune else RULES},
                         {"role": "user", "content": message}],
            "options": {"temperature": 0, "seed": 1, "num_predict": 120} if fine_tune else
                       {"temperature": 0, "seed": 1, "num_ctx": 8192, "num_predict": 120}}
    if not fine_tune:
        body["format"] = SCHEMA          # the prompted models ran with the schema, as in the lab
    req = urllib.request.Request("http://localhost:11434/api/chat", data=json.dumps(body).encode(),
                                 headers={"Content-Type": "application/json"})
    return json.loads(urllib.request.urlopen(req, timeout=900).read())["message"]["content"]


def wrong_fields(reply, gold):
    try:
        got = json.loads(reply)
    except ValueError:
        return "not JSON"
    bad = [f for f, g in zip(FIELDS, gold) if g != "?" and str(got.get(f)).lower() != str(g).lower()]
    return "all agreed fields right" if not bad else "wrong: " + ", ".join(bad)


for mid, (message, gold) in MESSAGES.items():
    print(f"{mid}: {message if len(message) <= 70 else message[:67] + '...'}")
    for model in sys.argv[1:] or ["ft-q15-f16", "qwen2.5:14b"]:
        reply = ask(model, message)
        print(f"  {model}: {reply}")
        print(f"  {' ' * len(model)}  {wrong_fields(reply, gold)}")

The Lab Report

A real terminal recording headed python newcases_report.py, titled every table in this lesson, from the stored files. Seven numbered sections: the new messages, with words, nearest training message, dropped, scored and fully gold per kind, the closest scored message, nc_cancel_01 at 0.791, and the 7 dropped messages; the annotators' agreement per kind and field, with cancel category 0; whole tickets per kind for the four set-ups, 22, 21, 24 and 25 of 53, beside the 140 with Fisher's test; per field, new and 140, and misses by kind; the 14 sign tests, none beyond 0.00357; cancel, with H delivery 15, I billing 15, what each set-up wrote and the training rows asking to cancel; and kinds of mistake: nouns kept inside longer names, words copied from the message, decoy numbers taken as orders and extra keys, the two-request counts of first, second and neither, and the training data behind two of the patterns. Beneath: the lab's own report. It calls no model.

The report lives in scripts/labs/finetune/newcases_report.py. It reads only stored files: the 90 messages, the two annotations and the key that hides which message is which, the gold, the similarity file, the four runs, the declared sign tests, the training data and the chapter's runs on the 140. It calls no model. A json mode writes the numbers to results/nc-report.json, which the figures read.

Before it prints anything, it checks the files against each other. The gold must be exactly what the two annotations give; every stored grade is graded again against the gold; each run's rows must be the scored messages in order; and the 14 sign tests are computed again and must match the ones the lab stored. If anything differs, the report stops.

The demo mode checks the student script's stored run: its prompt, rules and schema must be the lab's, its messages and gold must match the files, and the prompted model's replies must equal the stored ones. The box mode writes the playground below and checks that it reproduces every field of every reply's score.

What came after I saw the data: the per-field tables, the Fisher tests between the two sets, the counts of kinds of mistake, my reading of which action each two-request message asks for second, the cancel counts from the training data, the demo's five messages and every example in the figures. The kinds, the four set-ups, the leak line, the scoring and the 14 sign tests were fixed before the run, in .

Pick a Kind and Read the Replies

This box has no model in it. It holds the 83 scored new messages, their kind, the gold ticket (a ? where the annotators did not agree) and the four set-ups' real replies. A reply that is JSON with the five keys in order is stored as a ticket; any other reply is kept as its text, and the box reads it as JSON when it scores it.

As it is, the box prints the whole tickets of all four set-ups (22, 21, 24 and 25 of 53), then the decoy messages field by field, then one decoy message with every reply: the kettle with an item code, a date and a tracking number, where both fine-tunes wrote "kt-2210" as the order number, the 7B "19/09", and only the 14B null.

Try score('long') or score('two') for one kind. Compare fields('mixed') with fields('decoy'): on mixed messages the item field is where every set-up loses, and on decoys it is the order number, for the fine-tunes above all. Try fields('cancel') to see category scored on nothing, because no cancel message has a gold category, while the other four fields still score. Try show('nc_cancel_01') to see four set-ups agree on billing where the gold says ?, and show('nc_mixed_07') for the blanket that all four called "manta".

The Code, Part by Part

The prompts. SHORT is the one line the fine-tunes were trained with, word for word. RULES is the house rules, word for word as in TICKET-GUIDE.md, which the prompted models read as their system message. SCHEMA is the chapter's JSON schema.

The messages. MESSAGES holds five of the 90 new messages, keyed by their ids, each with the ticket both annotators agreed on. A ? marks a field they did not agree on; the script does not score it.

The request. ask decides the set-up from the model's name. A name that starts with ft- is treated as a fine-tune: the one-line prompt, no schema, temperature 0, seed 1, at most 120 tokens, as in lesson 9. Any other name gets the rules, the schema and the lab's context window of 8,192 tokens (the most text the model reads at once). The request goes to Ollama's local chat address, and the reply's text comes back.

The check. wrong_fields reads the reply as JSON and lists the agreed fields that differ from the gold. It is a small version of the lab's grader; it compares text in lower case, as the grader does.

The loop. For each message it prints the start of the message, then each model's reply and what it got wrong. Read the replies side by side: every finding in this lesson started that way.

How to Test Your Own Fine-Tune on New Cases

A hand-sketched column of six boxes joined by arrows, headed what held in this one lab, titled testing a fine-tune on new cases, in six steps. 1, list the kinds your real messages come in. 2, have someone else write each kind, from the rules only. 3, two people label blind; keep only what they agree on. 4, drop anything too close to the training data. 5, compare set-ups on the same messages, by kind. 6, where labellers split, fix the rules, then relabel. Beneath: four set-ups, one run each, 53 fully gold messages: a place to start, not a rule.

List the kinds. Look at real messages, or ask the people who read them, and list the shapes your training data does not have: long, mixed, several requests, strange products, numbers that are not what they seem. Write the list down before you score anything.

Have someone else write each kind. The writer should see the rules, not the training data or the test set. Aim for the same number of each kind, so no kind hides inside a total.

Label blind, twice. Two people label every message without seeing each other's work or the kind. Keep only the fields they agree on. Count where they split, kind by kind.

Drop leaks. Compare each new message with the training data and drop anything too close. Report how many you dropped and why.

Compare set-ups on the same messages. Put your fine-tune next to the best prompted model on the same messages, per kind, with a sign test declared in advance. Never compare a score on these messages with a score on other messages as if it were the same test.

Fix the rules where the labellers split. A split is a finding about your rules, not noise. Write the missing rule, relabel what it touches (the training data too), and only then test again.

When a New-Cases Test Is Worth It, and When It Is Not

Before you ship any fine-tune. A fine-tune learned from examples of one shape. Its test set, if it came from the same writer or source, has the same shape. A new-cases test is the only way to see how it does on the inputs that the examples never covered. Here, it held; on another task it may not.

When the fine-tune is replacing a prompt. The comparison that matters is on the same messages. On the 140 test messages, the fine-tuned 0.5B could not be told apart from the 14B; on the new kinds, it still could not. The second result tells you the first was not only about tidy messages.

When your rules are still changing. A prompted model follows a rule change with one edited sentence; a fine-tune needs relabelling and retraining. If your labellers still split often, fix the rules before you fine-tune, or you will train the split into the model, as the cancel labels were trained into these fine-tunes.

Not as a replacement for the ordinary test set. New cases tell you about shapes you chose to test. The ordinary test set tells you about the messages you usually get. You need both.

Not with a handful of messages per kind if you want a verdict. With 7 to 13 fully gold messages per kind, this lab can show patterns, not prove differences. For a decision that matters, write more of the kinds that matter most to you.

Not as a reason to distrust every fine-tune. The result here is reassuring: two small fine-tunes, trained on 500 short messages, could not be told apart from a prompted model almost thirty times larger, on messages unlike any they were trained on. A new-cases test is how you earn that confidence, not a ritual for finding fault.

What These Numbers Can and Cannot Tell You

A two-column page headed read before you trust these numbers, titled what these runs are, and what they are not. They are: one run of each set-up; 53 fully gold messages, 7 to 13 per kind; two fine-tunes, two prompted models; kinds and tests declared before the run; examples and extra tables chosen after; messages and labels written with AI help. They are not: not a spread across runs; too few to separate close scores; not a rule for other sizes or tasks; not chosen after seeing results; read them as illustrations, not tests; not real customers.

One run of each set-up. Every set-up answered every message once. Lesson 5 showed that retraining with a different seed flipped 15 to 21 tickets on the 140, so small differences here mean little.

53 fully gold messages. That is enough to see that the four set-ups are close, not enough to rank them. Per kind there are 7 to 13, which is why no per-kind difference can be told apart from luck.

Two fine-tunes and two prompted models, on one task. Qwen2.5 0.5B and 1.5B, trained once each in lesson 4, against Qwen2.5 7B and 14B. I make no claim about other sizes, other model families or other tasks.

The design was declared; the reading was not. The kinds, the set-ups, the leak line and the 14 sign tests were written down before any message existed; the record is the dated CHAPTER-PLAN.md entry. The per-field tables, the Fisher tests, the kinds of mistake, my reading of the two-request messages and every example were chosen after I saw the replies. Read those as illustrations.

The gold can be argued with. The annotators agreed on English nouns for foreign words and on billing for a question before buying, where the rules say nothing. Agreement makes a field scoreable; it does not make it the only reasonable answer.

The data was written and labelled with an AI model's help, as in every lesson of this chapter. Real customers write differently, and real annotators may split in other places.

What to Do Next

A hand-drawn list headed before you trust a fine-tune, titled five questions. Kinds?: what shapes do real messages come in that training did not show? Writer?: was the new set written by someone who never saw the training data? Agree?: where did two labellers split, and is that a rule the guide lacks? Paired?: are the set-ups compared on the same messages, message by message? Rules?: when a rule changes, can you relabel and retrain, or only re-prompt? Beneath: here, on 53 new messages, no fine-tune could be told apart from the prompted 14B.

If you have a fine-tune, or are thinking of training one, you can start this week. Write down the kinds of message your real users send that your training data does not show. Ask a colleague to write ten of each from your rules alone. Have two people label them without talking. Then score your fine-tune and your best prompt on the same messages, kind by kind, and read every reply where they differ. Look hardest at the places where your two labellers split: that is where your rules need a sentence, and no model can write it for you.

A closing card headed to keep, titled a model can learn the rules; it cannot write the missing ones. In large type: 22, 21, 24, 25 of 53. Beneath: whole tickets on new kinds of message: fine-tuned 0.5B and 1.5B, prompted 7B and 14B. Neither fine-tune told apart from the 14B. Cancel requests: the two annotators agreed on the category 0 times in 15. That is a gap in the rules. Then: test by kind before you trust a fine-tune, and fix the rules where the labellers split.

The chapter's wrap brings it together: when was the right decision in these labs, what it cost, and what it could not do.

Knowledge Check

Knowledge Check

4 questions - Score 80% to pass

Q1

On the 140 test messages the fine-tuned 0.5B got 93 of 131; on the new messages it got 22 of 53. What is the honest way to read those two numbers?

Q2

Annotator H called all 15 cancel requests delivery and annotator I called all 15 billing. What does that show?

Q3

On decoy messages, the fine-tunes wrote a tracking code or an item code as the order number. What is the lesson's possible reason?

Q4

The shop decides that a plain cancel goes to delivery. What must change for the fine-tune to follow the new rule?

Two fresh annotators, H and I, then labelled all 90 without seeing each other's work. The messages were shuffled and given meaningless ids, so the annotators could not see which kind each message was. A field is gold only where H and I agree, which is the chapter's rule since lesson 2.

Only one kind is longer than the training messages. The long emails ran from 130 to 227 words, and no training message was longer than 32. The other five kinds are about as long as the training messages, from 16 to 22 words at the median. They differ in shape, not in length: what is in them, and how it is said.

The cancel kind was in the design for a reason. Lesson 2 found that the rules have no category for a plain request to cancel, and I expected the annotators to split on it. I put it in on purpose, to see what a rule gap looks like from the outside.

So the answer to the chapter's question is: here, no fine-tune could be told apart from the prompted 14B, on messages shaped unlike their training. The direction did move. On the 140, the fine-tuned 1.5B was numerically ahead of the 14B (right alone on 21 messages, the 14B on 12, p 0.1628); here it was behind (5 against 9, p 0.4240). Neither split can be told apart from luck. With 53 messages, a real difference of a few tickets could hide inside this result, so "cannot be told apart" is the honest wording, not "equal". What 53 messages can do is catch a large collapse: if the fine-tunes had lost most of their skill on these shapes, that would likely have shown up.

Why billing? For the fine-tunes there is one possible reason. One training message, "Hi, if order 40127 hasn't left yet please cancel it immediately and keep it from shipping", was labelled delivery by the first two annotators in lesson 2 and billing by the two who relabelled it. It trained as billing. It is the closest training message to 5 of the 15 cancel messages. Of the training rows whose ticket asks to cancel, 33 were billing, 16 returns (cancelling a return) and 10 delivery. The fine-tunes may have learned the reading that their labels happened to take. This needs saying plainly. It is easy to read the models' agreement as evidence. All four set-ups writing billing on 11 or 12 of 13 messages looks like a strong answer. It is not an answer at all: it is one of two equally possible readings, baked into the training labels in one case and suggested by a nearby phrase in the rules in the other. Agreement between models tells you they share a habit, not that the habit is right.

The prompted models had no training data, and they also wrote billing. One possible reason is the boundary rule's phrase "the refund of a cancelled order is billing", which puts the word cancelled next to billing. That is a guess. The point does not depend on it: the models agreed with each other, but that does not make them right. On this question there is no right answer yet, because the rules do not give one.

Cancel is the clearest gap, but not the only place to look. In 16 cases, a gold field was wrong in all four set-ups. Some are plain model mistakes, like the invoice number taken as the order number. Others point back at the rules. "Is the 12-piece Villeroy & Boch stoneware dinner set dishwasher safe?" is a question before buying. The four categories do not cover it; both annotators chose billing, and the four set-ups spread over delivery, returns and account. When every model misses the same field, read the message and the rule before you blame the models.

This is a real run in VS Code's terminal (python newcases_demo.py).

A real screenshot of VS Code's terminal after running python newcases_demo.py. For each of five messages, nc_two_15, nc_mixed_07, nc_decoy_02, nc_rare_08 and nc_cancel_01, it prints the start of the message, then the reply from ft-q15-f16 and from qwen2.5:14b, each followed by the fields it got wrong. ft-q15-f16: all agreed fields right on nc_two_15 and nc_cancel_01; item mattress on nc_mixed_07; the tracking code as the order number on nc_decoy_02; bird bath fountain on nc_rare_08. qwen2.5:14b: wants refund on nc_two_15; item manta on nc_mixed_07; all agreed fields right on nc_decoy_02 and nc_cancel_01; bird bath fountain on nc_rare_08. Both wrote billing for the cancel.

When I ran it, all 5 of the 14B's replies were character for character the same as the lab's stored replies: same engine, same settings. The fine-tune is different. The lab ran it in mlx; the script runs lesson 9's 16-bit import in Ollama. 4 of its 5 replies matched the mlx ones exactly. On the Spanish blanket, mlx wrote "manta" and Ollama wrote "mattress": wrong either way, but a different wrong answer. Lesson 9 showed why: after the move, the 1.5B matched mlx on 136 of 140 replies, and borderline answers can tip. The report's demo mode checks all of this against the stored files. I chose these five messages after reading the lab's replies, because each shows one kind's typical mistake.

newcases.py

Four brand cards headed the tools, with their logos, titled what ran where. Apple mlx: ran the two fine-tunes. Ollama: ran the prompted 7B and 14B, and the leak check's embedder. Hugging Face: the two base models. Python: the lab, the report, the demo, the box.