Imagine a shop that takes on four new people for its customer service desk on the same Monday. On the first morning, the manager gives all four the same briefing. She reads out the house rules for writing up a customer's problem: which team it goes to, what the customer is asking for, the order number copied exactly, the product named in one plain word, and whether it is truly urgent. Everyone hears every rule. Everyone gets a printed copy to keep on the desk.
By Friday, the manager looks through the week's write-ups. They are all neat and in the right format, because the form on the screen only accepts the right format. But the content is another matter. One person writes "the blue ceramic mug" where the rule says "mug". Another marks every angry customer as urgent. The most experienced of the four gets almost everything right, and still, when a customer writes in only to say thank you, writes it up as a question.

The useful question for the manager is not "who is best?" It is "which rules are being dropped, and by whom?" A rule that the newest person drops but the experienced person keeps will fix itself with time. A rule that all four drop in the same way is a different problem: the briefing is not enough for it, and something else has to change.
This lesson does exactly that with four language models instead of four people. Each one gets the full house rules from lesson 2 as its instructions, and I check what it writes, field by field, on the 140 test messages. I wanted to know where a prompt alone stops being enough, because that is where training a model starts to be worth its cost.

Prompted model. A model used exactly as it was downloaded. Nothing about it is changed; everything it knows about the task comes from the text I send with each message. In this chapter, the opposite is a fine-tuned model, which is trained on examples.
House rules. The rules for the five ticket fields from lesson 2, written once in one file. Here the whole file is the model's instructions, sent again with every single message.
Schema. A short description of the shape the answer must have: which five keys, which words are allowed for category and wants, and which fields may be empty. Ollama, the program that runs the models on my computer, uses the schema to stop the model from writing anything that does not fit. The prompting chapter introduced this in its lesson on asking for a format.
Field score and whole ticket. A field score counts, for one field, how many messages the model got right. A whole ticket is right only when all five fields are right on the same message. A model can score well on every field and still get few whole tickets right, because its mistakes fall on different messages.
Gold. As in lesson 2, a field is gold only when the two annotators gave it the same value. Only gold fields are scored.
Sign test and Bonferroni. The two tools from the prompting chapter for asking whether a difference between two models is more than luck. The luck slide explains both again.
The floor is the weakest set-up in the lesson: a very small model given the same rules, to show what the rules can do when the model is too small to follow them.
The set-up is simple, and it is the same for all four models. For each of the 140 test messages, the lab sends Ollama one request with three things in it: the house rules as the system message (the instructions part of a chat), the customer's message as the user message, and the JSON schema. The model writes one ticket. A small grading step then compares each of the five fields with the gold ticket and marks it right or wrong.

The four models are qwen2.5:3b, llama3.2:3b, qwen2.5:7b and qwen2.5:14b. The number is the size of the model in billions of weights, the numbers a model learns during its original training; a bigger model usually follows instructions better but is slower and needs more memory. Two 3B models from different makers let me compare makers at the same size, and the three qwen models let me compare sizes from the same maker.
Every call used temperature 0 and seed 1. Temperature controls how much randomness goes into picking each next word; at 0 the model always takes its most likely choice, so the same message gives the same reply. I ran each model once on each message. That is one run, not an average over many, and I come back to what that means on the limits slide.
The rules each model read are the final house rules from lesson 2, word for word, including the tightened boundary rule, and the gold they are scored against follows the same rules. That was not true of my first runs, which read an older wording of one clause. I found that, reran all four models, and the history slide later in this lesson shows how much the scores moved. Everything else in this lesson uses the reruns.
A prompt with rules in it is not free. The rules are sent again with every message, and the model has to read all of them before it writes a word.

The house rules are 401 words long, counting every word that has a letter or a digit in it. A model does not read words, though; it reads tokens, small pieces of text, often part of a word. The stored results do not record how many tokens each prompt used, so I made one new call per model, on the first test message, and read Ollama's own count, prompt_eval_count. The whole prompt came to 648 tokens on all three qwen models and 653 on llama. That includes the rules, the 25-word message and a few tokens of chat formatting. Most of it is the rules.
Those same four calls did a second job. Each reply was compared with the stored reply for the same message from the lab's run, and all four were the same. So the stored results can be repeated, at least on that message, with this Ollama version.
I do not report how long any call took. Other labs were using Ollama on the same computer while these ran, so any time I measured would be partly someone else's work. Remember the token count instead: every ticket carries about 600 tokens of rules, and a fine-tuned model that has learned the rules would not need them.
Here is the first result, and it is the good news. Every reply from every model was valid JSON of the right shape: 140 of 140 for all four. That is the schema doing its job. No reply was a sentence to the customer, no key was missing, and no category or wants value was a word outside the list.

The fields tell different stories. The order number was right on 140 of 140 messages for qwen2.5:3b and qwen2.5:14b, and 139 for the other two. Urgent was right on 129 to 136 of 140. Category, scored on the 136 messages where it is gold, ranged from 112 on qwen2.5:3b to 121 on qwen2.5:7b. Wants, scored on 138, went from 105 to 123. Item, scored on 137, varied the most: 86 on qwen2.5:3b and 124 on qwen2.5:14b.
Now the whole ticket. A whole ticket counts only on the 131 messages where all five fields are gold, and it is right only if all five fields match.

qwen2.5:3b got 50 whole tickets right, llama3.2:3b 54, qwen2.5:7b 80 and qwen2.5:14b 93. The 0.5B floor got 3; it has its own slide later. Even the largest model, which gets each field right between 88% and 100% of the time, gets the whole ticket right on 93 of 131, about 71%. Small misses on different fields add up.
A field score of 90% sounds good. But a team that has to fix one ticket in four by hand does not have a working system. So the rest of this lesson reads the wrong answers themselves, field by field, to find which rules they break.
Wants is the field where the rule asks the model for something models find hard: to write down only what the customer asked for, not what they will probably want. I read every message where any model got wants wrong, 44 messages in all, and tagged each one with the rule it tests. These tags are my reading, done after I saw the replies; the next fields are sorted by code instead.

A comment that asks for nothing is other. 13 messages are praise, a complaint or a report with no request in them. The rules say that is "other". All four models got nearly all of them wrong: qwen2.5:3b missed 13, llama3.2:3b 12, qwen2.5:7b 13 and qwen2.5:14b 11. Most often they answered "information", as if a thank-you note were a question.
A question is information. 12 messages only ask a question, such as "why was i charged 45 ... when the price said 39". The rules say that is "information", even though the customer probably wants money back. Here size mattered a lot: 9 wrong on qwen2.5:3b, 6 on llama3.2:3b, 3 on qwen2.5:7b and 1 on qwen2.5:14b. qwen2.5:3b answered "refund" on 7 of its 9; llama3.2:3b answered "refund" on 3 of its 6 and "change" or "replacement" on the other 3.

A request that is not on the list is other. Four messages ask for something none of the six actions covers: deleting data, a copy of personal data, unlocking an account. qwen2.5:14b got 2 of the 4 wrong, llama3.2:3b 3 and the other two all 4, often answering "stop" for a deletion. Two requests, the first one counts. Seven messages ask for two things. llama3.2:3b missed 5 of them; qwen2.5:14b missed 1, the salad bowl, where it chose the second request, "cancel", over the first, "refund". A word used against its meaning covers eight odd answers, such as "stop" for a request to move a delivery slot, or "change" for a request to cancel a payment plan. qwen2.5:3b made 5 of them and llama3.2:3b 4; qwen2.5:7b and qwen2.5:14b made none.

The sketch puts the two biggest rules side by side, and it is the most important picture in this lesson. On the question rule, a bigger model was nearly perfect. On the comment rule, size made almost no difference. One possible reason: a model trained to be helpful treats every message as a request, and a thank-you with no request goes against that habit. I cannot test that reason here. What I can say is that the rule is written plainly in the prompt, all four models read it, and none of them kept it reliably: even qwen2.5:14b got it right only 2 times in 13.
Category is the field I expected to be easy, because the prompting chapter's models sorted messages into these four categories very well. With a ticket to fill in, it was harder. I tagged the 40 messages where any model got the category wrong, again by my reading.

The biggest group is payment details: 10 messages about changing a card, cancelling a subscription or a payment plan, paying in instalments, the company name on an invoice, and switching the payment method, all of which the rules put in billing. qwen2.5:3b got 9 of them wrong, every one by sending it to account, as if a card were part of a profile. llama3.2:3b got 6 wrong, and the larger qwen models 1 and 2. That is a rule that size fixed.
But category is also where a bigger model did not simply do better overall. qwen2.5:14b got 120 of 136 right, one fewer than qwen2.5:7b's 121. It made 5 mistakes in the exchange or return process group: four customers asking to swap an item, and one complaint about the courier who collected a return. The rules call all five returns; it called them delivery. It also made 4 mistakes on cancelling an order. qwen2.5:3b made 1 exchange mistake and llama3.2:3b none. Each model drew the lines between the four categories in its own places.

One group needs a warning. The cancelling group is six messages that ask to cancel something, and here the models all leaned the same way. The gold is returns for three of them (cancelling an exchange or a replacement), delivery for two (an order that is late or never sent) and billing for one. The models never said delivery: their 12 wrong answers were billing 9 times, account twice and returns once. So they read "cancel" as a money question, while the rules send it by what is being cancelled. Only one of the six, te125, a plain request to cancel an order before it is dispatched, is the unclear case lesson 2 found in the rules; there the gold says billing and only qwen2.5:3b missed it, with account.
The lost or damaged parcel group shows the models' own habits. Six messages are about a parcel that was lost or arrived damaged, which the rules say stays delivery, even when the customer wants money back. qwen2.5:7b got 5 of the 6 wrong, and every one of its wrong answers was returns, as if a broken item were the start of a return. qwen2.5:3b and qwen2.5:14b got 1 each wrong. The rule is one plain sentence in the prompt.
The item rule asks for the product as a short common noun: lower case, singular, with "no colour, size, brand, quantity or adjective". Because this is mechanical, I did not tag the mistakes by hand. The lab sorts each wrong item by comparing it with the gold word: did the answer end in the gold word with extra words in front, was it the gold word with an "s" added, was it empty, and so on. When extra words were kept, it sorts them once more, with a short list in the code: an adjective or material word (red, big, cordless, leather, wooden), which the rule clearly forbids, or a noun that is part of a compound name (patio heater, gaming chair, printer paper, salad bowl), which the rule does not mention. Which words count as part of a name is my reading, and the list is in the code for you to check.

The two kinds of kept word behave very differently. Adjectives and material words, the clear breaks of the rule, went from 23 on qwen2.5:3b ("red kettle", "silver necklace", "cordless drill") to 11 on llama3.2:3b, 4 on qwen2.5:7b and 1 on qwen2.5:14b ("electric toothbrush"). That part size fixed almost completely. Words that are part of a compound name went from 14 on qwen2.5:3b to 11, 5 and 4. Three of qwen2.5:14b's four are "dining table", "printer paper" and "salad bowl", where the gold keeps only the last noun; the fourth is "printer cartridges". Size fixed that part much less, and the next paragraph shows why it may not be the models' fault.
The next biggest mistake is answering null when a product is named, 9 times on qwen2.5:3b; for example, "what payment methods do you accept for the christmas hampers" gave no item, as if a product in a question did not count. llama3.2:3b's own weak spot was the plural: 6 times it kept an "s", as in "curtains" for curtain.

Two of these examples show a weakness in the rule, not in the models. The gold for "the printer cartridges" is "cartridge", and for "the printer toner" it is "toner", but the gold for "the baby monitor" is "baby monitor". "Printer" and "baby" are nouns, not colours, sizes, brands, quantities or adjectives, so the rule as written does not say whether they stay; the annotators chose differently for different names, and lesson 2 found the same thing with "yoga mat" and "mat". All four models kept "printer cartridges". On the baby monitor, which is in a message asking to stop marketing emails, all four answered null. The two tall glasses show the plural rule: all four models kept "glasses", three of them without the adjective. A fine-tune could learn the annotators' habit from examples; a prompt can only state the rule, and here the rule is not complete.
The urgent rule is also mechanical: true only when the customer gives a time limit or uses one of six words (urgent, asap, today, tonight, right now, immediately). The lab sorts each wrong answer into three kinds by code: a listed word the model missed, a stated deadline it missed, or a false alarm, where it said urgent and the gold says not.

The models failed in different directions. llama3.2:3b missed listed words: 8 times the message said "today" (3 times), "immediately" (twice), "right now", "tonight" or "Urgent", and it answered false. It raised only 1 false alarm. qwen2.5:14b did the reverse: it missed a listed word once and raised 3 false alarms. The favourite false alarm was an angry customer who has waited a long time, such as "I have been waiting 3 weeks for my money back ... Where is it??", which three models called urgent. The rule says anger alone is not urgency, and it says so in so many words.
One message beat every model: "I must ask that all telephone contact from your company ceases immediately." It contains "immediately", so it is urgent by the rules, and all four answered false. One possible reason is that the word describes how fast calls should stop rather than a problem that needs a fast answer, and the models judged the meaning rather than the word. The rule is about the word.

The order number is the rule that held. 77 of the 140 messages have one, and 27 messages also contain other numbers, such as a price, a date, a phone number or a quantity, which the rules say are never an order number. qwen2.5:3b and qwen2.5:14b got the field right on all 140 messages, and the other two on 139. The single miss is worth reading: a message with no order number that says "I ordered 3 packs of printer paper on 2 October". llama3.2:3b copied the date, "2 October". qwen2.5:7b answered "20231002", a date with a year that appears nowhere in the message. That model did not copy a wrong number; it made one up.
So far I have compared the models field by field. Now I look at the same question from the other side: take every field qwen2.5:3b got wrong, and see what qwen2.5:14b did with the same field of the same message. A "field" here means one field of one message; 691 fields are scored across the 140 messages.

qwen2.5:3b got 115 fields wrong. qwen2.5:14b, a model about five times the size, got 79 of those right. It still got 36 of them wrong. And it made 12 new mistakes on fields the small model had right. So the larger model went from 115 wrong fields to 48. That is a big improvement, and it is also far from zero.
The most useful number in this picture is the smallest box. 22 fields were wrong in all four models. No step in size, and no change of maker, fixed them. But before calling them the models' failure, look at where the gold itself is weak.

12 of the 22 are wants, all of them from the comment rule (10) and the not-on-the-list rule (2) on the wants slide. 5 are category: the new card on the account, the salad bowl, the photo of a child at the front door, and two cancellations. 4 are item: the printer cartridges, the boxes of tiles, the baby monitor and the two tall glasses. 1 is urgent: the "ceases immediately" message.
At least 6 of the 22 are problems in the rules or the labels as much as in the models. The printer cartridges and the baby monitor sit in the gap in the item rule. Two are cancellations, where the models' "billing" is a reading the rules' own wording invites. The salad bowl's category was set by the relabelling in lesson 2, after the first two annotators had split on it. And the photo at the front door moved to account in that relabelling as annotator variation, not because of any rule. The other 16 are clearer: mostly comments and requests that the rules plainly call "other". Those, and only those, are what I will look at first when the fine-tuned models arrive in the next lessons. If a small trained model gets them right, it has learned something that four prompted models, up to 14B, did not take from the written rules.
With 131 whole tickets and a few differences of a handful of messages, some of what I have described could be chance. The sign test answers that. It looks only at the messages where two models disagree: the ones where the second model is right and the first wrong ("fixed"), and the reverse ("broke"). If the two models were equally good, fixed and broke would be about even. The further apart they are, the smaller the p value, the chance of a split at least that uneven if the models were really equal.
I made 14 comparisons: the whole ticket for four pairs (llama3.2:3b against qwen2.5:3b, qwen2.5:3b against qwen2.5:7b, qwen2.5:7b against qwen2.5:14b and qwen2.5:3b against qwen2.5:14b), and each of the five fields for two pairs (qwen2.5:3b against qwen2.5:14b, and qwen2.5:7b against qwen2.5:14b). With 14 tests, one could reach the usual 0.05 by chance, so I used the Bonferroni correction: a result counts as beyond luck only if p is below 0.05 divided by 14, which is 0.0036. I chose these 14 comparisons after I had seen the score table, so treat them as a careful look at this data, not a test planned in advance.

Four comparisons are beyond the line. qwen2.5:7b beat qwen2.5:3b on the whole ticket (fixed 36, broke 6), and so did qwen2.5:14b (fixed 47, broke 4). Of the fields, qwen2.5:14b beat qwen2.5:3b on wants (fixed 19, broke 1) and on item (fixed 39, broke 1).
Everything else cannot be told apart from luck on these messages. llama3.2:3b's 54 against qwen2.5:3b's 50 cannot: fixed 13, broke 17, p 0.585, so two makers at 3B are not shown to differ. The step from qwen2.5:7b to qwen2.5:14b is the interesting one. On the whole ticket, 80 against 93, fixed 23 and broke 10, p is 0.035. On wants, fixed 9 and broke 1, p is 0.021. Both are below the usual 0.05, and if I had made only one comparison I might have called them real. But I made 14, and both are well above the corrected line of 0.0036, so neither can be told apart from luck. On category, the larger qwen was one message behind the 7B, 120 against 121, p 1.000. So the honest summary is: going from qwen2.5 3B to 7B or 14B clearly helps, on these messages; going from 7B to 14B may help, but these 131 messages cannot show it.
My first runs of all four models read the house rules before lesson 2's last change. The lab reads the rules file once, when it starts, and those runs started before I saved the tightened clause, so they sent the older wording: "the refund of a cancelled order or a missing item is billing", with no sentence saying that a lost or damaged parcel stays delivery. Their answers were then scored against gold that follows the new wording. I found this by comparing the times the files were written, reran every model with the current rules, and kept the first runs apart, in results/guide_v1/.

The whole-ticket scores barely moved: qwen2.5:3b from 49 to 50, llama3.2:3b from 56 to 54, qwen2.5:7b from 81 to 80, and qwen2.5:14b stayed at 93. On the chapter's 80 older messages the history has one more step, because those messages were run three times. The very first baselines, kept in results/oldguide/, used the rules from before any boundary rule existed: qwen2.5:3b got 10 and 10 on the two halves, llama3.2:3b 11 and 13, qwen2.5:7b 16 and 11. The second round read the older boundary wording: 11 and 13, 16 and 13, 15 and 11. The rerun with the current rules gave 12 and 12, 14 and 13, 15 and 11. qwen2.5:14b and the 0.5B were not rerun on the 80, because they had already read the current rules.
What did move is worth knowing. Out of 140 replies, the rerun changed 16 on qwen2.5:3b, 19 on llama3.2:3b, 3 on qwen2.5:7b and 11 on qwen2.5:14b, counting a reply as changed when any of its five values differs. Most of those messages have nothing to do with the clause that changed. One possible reason: a different sentence in the middle of a 400-word prompt changes what comes after it for the model, and on a close call that can be enough to tip the answer either way. I have not tested that. Right answers were lost about as often as they were gained, which is why the totals hardly moved. The lesson for your own work: a change to a prompt is a change to every reply, not only to the replies the change was meant for, and a score from one run carries that much wobble with it.
To see what the rules can do on their own, I gave the same house rules to the smallest model in the chapter, Qwen2.5 0.5B Instruct, the model later lessons will fine-tune. It ran through Apple's mlx library, not Ollama, and the mlx runner used here has no schema option, so nothing forced its replies into the right shape. It read the same current rules as the others, but its replies were limited to 80 tokens instead of 120, it always took its single most likely next word (greedy decoding), and it is a 4-bit build, compressed to save memory. So it differs from the others in size, in the program that ran it, in the schema and in those settings, and I cannot say how much of its result comes from each.

It got 3 whole tickets right out of 131. Every one of its 140 replies was valid JSON anyway, so the shape was not the problem. The content was. It said urgent on 122 of the 140 messages. 47 times its wants was not one of the seven allowed words: 14 times it wrote "urgent", 6 times it copied a line from the rules themselves ("what the customer is asking us to DO"), and the rest were phrases such as "replace billing address". 24 times its category was not one of the four. It wrote "patriotic heater" for a patio heater. It copied the order number well, though: 134 of 140.
So a 0.5B model that reads 401 words of rules keeps almost none of them. That is the starting point for the lessons: the same model, trained on 500 labelled tickets, has to get from 3 of 131 to somewhere near the 93 of the prompted 14B to be worth the trouble.
Lesson 2 kept the 80 test messages from the prompting chapter as a second, smaller ticket test. 78 of the 80 have all five fields gold, 39 in each half: the "old" 40 from prompting lesson 1 and the "new" 40 from prompting lesson 11. The same models were run on them, with the same current rules.

On these messages the whole ticket was right far less often: 12 of 39 old and 12 of 39 new on qwen2.5:3b, 14 and 13 on llama3.2:3b, 15 and 11 on qwen2.5:7b, and 19 and 14 on qwen2.5:14b. The reason is almost entirely wants, which was right on only 12 to 20 of 39. These messages were written for the sorting task, and most of them state a problem without asking for anything: "I was charged twice for the same order this month." The rules say that is "other". The models guessed the request instead, usually "refund" or "information". Of qwen2.5:14b's 44 wants misses on the 78 messages, 38 had the right answer "other"; for the other three models it was 41 of 48 to 50.
This is the same rule the comment slide found on the 140, showing up much more often because these messages were written differently. One note on the labels: the category gold here was first labelled before any boundary rule existed, so two fresh annotators labelled the category of all 80 again under the current rules, blind, and gave the same category as before on every one of the 80.
This script sends the full house rules and the schema to qwen2.5:3b for four of the 140 test messages, exactly as the lab did, and prints each field next to the right answer. I chose the four so that they show different kinds of mistake: te046, the "ceases immediately" message; te066, which it gets fully right; te114, the thank-you to the driver; and te139, a refund for a returned lamp. The rules in the script are the current TICKET-GUIDE.md, word for word, the same text the lab sent, so that your replies can be compared with the stored ones.

Before you run this lab. It uses qwen2.5:3b, running in Ollama on your own computer. If you have not set that up yet, the lab setup guide shows how to install Ollama, download the model and check that everything works, on macOS, Windows or Linux. You can use a different model instead by changing the MODEL line: llama3.2:3b (ollama pull llama3.2:3b) or, if your computer has the memory, qwen2.5:7b. Your numbers will differ from the ones in this lesson.
"""Four prompted tickets, field by field: the full house rules plus a JSON schema, sent to qwen2.5:3b.
Sends each message to a small model running in Ollama on your own computer, exactly the way the lab did
(the rules as the system message, a JSON schema in "format", temperature 0, seed 1), then prints each
field next to the right answer. Needs only Python 3 and Ollama with qwen2.5:3b:
ollama pull qwen2.5:3b
python prompted_demo.py
"""
import json
import urllib.request
MODEL = "qwen2.5:3b"
FIELDS = ["category", "wants", "order_number", "item", "urgent"]
# The house rules, word for word as in TICKET-GUIDE.md and as the lab sent them.
RULES = """Turn one customer message into one ticket, a JSON object with exactly these five fields.
## category (one of four, same definitions as the prompting chapter)
- billing: charges, payments, invoices, prices and refunds of money.
- delivery: shipping, couriers, tracking and parcels that have not arrived or arrived damaged.
- returns: sending an item back, exchanges and the return process.
- account: signing in, passwords, profile details, privacy and emails from us.
Boundary rule: a refund for an item the customer
has already sent back, or is sending back, is **returns**, because the refund is the last step of the return.
Money that is not tied to sending an item back (a wrong or double charge, an invoice, a price, the refund of a
cancelled order) is **billing**. A parcel that never arrived or arrived damaged stays **delivery**, even when the
customer asks for their money back.
## wants (one of seven): what the customer is asking us to DO
- refund: give money back.
- replacement: send the same item again, or swap it for another (an exchange).
- cancel: stop an order, subscription or payment before it happens or continues.
- change: edit details we hold (address, email, name, phone, password, payment card, delivery slot).
- stop: stop messages from us (emails, texts, calls, newsletters).
- information: they ask a question or for a status, and ask for no other action.
- other: none of the above (a complaint with no request, praise, deleting data, something else).
If the message asks for two actions, choose the one it asks for FIRST.
## order_number (string or null)
The order number exactly as digits, without "#", "order" or spaces ("order #4471" -> "4471").
null when no order number is written. A phone number, price or date is never an order number.
## item (string or null)
The product the message is about, as a short common noun: lower case, singular, no colour, size, brand,
quantity or adjective ("two blue ceramic mugs" -> "mug", "my new Sony headphones" -> "headphones",
"a pair of trainers" -> "trainers"). Keep nouns that are always plural in English (headphones, trainers,
jeans, scissors). null when no product is named ("my order", "the parcel", "it" are not products).
## urgent (true or false)
true only when the customer states a time limit or deadline, or says urgent, asap, today, tonight,
right now, or immediately. Anger alone is not urgency.
## Output
One line of JSON, keys in this order:
{"category": "...", "wants": "...", "order_number": "..." or null, "item": "..." or null, "urgent": true/false}"""
# The JSON schema: Ollama will only let the model write an object of this shape.
SCHEMA = {"type": "object", "properties": {
"category": {"type": "string", "enum": ["billing", "delivery", "returns", "account"]},
"wants": {"type": "string", "enum": ["refund", "replacement", "cancel", "change", "stop", "information", "other"]},
"order_number": {"type": ["string", "null"]}, "item": {"type": ["string", "null"]},
"urgent": {"type": "boolean"}},
"required": FIELDS}
# (id, message, the right ticket: category, wants, order_number, item, urgent)
MESSAGES = [
('te046', 'To whom it may concern, I must ask that all telephone contact from your company ceases immediately. I have registered my number with the preference service. Regards.',
['account', 'stop', None, None, True]),
('te066', 'The courier lost my parcel with the white backpack, order 38829. Refund me asap, I have already bought one elsewhere.',
['delivery', 'refund', '38829', 'backpack', True]),
('te114', 'Hi, please pass a thank you to the driver on order 12876, the big sofa. I paid extra for two-person delivery and it was worth it.',
['delivery', 'other', '12876', 'sofa', False]),
('te139', 'I need the refund for order no. 66390 before Friday, my card bill is due. It was for the grey metal lamp you already collected.',
['returns', 'refund', '66390', 'lamp', True]),
]
def ask(message):
body = {"model": MODEL, "stream": False, "format": SCHEMA,
"messages": [{"role": "system", "content": RULES}, {"role": "user", "content": message}],
"options": {"temperature": 0, "seed": 1, "num_ctx": 8192, "num_predict": 120}}
req = urllib.request.Request("http://localhost:11434/api/chat", data=json.dumps(body).encode(),
headers={"Content-Type": "application/json"})
return json.loads(urllib.request.urlopen(req, timeout=600).read())["message"]["content"]
def same(field, got, right):
"""The lab's scoring: text fields in lower case with spaces trimmed, and an empty text counts as null."""
if field in ("order_number", "item"):
got = None if got in (None, "", "null") else str(got).strip().lower()
return got == right
for mid, message, right in MESSAGES:
reply = ask(message)
ticket = json.loads(reply)
print(f"\n{mid}: {message}")
print(f" reply: {reply}")
for k, field in enumerate(FIELDS):
ok = same(field, ticket[field], right[k])
print(f" {field:<13} {str(ticket[field]):<24} right: {str(right[k]):<12} {'ok' if ok else 'WRONG'}")
This is a real run in VS Code's terminal.

I checked the script against the stored data in code, not by eye. I ran it, saved its output, and the lab's demo mode confirmed three things: the rules in the script are identical to TICKET-GUIDE.md; the four messages and their right answers match the test and gold files; and each of the four replies is the same JSON as the lab's stored reply for that message. All four matched. The prompting chapter found once that the order of calls can change a close answer, because Ollama reuses work from the previous call; here the script sends the four messages in the same order they sit in the test file, as the lab did, and no reply differed. On your own computer, a different Ollama version or chip can still change a close call.

The report lives in scripts/labs/finetune/prompted_report.py. It reads the stored replies of the four prompted models and the 0.5B floor, the test messages and the gold files, and prints every table in this lesson. A json mode writes the same numbers to results/ps-report.json, which is what the figures read.
Before it prints anything, it checks the files. It grades every stored reply again with the lab's own grading function against the current gold and stops if any stored mark differs, and it checks that the replies are in the same order as the test messages. For the history slide it reads the first runs in results/guide_v1/ and grades them twice: against the current gold, where the result must equal the score stored in each file, and against the gold from before lesson 2's relabelling, where it must equal the older score each of those files kept. It also checks that the saved copy of the older rules differs from the current rules only in the one clause, and that every wrong wants and category answer has exactly one of my tags, with no tag on a message nobody got wrong.
Only two modes call a model. The tokens mode made the one call per model that counted the prompt tokens, and stored the result. The student script calls qwen2.5:3b, and the demo mode checks its saved output. A box mode writes the playground on the next slide.
What came after I saw the data: all of this lesson's analysis. The rule tags for wants and category, the code that sorts item and urgent (including the list of name words), the choice of the 14 sign tests, the four demo messages and the examples in the figures were all chosen after I had read the replies. The schema and the settings were fixed before any model ran on the 140 messages. The rules text was not: it changed between the first runs and the reruns, which is what the history slide shows.

This box has no model in it. It holds all 140 test messages, the gold ticket for each, and the ticket each of the four prompted models actually wrote, taken from the stored replies with the lab's own clean-up (text fields in lower case). A gold value of '?' means the annotators disagreed on that field, so it is not scored.
As it is, the box shows two messages and then the score table. The first is te046, "ceases immediately": all four models got every field right except urgent, which all four answered False. The second is te139, the refund for a lamp that was already collected: qwen2.5:3b, llama3.2:3b and qwen2.5:7b said billing where the rules say returns, qwen2.5:3b gave no item at all, and llama3.2:3b missed the deadline "before Friday". Then score() prints the same field table as the lab, from the box's own data: 50, 54, 80 and 93 whole tickets.
Try wrong('wants', 'qwen14b') to list every wants answer qwen2.5:14b got wrong, and read how many of them have the gold answer other. Try wrong('item', 'qwen3b') to see the describing words. Try find('cancel') and then show() with one of the ids it prints, to see the models disagree about cancellations.
The rules and the schema. RULES is the full text of the house rules, word for word as in TICKET-GUIDE.md, the text the lab sent. SCHEMA describes the answer: an object with five keys, category limited to four words, wants limited to seven, order number and item either text or null, and urgent true or false. The schema goes in the request's format field; Ollama then only lets the model produce text that fits it.
The messages. MESSAGES holds four real test messages, each with its id and its gold ticket, the five right values in order. All four are fully gold, so every field is scored.
Asking the model. ask() builds one request, exactly as the lab did: the rules as the system message, the customer's message as the user message, the schema, temperature 0, seed 1, a context of 8,192 tokens (num_ctx, how much text the model can see at once) and a limit of 120 tokens on the reply. It sends the request to Ollama's local address with Python's urllib, so nothing needs installing, and returns the reply text.
Scoring. same() repeats the lab's comparison. For order number and item it turns an empty text or the word "null" into Python's None and puts the text in lower case, so "Lamp" and "lamp" count as the same; for the other fields it compares the value as it is. The loop prints the reply, then one line per field with the model's value, the right value, and ok or WRONG.
To use it on your own task, replace RULES with your own rules, SCHEMA with your own fields, and MESSAGES with a few of your own examples with their right answers. Run it on the model you plan to use, then on the next size up, and read the WRONG lines before you count anything.

Write the rules once, in full, and send them as they are. The same file your labellers used is the best first prompt, because it is the definition of a right answer. Do not shorten it until you have measured it.
Add a schema. Here it made every reply valid JSON of the right shape, on every model. It will not make the content right, but it takes one whole class of failure off the table. The 0.5B, which had no schema, wrote a wants value outside the list 47 times.
Score every field against gold, separately. A single "accuracy" number would have said 93 of 131 for the best model and hidden that the order number was perfect while wants was wrong on 15 of 138. Score only gold fields, and keep the whole-ticket count too, because that is what your users see.
Read the wrong answers and group them by the rule they break. This is the step that turns a score into a decision. Some groups will be the model's fault, some the rule's. The printer cartridges and the baby monitor were the rule's.
Try the next size up, and test the difference. If a bigger model fixes a group, you may not need training for it; you may need a bigger model, at a cost you can now weigh. Use a sign test on the same messages, and correct for how many tests you ran.
What no size fixed: check the rule first. Here that is 22 fields. At least 6 of them come from a gap in the rules or an unsettled label, and those need a better rule, not a trained model. The rest, mostly wants on messages that ask for nothing, are what a fine-tune has to get right.
A prompt is enough for mechanical rules on a big enough model. The order number, copied exactly and never confused with a price or a date, was right on 139 or 140 of 140 on every model, even the smallest prompted one. If your fields are like that, a prompt with a schema may be all you need.
A bigger model is enough for rules it half-knows. Adjectives kept in the item and questions read as requests went from common on the 3B models to rare on qwen2.5:14b. The sign tests covered the whole wants and item fields for qwen2.5:3b against qwen2.5:14b, and both were beyond luck; I did not test the smaller groups inside them, or llama. If your misses look like that, compare the cost of a bigger model with the cost of training a smaller one.
A prompt is not enough for rules that go against the model's habits. A comment that asks for nothing, "other", was wrong 11 times out of 13 even on qwen2.5:14b, and on the chapter's older messages it caused most of the wants misses on every model. More words in the rules may help a little, but the rules already say it plainly. This is the kind of rule is meant for: the model has to learn, from many examples, what your team means by a word.
A prompt is also not enough when the rule itself is incomplete. No model can follow a rule that does not say which words of a product name to keep. The fix there is the rule, or examples that show the annotators' habit, which is one more thing a training set gives you and a prompt cannot.
Remember the running cost. Every call here carried about 648 tokens of prompt for a 25-word message. At a few thousand tickets a day that is a real cost on a hosted model, and a model that has learned the rules would not need to read them every time.

One run, at temperature 0. Each model answered each message once. Temperature 0 makes the reply repeatable on the same set-up, and my one check call per model reproduced the stored replies, but a different seed, Ollama version or computer can change close calls. The numbers are one careful measurement, not an average.
Small prompt changes move replies. The history slide showed that changing one clause of the rules changed between 3 and 19 of 140 replies per model, mostly on messages the clause had nothing to do with, while the whole-ticket totals moved by 2 at most. Read any single score here as having at least that much wobble.
131 fully gold messages. One message is less than 1% of the whole-ticket score. Most differences between neighbouring models could not be told apart from luck, and the step from 7B to 14B on the whole ticket was one of them.
Written and labelled with an AI model's help. The 140 test messages were written for this course by an AI model, and the gold labels come from annotators that were runs of the same model family, as lesson 2 explained. Annotators of one model family probably agree with each other more than two people would, so the gold is cleaner than real labels. Real customers write differently, and people would label some fields differently.
The tags are my reading. The wants and category groups come from my reading of each wrong answer, after I saw the replies. Another reader might put a few messages in a different group; the counts per group could move by a message or two, and the per-model totals would not move at all, because those come straight from the grader.
The floor differs in several ways. The 0.5B ran in the mlx library, and the mlx runner used here has no schema option. It was limited to 80 tokens of reply against the others' 120, and it picked the most likely word every time (greedy decoding, the same idea as temperature 0). It is a 4-bit build of the model, compressed to save memory; the Ollama models are 4-bit builds too. And it is much smaller. So its 3 of 131 cannot be put down to size alone.

If you have a prompt that fills in structured fields from text, you can run this lesson's method on it this week. Put your rules in the system message and a schema in the request. Label 50 to 100 real examples twice, as lesson 2 showed, and score every field separately against the gold. Then read every wrong answer, and for each one write down which rule it breaks. A spreadsheet with one row per wrong field and one column for the rule is enough.
Then run the next size up on the same examples and count, for each group, how many it fixes. The groups a bigger model fixes are a pricing question. The groups it does not fix are your training question, once you have checked that the rule and the labels are clear, and the examples in those groups are the first thing to check after any fine-tune.

The next lesson is planned to train the 0.5B model from the floor slide on the 500 training tickets from lesson 2, with a short one-line prompt instead of the 401-word rules, and score it on the same 140 messages, field by field, against the four prompted models here.
4 questions - Score 80% to pass
Every reply from all four prompted models was valid JSON of the right shape. What made that happen, and what did it not fix?
On the 13 messages that only thank, complain or report and ask for nothing, the rules say wants is 'other'. How did the models do as they got bigger?
qwen2.5:7b got 80 whole tickets right and qwen2.5:14b got 93. The sign test gave fixed 23, broke 10, p 0.035, and I made 14 comparisons in all. What is the honest reading?
The gold item for 'the printer cartridges' is cartridge, but for 'the baby monitor' it is baby monitor. What does this show?