Fine Tuning

Where Prompts Slip: The Same House Rules for Four Models, Checked Field by Field

0 of 23 complete

0%

Contents

Back|Fine TuningWhere Prompts Slip: The Same House Rules for Four Models, Checked Field by Field
1/23
71 min left
Prerequisites
Building a Training Set You Can Trust: House Rules, Two Blind Labellers and a Leak Checkrequired
Related Topics
Asking for a Format: Words, JSON Mode and a SchemaPrompting as EngineeringPutting a Prompt Together: Which Parts Still Earn Their PlacePrompting as EngineeringSmall Rewordings: The Same Instruction, Seven WaysPrompting as EngineeringInstructions Around a Long Document: Where to Put the RulesPrompting as EngineeringPrompts as Code: Templates That a Customer Cannot BreakPrompting as Engineering
1 of 23

Four New Staff, One Briefing

Imagine a shop that takes on four new people for its customer service desk on the same Monday. On the first morning, the manager gives all four the same briefing. She reads out the house rules for writing up a customer's problem: which team it goes to, what the customer is asking for, the order number copied exactly, the product named in one plain word, and whether it is truly urgent. Everyone hears every rule. Everyone gets a printed copy to keep on the desk.

By Friday, the manager looks through the week's write-ups. They are all neat and in the right format, because the form on the screen only accepts the right format. But the content is another matter. One person writes "the blue ceramic mug" where the rule says "mug". Another marks every angry customer as urgent. The most experienced of the four gets almost everything right, and still, when a customer writes in only to say thank you, writes it up as a question.

An illustration of a man standing at a whiteboard with a simple flowchart on it, gesturing towards it, while two people sit at a table facing him, one of them taking notes. Headed one briefing, the same rules for everyone, titled everyone heard the rules; not everyone kept them. Beneath: 4 prompted models, the same 401-word house rules, 140 test messages. Every reply was valid JSON. The whole ticket was right on 50 to 93 of 131.

The useful question for the manager is not "who is best?" It is "which rules are being dropped, and by whom?" A rule that the newest person drops but the experienced person keeps will fix itself with time. A rule that all four drop in the same way is a different problem: the briefing is not enough for it, and something else has to change.

This lesson does exactly that with four language models instead of four people. Each one gets the full house rules from lesson 2 as its instructions, and I check what it writes, field by field, on the 140 test messages. I wanted to know where a prompt alone stops being enough, because that is where training a model starts to be worth its cost.

Eight Words for This Lesson

A hand-drawn list headed eight words for this lesson, titled scoring a prompt, one field at a time. Prompted model: a model used as it is, with only instructions in the prompt, no training. House rules: the written rules for the five fields, sent in full with every message. Schema: a description of the JSON shape; Ollama makes every reply fit it. Field score: for one field, how many messages the model got right. Whole ticket: all five fields right on the same message. Gold: the answer both annotators gave; only gold is scored. Sign test: counts the messages one model fixed and the other broke. Bonferroni: with many tests, divide 0.05 by how many you ran. Beneath: floor, the weakest set-up, to show what the rules alone do.

Prompted model. A model used exactly as it was downloaded. Nothing about it is changed; everything it knows about the task comes from the text I send with each message. In this chapter, the opposite is a fine-tuned model, which is trained on examples.

House rules. The rules for the five ticket fields from lesson 2, written once in one file. Here the whole file is the model's instructions, sent again with every single message.

Schema. A short description of the shape the answer must have: which five keys, which words are allowed for category and wants, and which fields may be empty. Ollama, the program that runs the models on my computer, uses the schema to stop the model from writing anything that does not fit. The prompting chapter introduced this in its lesson on asking for a format.

Field score and whole ticket. A field score counts, for one field, how many messages the model got right. A whole ticket is right only when all five fields are right on the same message. A model can score well on every field and still get few whole tickets right, because its mistakes fall on different messages.

Gold. As in lesson 2, a field is gold only when the two annotators gave it the same value. Only gold fields are scored.

Sign test and Bonferroni. The two tools from the prompting chapter for asking whether a difference between two models is more than luck. The luck slide explains both again.

The floor is the weakest set-up in the lesson: a very small model given the same rules, to show what the rules can do when the model is too small to follow them.

The Same Prompt for Every Model

The set-up is simple, and it is the same for all four models. For each of the 140 test messages, the lab sends Ollama one request with three things in it: the house rules as the system message (the instructions part of a chat), the customer's message as the user message, and the JSON schema. The model writes one ticket. A small grading step then compares each of the five fields with the gold ticket and marks it right or wrong.

A sequence diagram with three columns: the lab, Ollama, the grader. Step 1, the lab sends Ollama the house rules, 401 words. Step 2, the customer message. Step 3, the JSON schema. Step 4, Ollama sends back one ticket, as JSON. Step 5, the lab sends the grader the ticket and the gold. Step 6, the grader marks each field right or wrong. Headed one call, for each of the 140 messages, titled the rules, the message and a schema go in; one ticket comes out. Beneath: temperature 0, seed 1, one run per model. The same call for all four models.

The four models are qwen2.5:3b, llama3.2:3b, qwen2.5:7b and qwen2.5:14b. The number is the size of the model in billions of weights, the numbers a model learns during its original training; a bigger model usually follows instructions better but is slower and needs more memory. Two 3B models from different makers let me compare makers at the same size, and the three qwen models let me compare sizes from the same maker.

Every call used temperature 0 and seed 1. Temperature controls how much randomness goes into picking each next word; at 0 the model always takes its most likely choice, so the same message gives the same reply. I ran each model once on each message. That is one run, not an average over many, and I come back to what that means on the limits slide.

The rules each model read are the final house rules from lesson 2, word for word, including the tightened boundary rule, and the gold they are scored against follows the same rules. That was not true of my first runs, which read an older wording of one clause. I found that, reran all four models, and the history slide later in this lesson shows how much the scores moved. Everything else in this lesson uses the reruns.

What the Rules Cost on Every Call

A prompt with rules in it is not free. The rules are sent again with every message, and the model has to read all of them before it writes a word.

Two panels headed what every single call carries, titled the whole prompt came to about 648 tokens a message. The house rules: 401 words; words with a letter or digit; sent in full, every call. Prompt, message te001: 648 tokens, qwen; llama 653. Beneath: counted by Ollama on one real call per model. The message itself was 25 words.

The house rules are 401 words long, counting every word that has a letter or a digit in it. A model does not read words, though; it reads tokens, small pieces of text, often part of a word. The stored results do not record how many tokens each prompt used, so I made one new call per model, on the first test message, and read Ollama's own count, prompt_eval_count. The whole prompt came to 648 tokens on all three qwen models and 653 on llama. That includes the rules, the 25-word message and a few tokens of chat formatting. Most of it is the rules.

Those same four calls did a second job. Each reply was compared with the stored reply for the same message from the lab's run, and all four were the same. So the stored results can be repeated, at least on that message, with this Ollama version.

I do not report how long any call took. Other labs were using Ollama on the same computer while these ran, so any time I measured would be partly someone else's work. Remember the token count instead: every ticket carries about 600 tokens of rules, and a fine-tuned model that has learned the rules would not need them.

Field by Field

Here is the first result, and it is the good news. Every reply from every model was valid JSON of the right shape: 140 of 140 for all four. That is the schema doing its job. No reply was a sentence to the customer, no key was missing, and no category or wants value was a word outside the list.

A bar chart headed four prompted models, 140 test messages, field by field, titled order number held; wants and item slipped. Groups of four bars, one per model, in the order qwen 3B, llama 3B, qwen 7B, qwen 14B, showing percent right for each field. Category about 82, 88, 89 and 88. Wants about 76, 78, 83 and 89. Order number about 100, 99, 99 and 100. Item about 63, 74, 88 and 91. Urgent about 95, 92, 97 and 97. Beneath: valid JSON: 140 of 140 for every model. Item: 86 of 137 on qwen 3B, 124 on qwen 14B.

The fields tell different stories. The order number was right on 140 of 140 messages for qwen2.5:3b and qwen2.5:14b, and 139 for the other two. Urgent was right on 129 to 136 of 140. Category, scored on the 136 messages where it is gold, ranged from 112 on qwen2.5:3b to 121 on qwen2.5:7b. Wants, scored on 138, went from 105 to 123. Item, scored on 137, varied the most: 86 on qwen2.5:3b and 124 on qwen2.5:14b.

Now the whole ticket. A whole ticket counts only on the 131 messages where all five fields are gold, and it is right only if all five fields match.

Five isometric columns headed all five fields right, of the 131 fully gold messages, titled 3, 50, 54, 80 and 93 of 131. From left: qwen 0.5B, 3, a flat tile; qwen 3B, 50; llama 3B, 54; qwen 7B, 80; qwen 14B, 93, the tallest. Beneath: height is whole tickets right. The 0.5B read the rules with no schema; the other four had one.

qwen2.5:3b got 50 whole tickets right, llama3.2:3b 54, qwen2.5:7b 80 and qwen2.5:14b 93. The 0.5B floor got 3; it has its own slide later. Even the largest model, which gets each field right between 88% and 100% of the time, gets the whole ticket right on 93 of 131, about 71%. Small misses on different fields add up.

A field score of 90% sounds good. But a team that has to fix one ticket in four by hand does not have a working system. So the rest of this lesson reads the wrong answers themselves, field by field, to find which rules they break.

Wants: Guessing What the Customer Will Ask Next

Wants is the field where the rule asks the model for something models find hard: to write down only what the customer asked for, not what they will probably want. I read every message where any model got wants wrong, 44 messages in all, and tagged each one with the rule it tests. These tags are my reading, done after I saw the replies; the next fields are sorted by code instead.

A table headed wants: every wrong answer, by the rule it breaks, titled a comment that asks for nothing: wrong in every size. A comment, thanks or complaint that asks for nothing is other: 13 messages; wrong: qwen 3B 13, llama 3B 12, qwen 7B 13, qwen 14B 11. A question, with no other request, is information: 12 messages; wrong: qwen 3B 9, llama 3B 6, qwen 7B 3, qwen 14B 1. A request that is not one of the six actions (delete, unlock, a copy of data) is other: 4 messages; wrong: 4, 3, 4 and 2. Two requests: choose the one asked for first: 7 messages; wrong: 2, 5, 3 and 1. A word used against its meaning: 8 messages; wrong: 5, 4, 0 and 0. Beneath: tagged by me, reading every wrong answer.

A comment that asks for nothing is other. 13 messages are praise, a complaint or a report with no request in them. The rules say that is "other". All four models got nearly all of them wrong: qwen2.5:3b missed 13, llama3.2:3b 12, qwen2.5:7b 13 and qwen2.5:14b 11. Most often they answered "information", as if a thank-you note were a question.

A question is information. 12 messages only ask a question, such as "why was i charged 45 ... when the price said 39". The rules say that is "information", even though the customer probably wants money back. Here size mattered a lot: 9 wrong on qwen2.5:3b, 6 on llama3.2:3b, 3 on qwen2.5:7b and 1 on qwen2.5:14b. qwen2.5:3b answered "refund" on 7 of its 9; llama3.2:3b answered "refund" on 3 of its 6 and "change" or "replacement" on the other 3.

A two-column page headed real messages where wants went wrong, titled comments read as questions, questions read as requests. te114: Hi, please pass a thank you to the driver on order 12876, the big sofa. I paid extra for two-person delivery and it was ... Right: other. All four models: information. te094: Under data protection law I request a copy of all personal data you hold about me. Please confirm receipt of this ... Right: other. All four: information. te103: I have been waiting 3 weeks for my money back on order 20931, the broken radio. Where is it?? Absolutely useless ... Right: information. qwen 3B, qwen 7B and qwen 14B: refund. te126: Please cancel the exchange order for the lamp, order 16677, and refund the postage. I have decided to keep the one I ... Right: cancel. qwen 3B, llama 3B and qwen 7B: refund. Beneath: the first two ask for nothing: information is wrong. In the last two, a guessed or later request won.

A request that is not on the list is other. Four messages ask for something none of the six actions covers: deleting data, a copy of personal data, unlocking an account. qwen2.5:14b got 2 of the 4 wrong, llama3.2:3b 3 and the other two all 4, often answering "stop" for a deletion. Two requests, the first one counts. Seven messages ask for two things. llama3.2:3b missed 5 of them; qwen2.5:14b missed 1, the salad bowl, where it chose the second request, "cancel", over the first, "refund". A word used against its meaning covers eight odd answers, such as "stop" for a request to move a delivery slot, or "change" for a request to cancel a payment plan. qwen2.5:3b made 5 of them and llama3.2:3b 4; qwen2.5:7b and qwen2.5:14b made none.

A hand-drawn sketch of two rows, headed sketched: two wants rules, four model sizes, titled size fixed one rule and not the other. Top row labelled a question: information (12 messages), with four boxes joined by arrows: qwen 3B: 9, llama 3B: 6, qwen 7B: 3, qwen 14B: 1. Bottom row labelled a comment: other (13 messages): qwen 3B: 13, llama 3B: 12, qwen 7B: 13, qwen 14B: 11. Beneath: numbers are wrong answers, smallest model on the left.

The sketch puts the two biggest rules side by side, and it is the most important picture in this lesson. On the question rule, a bigger model was nearly perfect. On the comment rule, size made almost no difference. One possible reason: a model trained to be helpful treats every message as a request, and a thank-you with no request goes against that habit. I cannot test that reason here. What I can say is that the rule is written plainly in the prompt, all four models read it, and none of them kept it reliably: even qwen2.5:14b got it right only 2 times in 13.

Category: Lines Drawn in Different Places

Category is the field I expected to be easy, because the prompting chapter's models sorted messages into these four categories very well. With a ticket to fill in, it was harder. I tagged the 40 messages where any model got the category wrong, again by my reading.

A grouped bar chart headed category: every wrong answer, by the rule it breaks, titled payment misses fell from 9 to 2; exchange misses rose from 1 to 5. Seven groups of four bars, one per model in the order qwen 3B, llama 3B, qwen 7B, qwen 14B, counting wrong answers. Payment: 9, 6, 1, 2. Cancel: 3, 3, 2, 4. Lost parcel: 1, 3, 5, 1. Exchange: 1, 0, 1, 5. Sent back: 3, 1, 2, 1. Delivery detail: 2, 0, 3, 2. One-off: 5, 3, 1, 1. Beneath: of 136 scored. Payment: qwen 3B 9, qwen 14B 2. Exchange: qwen 3B 1, qwen 14B 5.

The biggest group is payment details: 10 messages about changing a card, cancelling a subscription or a payment plan, paying in instalments, the company name on an invoice, and switching the payment method, all of which the rules put in billing. qwen2.5:3b got 9 of them wrong, every one by sending it to account, as if a card were part of a profile. llama3.2:3b got 6 wrong, and the larger qwen models 1 and 2. That is a rule that size fixed.

But category is also where a bigger model did not simply do better overall. qwen2.5:14b got 120 of 136 right, one fewer than qwen2.5:7b's 121. It made 5 mistakes in the exchange or return process group: four customers asking to swap an item, and one complaint about the courier who collected a return. The rules call all five returns; it called them delivery. It also made 4 mistakes on cancelling an order. qwen2.5:3b made 1 exchange mistake and llama3.2:3b none. Each model drew the lines between the four categories in its own places.

A two-column page headed real messages where category went wrong, titled each model drew the lines in its own places. te054: i need to change the card on my account, new card, old one lost. Right: billing. All four: account. te139: I need the refund for order no. 66390 before Friday, my card bill is due. It was for the grey metal lamp you already ... Right: returns. qwen 3B, llama 3B and qwen 7B: billing. te133: My watch came with a loose buckle. Could you exchange it for a new one? Order 83012. Right: returns. qwen 14B: delivery. te108: Please refund the cracked salad bowl and cancel the other bowl in the same order, #61047. Right: delivery. qwen 3B, llama 3B and qwen 7B: returns; qwen 14B: billing. Beneath: a new card: account to all four. The salad bowl: none of the four said delivery.

One group needs a warning. The cancelling group is six messages that ask to cancel something, and here the models all leaned the same way. The gold is returns for three of them (cancelling an exchange or a replacement), delivery for two (an order that is late or never sent) and billing for one. The models never said delivery: their 12 wrong answers were billing 9 times, account twice and returns once. So they read "cancel" as a money question, while the rules send it by what is being cancelled. Only one of the six, te125, a plain request to cancel an order before it is dispatched, is the unclear case lesson 2 found in the rules; there the gold says billing and only qwen2.5:3b missed it, with account.

The lost or damaged parcel group shows the models' own habits. Six messages are about a parcel that was lost or arrived damaged, which the rules say stays delivery, even when the customer wants money back. qwen2.5:7b got 5 of the 6 wrong, and every one of its wrong answers was returns, as if a broken item were the start of a return. qwen2.5:3b and qwen2.5:14b got 1 each wrong. The rule is one plain sentence in the prompt.

Item: One Word, Not a Description

The item rule asks for the product as a short common noun: lower case, singular, with "no colour, size, brand, quantity or adjective". Because this is mechanical, I did not tag the mistakes by hand. The lab sorts each wrong item by comparing it with the gold word: did the answer end in the gold word with extra words in front, was it the gold word with an "s" added, was it empty, and so on. When extra words were kept, it sorts them once more, with a short list in the code: an adjective or material word (red, big, cordless, leather, wooden), which the rule clearly forbids, or a noun that is part of a compound name (patio heater, gaming chair, printer paper, salad bowl), which the rule does not mention. Which words count as part of a name is my reading, and the list is in the code for you to check.

A grouped bar chart headed item: every wrong answer, sorted by code, titled adjectives kept: 23 on qwen 3B, 1 on qwen 14B. Six groups of four bars, one per model in the order qwen 3B, llama 3B, qwen 7B, qwen 14B. Adjective: 23, 11, 4, 1. Name word: 14, 11, 5, 4. Null: 9, 4, 5, 4. Plural: 4, 6, 2, 3. Other noun: 0, 2, 1, 1. Not a product: 1, 2, 0, 0. Beneath: item right: qwen 3B 86, llama 3B 101, qwen 7B 120, qwen 14B 124 of 137.

The two kinds of kept word behave very differently. Adjectives and material words, the clear breaks of the rule, went from 23 on qwen2.5:3b ("red kettle", "silver necklace", "cordless drill") to 11 on llama3.2:3b, 4 on qwen2.5:7b and 1 on qwen2.5:14b ("electric toothbrush"). That part size fixed almost completely. Words that are part of a compound name went from 14 on qwen2.5:3b to 11, 5 and 4. Three of qwen2.5:14b's four are "dining table", "printer paper" and "salad bowl", where the gold keeps only the last noun; the fourth is "printer cartridges". Size fixed that part much less, and the next paragraph shows why it may not be the models' fault.

The next biggest mistake is answering null when a product is named, 9 times on qwen2.5:3b; for example, "what payment methods do you accept for the christmas hampers" gave no item, as if a product in a question did not count. llama3.2:3b's own weak spot was the plural: 6 times it kept an "s", as in "curtains" for curtain.

An editorial frame headed the item rule: a short common noun, singular, titled no colour, size, brand, quantity or adjective. A zone labelled four real items holds four boxes. te005, the patio heater: right, heater. qwen 3B, llama 3B and qwen 7B: patio heater. te107, the two tall glasses: right, glass. qwen 3B: tall glasses; llama 3B, qwen 7B and qwen 14B: glasses. te040, the printer cartridges: right, cartridge. All four models: printer cartridges. te061, the baby monitor: right, baby monitor. All four models: null. Beneath: printer and baby are nouns, not adjectives. The rule does not say whether a noun in front of a noun stays.

Two of these examples show a weakness in the rule, not in the models. The gold for "the printer cartridges" is "cartridge", and for "the printer toner" it is "toner", but the gold for "the baby monitor" is "baby monitor". "Printer" and "baby" are nouns, not colours, sizes, brands, quantities or adjectives, so the rule as written does not say whether they stay; the annotators chose differently for different names, and lesson 2 found the same thing with "yoga mat" and "mat". All four models kept "printer cartridges". On the baby monitor, which is in a message asking to stop marketing emails, all four answered null. The two tall glasses show the plural rule: all four models kept "glasses", three of them without the adjective. A fine-tune could learn the annotators' habit from examples; a prompt can only state the rule, and here the rule is not complete.

Urgent, and the Order Number That Held

The urgent rule is also mechanical: true only when the customer gives a time limit or uses one of six words (urgent, asap, today, tonight, right now, immediately). The lab sorts each wrong answer into three kinds by code: a listed word the model missed, a stated deadline it missed, or a false alarm, where it said urgent and the gold says not.

Three panels headed urgent: every wrong answer, sorted by code, titled a listed word missed: llama 3B 8 times, qwen 14B 1. Missed a listed word: 10 msgs; q3B 4, l3B 8, q7B 2, q14B 1. Missed a deadline: 2 msgs; q3B 1, l3B 2, q7B 0, q14B 0. False alarm: 6 msgs; q3B 2, l3B 1, q7B 2, q14B 3. Beneath: te046, "ceases immediately": all four said not urgent. q = qwen, l = llama; the number is wrong answers.

The models failed in different directions. llama3.2:3b missed listed words: 8 times the message said "today" (3 times), "immediately" (twice), "right now", "tonight" or "Urgent", and it answered false. It raised only 1 false alarm. qwen2.5:14b did the reverse: it missed a listed word once and raised 3 false alarms. The favourite false alarm was an angry customer who has waited a long time, such as "I have been waiting 3 weeks for my money back ... Where is it??", which three models called urgent. The rule says anger alone is not urgency, and it says so in so many words.

One message beat every model: "I must ask that all telephone contact from your company ceases immediately." It contains "immediately", so it is urgent by the rules, and all four answered false. One possible reason is that the word describes how fast calls should stop rather than a problem that needs a fast answer, and the models judged the meaning rather than the word. The rule is about the word.

A sidebar headed order number: the rule that held, titled 140 of 140 on qwen 14B. Has one: 77 of 140 messages. Other digits: 27 messages also carry a price, date, phone number or quantity. Right: qwen 3B 140, llama 3B 139, qwen 7B 139, qwen 14B 140. The one miss: te056, no order number, says "on 2 October": llama 3B 2 October; qwen 7B 20231002. Beneath: one model wrote a year the message never gave.

The order number is the rule that held. 77 of the 140 messages have one, and 27 messages also contain other numbers, such as a price, a date, a phone number or a quantity, which the rules say are never an order number. qwen2.5:3b and qwen2.5:14b got the field right on all 140 messages, and the other two on 139. The single miss is worth reading: a message with no order number that says "I ordered 3 packs of printer paper on 2 October". llama3.2:3b copied the date, "2 October". qwen2.5:7b answered "20231002", a date with a year that appears nowhere in the message. That model did not copy a wrong number; it made one up.

Does a Bigger Model Fix It?

So far I have compared the models field by field. Now I look at the same question from the other side: take every field qwen2.5:3b got wrong, and see what qwen2.5:14b did with the same field of the same message. A "field" here means one field of one message; 691 fields are scored across the 140 messages.

A flowchart headed qwen 3B's wrong fields, looked up in qwen 14B, titled the bigger model fixed 79 of 115. qwen 3B: 115 wrong fields leads to right in qwen 14B: 79, and to still wrong in qwen 14B: 36, which leads to wrong in all four models: 22. A separate box: right in 3B, wrong in 14B: 12. Beneath: a field is one field of one message: 691 fields are scored across the 140 messages.

qwen2.5:3b got 115 fields wrong. qwen2.5:14b, a model about five times the size, got 79 of those right. It still got 36 of them wrong. And it made 12 new mistakes on fields the small model had right. So the larger model went from 115 wrong fields to 48. That is a big improvement, and it is also far from zero.

The most useful number in this picture is the smallest box. 22 fields were wrong in all four models. No step in size, and no change of maker, fixed them. But before calling them the models' failure, look at where the gold itself is weak.

A two-column page headed the 22 fields every model got wrong, titled wants 12, category 5, item 4, urgent 1. te094, wants: Under data protection law I request a copy of all personal data you hold about me. Please confirm receipt of ... Right: other. All four models: information. te064, wants: I just want to report that your courier has been parking across our driveway every time he delivers on our ... Right: other. All four: information. te040, item: courier lost order 44102, i need the printer cartridges by tomorrow for my exams, send again right now. Right: cartridge. All four: printer cartridges. te046, urgent: To whom it may concern, I must ask that all telephone contact from your company ceases immediately. I have ... Right: true. All four: false. Beneath: wants: 12 of the 22. At least 6 of the 22 are gaps in the rules or labels.

12 of the 22 are wants, all of them from the comment rule (10) and the not-on-the-list rule (2) on the wants slide. 5 are category: the new card on the account, the salad bowl, the photo of a child at the front door, and two cancellations. 4 are item: the printer cartridges, the boxes of tiles, the baby monitor and the two tall glasses. 1 is urgent: the "ceases immediately" message.

At least 6 of the 22 are problems in the rules or the labels as much as in the models. The printer cartridges and the baby monitor sit in the gap in the item rule. Two are cancellations, where the models' "billing" is a reading the rules' own wording invites. The salad bowl's category was set by the relabelling in lesson 2, after the first two annotators had split on it. And the photo at the front door moved to account in that relabelling as annotator variation, not because of any rule. The other 16 are clearer: mostly comments and requests that the rules plainly call "other". Those, and only those, are what I will look at first when the fine-tuned models arrive in the next lessons. If a small trained model gets them right, it has learned something that four prompted models, up to 14B, did not take from the written rules.

Which Differences Are More Than Luck

With 131 whole tickets and a few differences of a handful of messages, some of what I have described could be chance. The sign test answers that. It looks only at the messages where two models disagree: the ones where the second model is right and the first wrong ("fixed"), and the reverse ("broke"). If the two models were equally good, fixed and broke would be about even. The further apart they are, the smaller the p value, the chance of a split at least that uneven if the models were really equal.

I made 14 comparisons: the whole ticket for four pairs (llama3.2:3b against qwen2.5:3b, qwen2.5:3b against qwen2.5:7b, qwen2.5:7b against qwen2.5:14b and qwen2.5:3b against qwen2.5:14b), and each of the five fields for two pairs (qwen2.5:3b against qwen2.5:14b, and qwen2.5:7b against qwen2.5:14b). With 14 tests, one could reach the usual 0.05 by chance, so I used the Bonferroni correction: a result counts as beyond luck only if p is below 0.05 divided by 14, which is 0.0036. I chose these 14 comparisons after I had seen the score table, so treat them as a careful look at this data, not a test planned in advance.

A table headed message by message, 14 sign tests, titled beyond luck needs p below 0.0036, with one line per test, all 14: ticket, llama 3B to qwen 3B: 54 to 50, fixed 13, broke 17, p 0.585; ticket, qwen 3B to qwen 7B: 50 to 80, fixed 36, broke 6, p below 0.0001, beyond; ticket, qwen 7B to qwen 14B: 80 to 93, fixed 23, broke 10, p 0.035; ticket, qwen 3B to qwen 14B: 50 to 93, fixed 47, broke 4, p below 0.0001, beyond; category, qwen 3B to qwen 14B: 112 to 120, fixed 16, broke 8, p 0.152; wants, qwen 3B to qwen 14B: 105 to 123, fixed 19, broke 1, p below 0.0001, beyond; order number, qwen 3B to qwen 14B: 140 to 140, fixed 0, broke 0, p 1.000; item, qwen 3B to qwen 14B: 86 to 124, fixed 39, broke 1, p below 0.0001, beyond; urgent, qwen 3B to qwen 14B: 133 to 136, fixed 5, broke 2, p 0.453; category, qwen 7B to qwen 14B: 121 to 120, fixed 6, broke 7, p 1.000; wants, qwen 7B to qwen 14B: 115 to 123, fixed 9, broke 1, p 0.021; order number, qwen 7B to qwen 14B: 139 to 140, fixed 1, broke 0, p 1.000; item, qwen 7B to qwen 14B: 120 to 124, fixed 12, broke 8, p 0.503; urgent, qwen 7B to qwen 14B: 136 to 136, fixed 2, broke 2, p 1.000. Beneath: all 14 shown. 4 are beyond the line 0.05 / 14 = 0.0036. q = qwen, l = llama.

Four comparisons are beyond the line. qwen2.5:7b beat qwen2.5:3b on the whole ticket (fixed 36, broke 6), and so did qwen2.5:14b (fixed 47, broke 4). Of the fields, qwen2.5:14b beat qwen2.5:3b on wants (fixed 19, broke 1) and on item (fixed 39, broke 1).

Everything else cannot be told apart from luck on these messages. llama3.2:3b's 54 against qwen2.5:3b's 50 cannot: fixed 13, broke 17, p 0.585, so two makers at 3B are not shown to differ. The step from qwen2.5:7b to qwen2.5:14b is the interesting one. On the whole ticket, 80 against 93, fixed 23 and broke 10, p is 0.035. On wants, fixed 9 and broke 1, p is 0.021. Both are below the usual 0.05, and if I had made only one comparison I might have called them real. But I made 14, and both are well above the corrected line of 0.0036, so neither can be told apart from luck. On category, the larger qwen was one message behind the 7B, 120 against 121, p 1.000. So the honest summary is: going from qwen2.5 3B to 7B or 14B clearly helps, on these messages; going from 7B to 14B may help, but these 131 messages cannot show it.

History: My First Runs Read an Older Wording

My first runs of all four models read the house rules before lesson 2's last change. The lab reads the rules file once, when it starts, and those runs started before I saved the tightened clause, so they sent the older wording: "the refund of a cancelled order or a missing item is billing", with no sentence saying that a lost or damaged parcel stays delivery. Their answers were then scored against gold that follows the new wording. I found this by comparing the times the files were written, reran every model with the current rules, and kept the first runs apart, in results/guide_v1/.

Four panels headed whole ticket: the first runs, with the older rules, then the rerun, titled rerun with the current rules: 50, 54, 80, 93. Qwen 3B: 50; first run 49; 16 replies changed. Llama 3B: 54; first run 56; 19 replies changed. Qwen 7B: 80; first run 81; 3 replies changed. Qwen 14B: 93; first run 93; 11 replies changed. Beneath: of 131 fully gold messages. Only one clause of the rules differed between the two runs.

The whole-ticket scores barely moved: qwen2.5:3b from 49 to 50, llama3.2:3b from 56 to 54, qwen2.5:7b from 81 to 80, and qwen2.5:14b stayed at 93. On the chapter's 80 older messages the history has one more step, because those messages were run three times. The very first baselines, kept in results/oldguide/, used the rules from before any boundary rule existed: qwen2.5:3b got 10 and 10 on the two halves, llama3.2:3b 11 and 13, qwen2.5:7b 16 and 11. The second round read the older boundary wording: 11 and 13, 16 and 13, 15 and 11. The rerun with the current rules gave 12 and 12, 14 and 13, 15 and 11. qwen2.5:14b and the 0.5B were not rerun on the 80, because they had already read the current rules.

What did move is worth knowing. Out of 140 replies, the rerun changed 16 on qwen2.5:3b, 19 on llama3.2:3b, 3 on qwen2.5:7b and 11 on qwen2.5:14b, counting a reply as changed when any of its five values differs. Most of those messages have nothing to do with the clause that changed. One possible reason: a different sentence in the middle of a 400-word prompt changes what comes after it for the model, and on a close call that can be enough to tip the answer either way. I have not tested that. Right answers were lost about as often as they were gained, which is why the totals hardly moved. The lesson for your own work: a change to a prompt is a change to every reply, not only to the replies the change was meant for, and a score from one run carries that much wobble with it.

The Floor: a 0.5B Model With the Same Rules

To see what the rules can do on their own, I gave the same house rules to the smallest model in the chapter, Qwen2.5 0.5B Instruct, the model later lessons will fine-tune. It ran through Apple's mlx library, not Ollama, and the mlx runner used here has no schema option, so nothing forced its replies into the right shape. It read the same current rules as the others, but its replies were limited to 80 tokens instead of 120, it always took its single most likely next word (greedy decoding), and it is a 4-bit build, compressed to save memory. So it differs from the others in size, in the program that ran it, in the schema and in those settings, and I cannot say how much of its result comes from each.

A two-column page headed the floor: qwen 0.5B, the same rules, no schema, titled 3 whole tickets of 131. te001: Please change the billing address on the invoice for order 51234, the ergonomic office chair, to my company ... The 0.5B: category billing; wants replace billing address; item invoice; urgent true. te005: you charged me 64.00 for order 8812 the patio heater but it was cancelled. refund today please i need that ... Category billing; wants refund; item patriotic heater; urgent true. te002: The courier leave the package in the rain and the book is all wet now. Please can you send me new one? Order ... Category delivery; wants urgent; item book; urgent true. Beneath: urgent true on 122 of 140. Wants not one of the seven: 47.

It got 3 whole tickets right out of 131. Every one of its 140 replies was valid JSON anyway, so the shape was not the problem. The content was. It said urgent on 122 of the 140 messages. 47 times its wants was not one of the seven allowed words: 14 times it wrote "urgent", 6 times it copied a line from the rules themselves ("what the customer is asking us to DO"), and the rest were phrases such as "replace billing address". 24 times its category was not one of the four. It wrote "patriotic heater" for a patio heater. It copied the order number well, though: 134 of 140.

So a 0.5B model that reads 401 words of rules keeps almost none of them. That is the starting point for the lessons: the same model, trained on 500 labelled tickets, has to get from 3 of 131 to somewhere near the 93 of the prompted 14B to be worth the trouble.

The Chapter's 80 Older Messages

Lesson 2 kept the 80 test messages from the prompting chapter as a second, smaller ticket test. 78 of the 80 have all five fields gold, 39 in each half: the "old" 40 from prompting lesson 1 and the "new" 40 from prompting lesson 11. The same models were run on them, with the same current rules.

A bar chart headed the chapter's 80 older sorting messages, 39 fully gold in each half, titled wants fell apart on messages that ask for nothing. Pairs of bars for each model on the old 40, whole ticket then wants, out of 39: qwen 3B 12 and 16; llama 3B 14 and 16; qwen 7B 15 and 16; qwen 14B 19 and 20; qwen 0.5B 1 and 4. Beneath: qwen 14B: 38 of its 44 wants misses (both halves) have the right answer other.

On these messages the whole ticket was right far less often: 12 of 39 old and 12 of 39 new on qwen2.5:3b, 14 and 13 on llama3.2:3b, 15 and 11 on qwen2.5:7b, and 19 and 14 on qwen2.5:14b. The reason is almost entirely wants, which was right on only 12 to 20 of 39. These messages were written for the sorting task, and most of them state a problem without asking for anything: "I was charged twice for the same order this month." The rules say that is "other". The models guessed the request instead, usually "refund" or "information". Of qwen2.5:14b's 44 wants misses on the 78 messages, 38 had the right answer "other"; for the other three models it was 41 of 48 to 50.

This is the same rule the comment slide found on the 140, showing up much more often because these messages were written differently. One note on the labels: the category gold here was first labelled before any boundary rule existed, so two fresh annotators labelled the category of all 80 again under the current rules, blind, and gave the same category as before on every one of the 80.

Try It Yourself

This script sends the full house rules and the schema to qwen2.5:3b for four of the 140 test messages, exactly as the lab did, and prints each field next to the right answer. I chose the four so that they show different kinds of mistake: te046, the "ceases immediately" message; te066, which it gets fully right; te114, the thank-you to the driver; and te139, a refund for a returned lamp. The rules in the script are the current TICKET-GUIDE.md, word for word, the same text the lab sent, so that your replies can be compared with the stored ones.

A real screenshot of VS Code with prompted_demo.py open, showing the top of the file: the docstring, the imports, and the start of the house rules text. The rest of the file is in the code box on this slide. Beneath: copy it from the box on the slide.

Before you run this lab. It uses qwen2.5:3b, running in Ollama on your own computer. If you have not set that up yet, the lab setup guide shows how to install Ollama, download the model and check that everything works, on macOS, Windows or Linux. You can use a different model instead by changing the MODEL line: llama3.2:3b (ollama pull llama3.2:3b) or, if your computer has the memory, qwen2.5:7b. Your numbers will differ from the ones in this lesson.

"""Four prompted tickets, field by field: the full house rules plus a JSON schema, sent to qwen2.5:3b.

Sends each message to a small model running in Ollama on your own computer, exactly the way the lab did
(the rules as the system message, a JSON schema in "format", temperature 0, seed 1), then prints each
field next to the right answer. Needs only Python 3 and Ollama with qwen2.5:3b:
    ollama pull qwen2.5:3b
    python prompted_demo.py
"""
import json
import urllib.request

MODEL = "qwen2.5:3b"
FIELDS = ["category", "wants", "order_number", "item", "urgent"]

# The house rules, word for word as in TICKET-GUIDE.md and as the lab sent them.
RULES = """Turn one customer message into one ticket, a JSON object with exactly these five fields.

## category (one of four, same definitions as the prompting chapter)
- billing: charges, payments, invoices, prices and refunds of money.
- delivery: shipping, couriers, tracking and parcels that have not arrived or arrived damaged.
- returns: sending an item back, exchanges and the return process.
- account: signing in, passwords, profile details, privacy and emails from us.

Boundary rule: a refund for an item the customer
has already sent back, or is sending back, is **returns**, because the refund is the last step of the return.
Money that is not tied to sending an item back (a wrong or double charge, an invoice, a price, the refund of a
cancelled order) is **billing**. A parcel that never arrived or arrived damaged stays **delivery**, even when the
customer asks for their money back.

## wants (one of seven): what the customer is asking us to DO
- refund: give money back.
- replacement: send the same item again, or swap it for another (an exchange).
- cancel: stop an order, subscription or payment before it happens or continues.
- change: edit details we hold (address, email, name, phone, password, payment card, delivery slot).
- stop: stop messages from us (emails, texts, calls, newsletters).
- information: they ask a question or for a status, and ask for no other action.
- other: none of the above (a complaint with no request, praise, deleting data, something else).
If the message asks for two actions, choose the one it asks for FIRST.

## order_number (string or null)
The order number exactly as digits, without "#", "order" or spaces ("order #4471" -> "4471").
null when no order number is written. A phone number, price or date is never an order number.

## item (string or null)
The product the message is about, as a short common noun: lower case, singular, no colour, size, brand,
quantity or adjective ("two blue ceramic mugs" -> "mug", "my new Sony headphones" -> "headphones",
"a pair of trainers" -> "trainers"). Keep nouns that are always plural in English (headphones, trainers,
jeans, scissors). null when no product is named ("my order", "the parcel", "it" are not products).

## urgent (true or false)
true only when the customer states a time limit or deadline, or says urgent, asap, today, tonight,
right now, or immediately. Anger alone is not urgency.

## Output
One line of JSON, keys in this order:
{"category": "...", "wants": "...", "order_number": "..." or null, "item": "..." or null, "urgent": true/false}"""

# The JSON schema: Ollama will only let the model write an object of this shape.
SCHEMA = {"type": "object", "properties": {
    "category": {"type": "string", "enum": ["billing", "delivery", "returns", "account"]},
    "wants": {"type": "string", "enum": ["refund", "replacement", "cancel", "change", "stop", "information", "other"]},
    "order_number": {"type": ["string", "null"]}, "item": {"type": ["string", "null"]},
    "urgent": {"type": "boolean"}},
    "required": FIELDS}

# (id, message, the right ticket: category, wants, order_number, item, urgent)
MESSAGES = [
    ('te046', 'To whom it may concern, I must ask that all telephone contact from your company ceases immediately. I have registered my number with the preference service. Regards.',
     ['account', 'stop', None, None, True]),
    ('te066', 'The courier lost my parcel with the white backpack, order 38829. Refund me asap, I have already bought one elsewhere.',
     ['delivery', 'refund', '38829', 'backpack', True]),
    ('te114', 'Hi, please pass a thank you to the driver on order 12876, the big sofa. I paid extra for two-person delivery and it was worth it.',
     ['delivery', 'other', '12876', 'sofa', False]),
    ('te139', 'I need the refund for order no. 66390 before Friday, my card bill is due. It was for the grey metal lamp you already collected.',
     ['returns', 'refund', '66390', 'lamp', True]),
]


def ask(message):
    body = {"model": MODEL, "stream": False, "format": SCHEMA,
            "messages": [{"role": "system", "content": RULES}, {"role": "user", "content": message}],
            "options": {"temperature": 0, "seed": 1, "num_ctx": 8192, "num_predict": 120}}
    req = urllib.request.Request("http://localhost:11434/api/chat", data=json.dumps(body).encode(),
                                 headers={"Content-Type": "application/json"})
    return json.loads(urllib.request.urlopen(req, timeout=600).read())["message"]["content"]


def same(field, got, right):
    """The lab's scoring: text fields in lower case with spaces trimmed, and an empty text counts as null."""
    if field in ("order_number", "item"):
        got = None if got in (None, "", "null") else str(got).strip().lower()
    return got == right


for mid, message, right in MESSAGES:
    reply = ask(message)
    ticket = json.loads(reply)
    print(f"\n{mid}: {message}")
    print(f"  reply: {reply}")
    for k, field in enumerate(FIELDS):
        ok = same(field, ticket[field], right[k])
        print(f"  {field:<13} {str(ticket[field]):<24} right: {str(right[k]):<12} {'ok' if ok else 'WRONG'}")

This is a real run in VS Code's terminal.

A real screenshot of VS Code's terminal after running python prompted_demo.py. For each of four messages it prints the message, the raw reply and one line per field with the model's value, the right value and ok or WRONG. te046: urgent False, right True, WRONG; the other four fields ok. te066: all five fields ok. te114: wants information, right other, WRONG; item big sofa, right sofa, WRONG. te139: category billing, right returns, WRONG; item None, right lamp, WRONG.

I checked the script against the stored data in code, not by eye. I ran it, saved its output, and the lab's demo mode confirmed three things: the rules in the script are identical to TICKET-GUIDE.md; the four messages and their right answers match the test and gold files; and each of the four replies is the same JSON as the lab's stored reply for that message. All four matched. The prompting chapter found once that the order of calls can change a close answer, because Ollama reuses work from the previous call; here the script sends the four messages in the same order they sit in the test file, as the lab did, and no reply differed. On your own computer, a different Ollama version or chip can still change a close call.

The Lab Report

A real terminal session of python prompted_report.py, recorded with asciinema and rendered to an image; twelve numbered sections. Section 1, the prompt: rules 401 words, temperature 0, seed 1, one run per model; 131 of 140 messages fully gold; the 0.5B had no schema; prompt tokens on te001, qwen 648 and llama 653; all four replies the same as stored. Section 2, whole ticket of 131: qwen3B 50, llama3B 54, qwen7B 80, qwen14B 93, 0.5B 3, with every field's score beside it and valid JSON 140 for all. Section 3, history: first run to rerun, qwen3B 49 to 50, llama3B 56 to 54, qwen7B 81 to 80, qwen14B 93 to 93, with 16, 19, 3 and 11 replies changed; then, for the 80 older messages, the baselines from before the boundary rule and the first run to the rerun. Sections 4 to 7, tables of wrong answers by rule for wants, category, item and urgent, the same counts as the figures above; item now splits kept adjectives, 23, 11, 4 and 1, from kept name words, 14, 11, 5 and 4. Section 8, order number: 77 have one, 27 carry other digits, the only miss te056. Section 9: qwen3B wrong on 115 fields, qwen14B fixes 79, keeps 36, adds 12; 22 wrong in all four. Section 10, 14 sign tests with the line 0.0036, four marked beyond luck. Section 11, the 80 older messages, ticket old and new: qwen3B 12 and 12, llama3B 14 and 13, qwen7B 15 and 11, qwen14B 19 and 14, 0.5B 1 and 0. Section 12, the floor: urgent true on 122 of 140. Beneath: the lab's own report. It calls no model.

The report lives in scripts/labs/finetune/prompted_report.py. It reads the stored replies of the four prompted models and the 0.5B floor, the test messages and the gold files, and prints every table in this lesson. A json mode writes the same numbers to results/ps-report.json, which is what the figures read.

Before it prints anything, it checks the files. It grades every stored reply again with the lab's own grading function against the current gold and stops if any stored mark differs, and it checks that the replies are in the same order as the test messages. For the history slide it reads the first runs in results/guide_v1/ and grades them twice: against the current gold, where the result must equal the score stored in each file, and against the gold from before lesson 2's relabelling, where it must equal the older score each of those files kept. It also checks that the saved copy of the older rules differs from the current rules only in the one clause, and that every wrong wants and category answer has exactly one of my tags, with no tag on a message nobody got wrong.

Only two modes call a model. The tokens mode made the one call per model that counted the prompt tokens, and stored the result. The student script calls qwen2.5:3b, and the demo mode checks its saved output. A box mode writes the playground on the next slide.

What came after I saw the data: all of this lesson's analysis. The rule tags for wants and category, the code that sorts item and urgent (including the list of name words), the choice of the 14 sign tests, the four demo messages and the examples in the figures were all chosen after I had read the replies. The schema and the settings were fixed before any model ran on the 140 messages. The rules text was not: it changed between the first runs and the reruns, which is what the history slide shows.

Three brand cards headed the tools, with their logos, titled what ran where. Ollama: the four prompted models, with the schema. Apple mlx: the 0.5B floor, no schema. Python: the report, the demo, the box.

Pick a Message and Compare the Models

This box has no model in it. It holds all 140 test messages, the gold ticket for each, and the ticket each of the four prompted models actually wrote, taken from the stored replies with the lab's own clean-up (text fields in lower case). A gold value of '?' means the annotators disagreed on that field, so it is not scored.

As it is, the box shows two messages and then the score table. The first is te046, "ceases immediately": all four models got every field right except urgent, which all four answered False. The second is te139, the refund for a lamp that was already collected: qwen2.5:3b, llama3.2:3b and qwen2.5:7b said billing where the rules say returns, qwen2.5:3b gave no item at all, and llama3.2:3b missed the deadline "before Friday". Then score() prints the same field table as the lab, from the box's own data: 50, 54, 80 and 93 whole tickets.

Try wrong('wants', 'qwen14b') to list every wants answer qwen2.5:14b got wrong, and read how many of them have the gold answer other. Try wrong('item', 'qwen3b') to see the describing words. Try find('cancel') and then show() with one of the ids it prints, to see the models disagree about cancellations.

The Code, Part by Part

The rules and the schema. RULES is the full text of the house rules, word for word as in TICKET-GUIDE.md, the text the lab sent. SCHEMA describes the answer: an object with five keys, category limited to four words, wants limited to seven, order number and item either text or null, and urgent true or false. The schema goes in the request's format field; Ollama then only lets the model produce text that fits it.

The messages. MESSAGES holds four real test messages, each with its id and its gold ticket, the five right values in order. All four are fully gold, so every field is scored.

Asking the model. ask() builds one request, exactly as the lab did: the rules as the system message, the customer's message as the user message, the schema, temperature 0, seed 1, a context of 8,192 tokens (num_ctx, how much text the model can see at once) and a limit of 120 tokens on the reply. It sends the request to Ollama's local address with Python's urllib, so nothing needs installing, and returns the reply text.

Scoring. same() repeats the lab's comparison. For order number and item it turns an empty text or the word "null" into Python's None and puts the text in lower case, so "Lamp" and "lamp" count as the same; for the other fields it compares the value as it is. The loop prints the reply, then one line per field with the model's value, the right value, and ok or WRONG.

To use it on your own task, replace RULES with your own rules, SCHEMA with your own fields, and MESSAGES with a few of your own examples with their right answers. Run it on the model you plan to use, then on the next size up, and read the WRONG lines before you count anything.

How to Find Where Your Prompt Slips

A hand-sketched column of six boxes joined by arrows, headed before you decide a prompt is not enough, titled six steps, in this order. 1, write the rules once, in full. 2, add a schema for the shape. 3, score every field against gold. 4, read the misses; group them by rule. 5, try a bigger model; sign test it. 6, what still fails: fix the rule first. Beneath: a fine-tune is for rules that are clear and still not kept.

Write the rules once, in full, and send them as they are. The same file your labellers used is the best first prompt, because it is the definition of a right answer. Do not shorten it until you have measured it.

Add a schema. Here it made every reply valid JSON of the right shape, on every model. It will not make the content right, but it takes one whole class of failure off the table. The 0.5B, which had no schema, wrote a wants value outside the list 47 times.

Score every field against gold, separately. A single "accuracy" number would have said 93 of 131 for the best model and hidden that the order number was perfect while wants was wrong on 15 of 138. Score only gold fields, and keep the whole-ticket count too, because that is what your users see.

Read the wrong answers and group them by the rule they break. This is the step that turns a score into a decision. Some groups will be the model's fault, some the rule's. The printer cartridges and the baby monitor were the rule's.

Try the next size up, and test the difference. If a bigger model fixes a group, you may not need training for it; you may need a bigger model, at a cost you can now weigh. Use a sign test on the same messages, and correct for how many tests you ran.

What no size fixed: check the rule first. Here that is 22 fields. At least 6 of them come from a gap in the rules or an unsettled label, and those need a better rule, not a trained model. The rest, mostly wants on messages that ask for nothing, are what a fine-tune has to get right.

When a Prompt Is Enough, and When It Is Not

A prompt is enough for mechanical rules on a big enough model. The order number, copied exactly and never confused with a price or a date, was right on 139 or 140 of 140 on every model, even the smallest prompted one. If your fields are like that, a prompt with a schema may be all you need.

A bigger model is enough for rules it half-knows. Adjectives kept in the item and questions read as requests went from common on the 3B models to rare on qwen2.5:14b. The sign tests covered the whole wants and item fields for qwen2.5:3b against qwen2.5:14b, and both were beyond luck; I did not test the smaller groups inside them, or llama. If your misses look like that, compare the cost of a bigger model with the cost of training a smaller one.

A prompt is not enough for rules that go against the model's habits. A comment that asks for nothing, "other", was wrong 11 times out of 13 even on qwen2.5:14b, and on the chapter's older messages it caused most of the wants misses on every model. More words in the rules may help a little, but the rules already say it plainly. This is the kind of rule is meant for: the model has to learn, from many examples, what your team means by a word.

A prompt is also not enough when the rule itself is incomplete. No model can follow a rule that does not say which words of a product name to keep. The fix there is the rule, or examples that show the annotators' habit, which is one more thing a training set gives you and a prompt cannot.

Remember the running cost. Every call here carried about 648 tokens of prompt for a 25-word message. At a few thousand tickets a day that is a real cost on a hosted model, and a model that has learned the rules would not need to read them every time.

What These Numbers Can and Cannot Tell You

A two-column page headed read before you trust these numbers, titled what this test is, and what it is not. It is: one run per model, temperature 0; 131 fully gold test messages; written and labelled with an AI model; tags on wants and category: my reading. It is not: not repeated with other seeds; not large, one message is under 1%; not real customers; not fixed to the letter, a one-clause edit changed 3 to 19 replies.

One run, at temperature 0. Each model answered each message once. Temperature 0 makes the reply repeatable on the same set-up, and my one check call per model reproduced the stored replies, but a different seed, Ollama version or computer can change close calls. The numbers are one careful measurement, not an average.

Small prompt changes move replies. The history slide showed that changing one clause of the rules changed between 3 and 19 of 140 replies per model, mostly on messages the clause had nothing to do with, while the whole-ticket totals moved by 2 at most. Read any single score here as having at least that much wobble.

131 fully gold messages. One message is less than 1% of the whole-ticket score. Most differences between neighbouring models could not be told apart from luck, and the step from 7B to 14B on the whole ticket was one of them.

Written and labelled with an AI model's help. The 140 test messages were written for this course by an AI model, and the gold labels come from annotators that were runs of the same model family, as lesson 2 explained. Annotators of one model family probably agree with each other more than two people would, so the gold is cleaner than real labels. Real customers write differently, and people would label some fields differently.

The tags are my reading. The wants and category groups come from my reading of each wrong answer, after I saw the replies. Another reader might put a few messages in a different group; the counts per group could move by a message or two, and the per-model totals would not move at all, because those come straight from the grader.

The floor differs in several ways. The 0.5B ran in the mlx library, and the mlx runner used here has no schema option. It was limited to 80 tokens of reply against the others' 120, and it picked the most likely word every time (greedy decoding, the same idea as temperature 0). It is a 4-bit build of the model, compressed to save memory; the Ollama models are 4-bit builds too. And it is much smaller. So its 3 of 131 cannot be put down to size alone.

What to Do Next

A hand-drawn list headed for your own prompt with rules in it, titled five checks before you train anything. Per field?: score each field on its own, never only the whole answer. Which rule?: read every wrong answer and name the rule it breaks. Bigger?: try the next size up on the same messages. Luck?: sign test each comparison; divide 0.05 by how many. What is left?: for what no size kept, check the rule and the labels, then train. Beneath: then write down the cost: the rules ride along on every call.

If you have a prompt that fills in structured fields from text, you can run this lesson's method on it this week. Put your rules in the system message and a schema in the request. Label 50 to 100 real examples twice, as lesson 2 showed, and score every field separately against the gold. Then read every wrong answer, and for each one write down which rule it breaks. A spreadsheet with one row per wrong field and one column for the rule is enough.

Then run the next size up on the same examples and count, for each group, how many it fixes. The groups a bigger model fixes are a pricing question. The groups it does not fix are your training question, once you have checked that the rule and the labels are clear, and the examples in those groups are the first thing to check after any fine-tune.

A closing card headed to keep, titled size did not fix every rule. In large type: 22 fields. Beneath: wrong in all four prompted models, from 3B to 14B; at least 6 are gaps in the rules or labels. Then: fix those first; the rest are a fine-tune's target.

The next lesson is planned to train the 0.5B model from the floor slide on the 500 training tickets from lesson 2, with a short one-line prompt instead of the 401-word rules, and score it on the same 140 messages, field by field, against the four prompted models here.

Knowledge Check

Knowledge Check

4 questions - Score 80% to pass

Q1

Every reply from all four prompted models was valid JSON of the right shape. What made that happen, and what did it not fix?

Q2

On the 13 messages that only thank, complain or report and ask for nothing, the rules say wants is 'other'. How did the models do as they got bigger?

Q3

qwen2.5:7b got 80 whole tickets right and qwen2.5:14b got 93. The sign test gave fixed 23, broke 10, p 0.035, and I made 14 comparisons in all. What is the honest reading?

Q4

The gold item for 'the printer cartridges' is cartridge, but for 'the baby monitor' it is baby monitor. What does this show?