Fine Tuning

What Each Training Setting Does: More Examples, More Passes, Learning Rate and Masking, Tested One at a Time

0 of 19 complete

0%

Contents

Back|Fine TuningWhat Each Training Setting Does: More Examples, More Passes, Learning Rate and Masking, Tested One at a Time
1/19
68 min left
Prerequisites
Training a Small Model on 500 Tickets: What LoRA Changes, and How It Compares With a Promptrequired
Related Topics
Testing a Prompt on New Cases: Why the Score You Tuned On May Be OptimisticPrompting as EngineeringParameter-Efficient Fine-Tuning: LoRA and QLoRALLM and GenAI OpsSmall Rewordings: The Same Instruction, Seven WaysPrompting as EngineeringInstructions Around a Long Document: Where to Put the RulesPrompting as EngineeringPrompts as Code: Templates That a Customer Cannot BreakPrompting as Engineering
1 of 19

Bake the Same Loaf Twice Before You Change the Recipe

Imagine you bake bread at home and you want a better loaf. You could add more yeast, bake it longer, turn the oven up, or use a different flour. The sensible way is to change one thing at a time and see what happens. But there is a step most people skip. Before you change anything, bake the same recipe twice, exactly the same way. The two loaves will not come out identical. One will rise a little higher, one will brown a little more. That difference is not caused by anything you did. It is just how baking goes.

That small, natural difference is the most useful thing to know before you test anything. If a new flour gives you a loaf a little higher than usual, but two loaves from the old recipe already differ by that much, you have learned nothing about the flour. Only a change bigger than the natural difference tells you something.

An illustration of a woman at a wooden table reaching into a large glass jar full of coloured marbles, with a small kitchen timer on the table beside the jar, a window with plants on one side and a bookshelf on the other. Headed the same recipe, three times, titled every run is a fresh handful. Beneath: the same training settings, three seeds: 93, 92 and 93 whole tickets of 131, yet 15 to 21 tickets flipped between any two. Those flips are the wobble a change has to beat.

Training a model works the same way. The last lesson trained one small model once, with one set of choices, and it got 93 of 131 tickets fully right. In this lesson I go back and change those choices, one at a time. How many examples does the model need? How many times should it read them? How big should each learning step be? Should it be scored on the whole example, or only on the answer?

And first, before any of that, I trained the exact same recipe two more times, to see how much the score moves on its own. Every other result in this lesson is judged against that number.

Nine Words for This Lesson

A hand-drawn list headed nine words for this lesson, titled the settings, and how to judge them. Setting: a choice made before training starts, not learned from the data. Seed: a number that fixes the random shuffle of the examples. Wobble: the tickets that flip when only the seed changes. Learning curve: the score plotted against the number of training examples. Pass, or epoch: one read of every training example. Overfitting: getting better on the training examples while getting worse on new ones. Learning rate: how big each small change to the weights is. Masking the prompt: the loss counts only the answer, not the question. Checkpoint: a copy of the add-on saved part way through training. Beneath: one setting changed at a time, against a base run.

A setting is a choice you make before training starts, one the model does not learn from the data. You will also see the word hyperparameter for it. The number of examples, the number of passes and the learning rate are all settings.

A seed is a number that fixes the random choices a training program makes, such as the order in which the examples are shuffled. Two runs with the same seed and the same settings make the same choices. Two runs with different seeds shuffle the examples differently, and end up with slightly different models. The answers that change when only the seed changes are what I call the wobble: the tickets that flip from right to wrong, or back, between two runs that differ in nothing else.

A learning curve is the score plotted against the number of training examples. A pass, also called an epoch, is one read through every training example. Overfitting is when a model keeps getting better on its own training examples while it gets worse on examples it has not seen.

The learning rate is how big each small change to the weights is. Masking the prompt means the loss, the score training tries to lower, counts only the answer part of each example and ignores the question. A checkpoint is a copy of the model's add-on saved part way through training, so you can go back to it later.

One Base Run, Twelve Changes

Everything in this lesson starts from the fine-tune of lesson 4. That base run trained the smallest model, Qwen2.5 0.5B Instruct in its 4-bit version, with : a small add-on trained beside the frozen model, under 1% of its size. It read all 500 training examples, four at a time, for 378 steps, which is just over three passes. The learning rate was 0.00001 and the seed was 1. Every example was the one-line prompt, a customer message and the ticket we want back.

A table headed one base run, and one setting changed at a time, titled whole tickets right, of 131. Base: 500 rows, 378 steps, 3.02 passes, learning rate 0.00001, seed 1: 93. Seed: seed 2: 92; seed 3: 93. Rows: 50, 100, 200, each for 3 passes, so 38, 75 and 150 steps: 26, 58, 78. Passes: one run of 630 steps, a copy after each pass: 82, 88, 93, 88, 91. Learning rate: 0.000001: 58; 0.0001: 96. Mask the prompt: the loss counts only the ticket: 97. Beneath: all Qwen2.5 0.5B, LoRA, batch 4, the one-line prompt. One run per setting.

From that base I made twelve changes. Two runs changed only the seed, to 2 and 3. Three runs used fewer examples: the first 50, 100 and 200 of the 500, each read three times, like the base. One long run read the 500 examples five times, and saved a copy of the add-on after every pass, so each pass could be scored on its own. That gives five results from one run, but the one after three passes turns out to be the base run itself, so it adds four new ones. Two runs used a learning rate ten times smaller and ten times bigger. And one run switched on a setting called --mask-prompt, which a later slide explains.

A flowchart headed how the runs were planned, titled each change starts from the base, alone. A box at the top, base run: 500 rows, 3 passes, rate 0.00001, seed 1, with arrows down to five boxes: seed, 2 or 3; rows, 50, 100, 200; passes, 1 to 5; rate, 10x smaller or bigger; mask the prompt. Beneath: never two changes together, so no run can tell you which combination is best.

Each change was made alone. I never changed two settings together. That keeps each result easy to read, but it has a cost you should keep in mind for the whole lesson: nothing here can tell you which combination of settings is best. A higher learning rate together with masking might be better than either one, or worse. I did not test it.

A sequence diagram with three columns: batch plan, mlx-lm, the grader. Step 1, the batch plan sends mlx-lm one setting changed. Step 2, mlx-lm trains and logs every loss. Step 3, mlx-lm sends back the add-on file. Step 4, the batch plan sends mlx-lm the 140 test messages. Step 5, mlx-lm sends 140 tickets to the grader. Headed what happened for each setting, titled train once, then score the same 140 messages. Beneath: the plan was written before any run started. Greedy decoding, always the most likely next token, at most 80 tokens, one run per setting.

I wrote the plan down before any of these runs started, in the description at the top of the batch script, ticket_batch2.py. It lists every run and every setting.

Each run was then scored exactly as in lessons 3 and 4. It answered the same 140 test messages, never trained on, and the same grader compared each answer with the same gold tickets, the answers the two annotators of lesson 2 agreed on. A field counted only where both annotators agreed, and a whole ticket only on the 131 messages where all five fields are gold. The model always took its most likely next word, which is called greedy decoding, with at most 80 tokens of reply. A token is a small piece of text, often part of a word; a model reads and writes text one token at a time.

All the training and test messages, and their labels, were written with an AI model's help, as lesson 2 explained. They are not real customers' messages. Everything ran on one Mac with mlx; a later slide says what that means for Windows and Linux readers.

Train the Same Thing Three Times

The first question is the one from the bread. How much does the score move when nothing changes except the seed?

Three panels headed the base settings, three different seeds, titled whole tickets right, of 131. Seed 1: 93; category 124, item 123, urgent 134. Seed 2: 92; category 119, item 123, urgent 136. Seed 3: 93; category 120, item 122, urgent 137. Beneath: nothing changed but the shuffle. Category moved from 124 to 119 of 136.

The three runs got 93, 92 and 93 whole tickets right. On totals alone that looks almost perfectly steady: a spread of one ticket. The single fields move more. Category went from 124 with seed 1 to 119 with seed 2, out of 136 scored. Urgent went the other way, from 134 to 136 and 137. The shuffle changed which examples the model saw early and which late, and that was enough to move some fields by five.

Before I trust any of that, one check. Is a run with the same seed really the same run? The long five-pass run used seed 1 and the base settings, so its copy after three passes should be the base run itself. It is: the add-on file saved at step 378 is byte for byte the same file as the base run's add-on, and it wrote the same 140 replies. On this Mac, with these versions of mlx and mlx-lm, the seed fixes the whole run. So everything that differs between the seeds comes from the seed.

A hand-drawn sketch headed sketched: which tickets the three seeds agree on, titled 26 tickets change with the seed alone. A box at the top, 131 fully gold test messages, three seeds, with three arrows down to three boxes: right in all three: 79; right in some: 26; wrong in all three: 26. Below the middle box, another: 15 in two seeds, 11 in one. Beneath: seed 1 to seed 2: 10 tickets fixed and 11 broken, for a total that moved by 1.

Now look message by message, and the steady total turns out to hide a lot of movement. Of the 131 tickets, 79 were right in all three seeds and 26 were wrong in all three. The other 26 were right in some seeds and wrong in others: 15 in two of the three, 11 in only one. From seed 1 to seed 2, 10 tickets went from wrong to right and 11 from right to wrong. The total moved by one, but 21 tickets changed.

This is the most important number in the lesson. There is a group of about 26 messages where this model's answer is close to a coin toss: which way it falls depends on the shuffle. Any change to a setting will also move some of those, just because it changes the training a little.

So in this lesson the wobble is not the one-ticket spread of the totals. It is the tickets that flip: 21 between seeds 1 and 2, 16 between seeds 1 and 3, and 15 between seeds 2 and 3. When that many tickets change at random, the total can easily land a few tickets higher or lower by chance. A new setting that moves the total by three or four, by fixing a few more of those coin tosses than it breaks, has not yet shown anything. The first question is whether the tickets it fixed are the same tickets that move anyway.

A bar chart headed fields whose answer was right in some seeds, wrong in others, titled category moved most with the seed. Five bars, one per field, counting test messages on a scale up to 15: category about 13, wants about 8, order number none, item about 7, urgent about 5. Beneath: category 13, wants 8, order no. 0, item 7, urgent 5. Out of 136 to 140 scored per field.

Field by field, the unstable answers are mostly category: 13 messages where the three seeds did not agree on the team. Wants had 8, item 7, urgent 5, and the order number none at all. Copying digits from the message is settled; deciding which team a borderline message belongs to is where the model is least sure.

The sign test, the message-by-message check from earlier lessons, agrees. It looks only at the messages where two runs disagree, and asks how likely a split that uneven would be if the two runs were equally good. That chance is the p value: a small p means the difference is unlikely to be luck, and p = 1 means the split is as even as it can be. Between any two seeds it gives p = 1. The three seeds are one model trained three times, and they cannot be told apart.

How Many Examples: the Learning Curve

The first real setting is the one you control most directly when you build your own task: how many examples you write.

A line chart headed how many training examples: the learning curve, titled 26, 58, 78, 93: still rising at 500. Whole tickets right against training rows, on a scale up to 131 with a dashed line at all 131. Four points joined by a line: 50 rows about 26, 100 rows about 58, 200 rows about 78, 500 rows about 93. The line climbs steeply at first, then more slowly, and is still going up at 500. Beneath: 50 rows: 26. 100: 58. 200: 78. 500: 93. Each size ran 3 passes, so the small sets also got fewer steps.

With the first 50 of the 500 training examples, the model got 26 whole tickets right. With 100 it got 58, with 200 it got 78, and with all 500 it got 93. This is the biggest effect in the lesson, far bigger than the 15 to 21 tickets that flip between seeds. It is one of only two settings whose effect was clearly beyond luck; the other, a learning rate far too small, comes later and pushed the score down. The curve climbs steeply at first and then more slowly, but it is still rising at 500. I cannot tell from this data where it would stop.

One detail of the design matters here. Each size was read three times, like the base. So the 50-row run took only 38 steps of 4 examples, the 100-row run 75 and the 200-row run 150, against the base's 378. The small sets had fewer examples and fewer steps. This design cannot separate the two. It answers the practical question, "if I write fewer examples and train for three passes, what do I get?", not the narrower one, "what do more examples do at a fixed number of steps?".

A line chart headed the learning curve, field by field, titled order numbers were right from 50 rows; the other four kept rising. Percent right against training rows, from 50 to 100 percent, with five lines: order number flat near 100 percent; urgent rising from about 80 to about 96; wants from about 75 to about 94; category from about 54 to about 91, the steepest climb; item from about 77 to about 82, dipping at 200, then about 90. Beneath: category 74, 111, 121, 124. Wants 104, 117, 121, 130. Order no. 140, 140, 139, 140. Item 105, 112, 109, 123. Urgent 112, 118, 132, 134.

Split by field, the curve tells you which rules need examples. The order number was right on 140 of 140 already with 50 rows: copying digits needs very little. Category went from 74 of 136 to 124, the biggest climb, and most of it came between 50 and 100 rows. Urgent climbed from 112 to 134, most of it between 100 and 200. Wants rose from 104 to 130 in steady steps. Item is the odd one: 105, 112, then 109 with 200 rows, then 123 with 500. That dip at 200 is three fields, well inside the kind of movement the seeds showed, so I would not read anything into it.

An editorial frame headed what the 50-row run had to learn from, titled it wrote the most common answer most often. A zone labelled the first 50 training rows, by category: billing 22, delivery 14, account 7, returns 7. A zone labelled and by wants, the rarest: stop, 1 of 50; in all 500, 60. A zone labelled what the 50-row run wrote, on the test: returns written as billing, 17; account written as billing, 16; stop written as other, 9; a category that is not one of the four, 6. Beneath: category right: 74 of 136; with 500 rows, 124. This run also had only 38 steps.

Reading the 50-row run's wrong answers gives one possible reason. The first 50 training rows are not balanced: 22 of them are billing, and only 7 are account and 7 returns. On the test the model wrote billing for 17 messages whose right answer was returns and for 16 that were account, as if billing were the safe answer. Only one of the first 50 rows has wants "stop", against 60 in all 500, and the 50-row run answered "other" for 9 of the stop requests.

Six times it wrote a category that is not one of the four at all. Four of those were "email": three for customers asking to change their own email address (te010, te116, te132), which is account, and one for "please stop emailing me now" (te119). But this run also trained for only 38 steps, and I did not test a balanced set of 50, so I cannot say how much of this comes from the lopsided rows and how much from too little training.

Still, it is worth checking in your own data. A small set is not only small; it can also be lopsided in ways you did not choose, and a rule that appears in one example is hard for any model to learn. Counting your answers by label costs nothing.

The sign test puts two of the three steps beyond luck: 50 to 100 rows (35 tickets fixed, 3 broken) and 100 to 200 (29 fixed, 9 broken, p 0.0017). The last step, 200 to 500, fixed 22 and broke 7, with p 0.0081. That is below the usual 0.05, but not below the stricter line this lesson uses. Because I ran 11 such tests, some would look good by chance, so a result only counts as beyond luck if p is below 0.05 divided by 11, which is 0.0045. That correction is called Bonferroni. Its biggest gain was on item: 17 item fields fixed, 3 broken. The 200-row run had kept describing words ("oak bookshelf", "scented candles") 6 times and named things that are not products ("van", "payment", "order") 5 times, and the 500-row run got those fields right.

How Many Passes, and What Overfitting Looks Like

The second setting is how many times the model reads the training examples. I trained one run for five passes, 630 steps, and saved a copy of the add-on at the end of each pass, every 126 steps. Then I scored all five copies.

A line chart headed five passes over the 500 rows, one run, titled training loss kept falling; validation loss turned up. Loss against training step from 0 to 650, with dotted vertical lines marked 1 to 5 at the end of each pass. The training line starts above 2 at step 10, drops to about 0.9, then falls in steps at each pass, to about 0.75, 0.6, 0.5 and finally about 0.35. The validation line starts near 1.05, falls slowly to about 0.87 in pass 2, stays near 0.9 through pass 3, then rises to about 0.97 in pass 4 and about 1.1 in pass 5. Beneath: dotted lines: the end of each pass. Validation lowest 0.870 at step 240; 1.089 at step 630. Training 0.32 at the end.

The two loss lines tell a textbook story. The training loss, on the examples the model learns from, kept falling with every pass. It even drops in a visible step each time a new pass starts, because the model is now seeing examples it has already seen, and it remembers them. By step 630 it was down to 0.32. The validation loss, on the 60 examples it never learns from, fell to its lowest, 0.870, at step 240, in the second pass. Then it turned and rose, to 1.089 by the end of pass 5.

That shape is overfitting in the loss. The model keeps getting better at its own 500 examples, including their exact wording, and on the validation examples it becomes more surprised, not less. If you judged by the loss alone, you would stop training during pass 2. Whether the model also got worse at the task is a different question, and only the score can answer it.

A two-column page headed the loss and the score, pass by pass, titled validation loss was lowest in pass 2; the score, at pass 3. Left, the loss: pass 1, step 126: training 1.039, validation 0.883. Pass 2, step 252: training 0.742, validation 0.884. Pass 3, step 378: training 0.616, validation 0.909. Pass 4, step 504: training 0.499, validation 0.968. Pass 5, step 630: training 0.372, validation 1.089. Right, the score: 82, 88, 93, 88 and 91 of 131 whole tickets. Beneath, left: training: the mean of that pass. Validation: the nearest measured step. Beneath, right: no pair of passes compared is beyond luck.

But the score did not follow the loss. After one pass the model got 82 whole tickets right. After two, 88. After three, 93, which is the base run itself. After four, 88, and after five, 91. The validation loss rose by about a fifth from pass 3 to pass 5, while the score went from 93 to 91, a drop the size of the seed wobble.

The sign test agrees with that reading. From pass 3 to pass 5, 7 tickets were fixed and 9 broken: p 0.8036, nothing. From pass 1 to pass 3, 20 fixed and 9 broken: p 0.0614, not beyond luck. And pass 2 against pass 3, 88 against 93, a comparison I added after an independent review asked for it, is 11 fixed and 6 broken: p 0.3323. The fields that moved from pass 3 to pass 5 were mostly item (1 fixed, 6 broken), where the later passes started keeping describing words again: "gaming chair", "leather jacket", "garden bench". That could be a small sign of the model copying the message more closely, but six fields is inside the seed wobble for item, so I note it and do not rely on it.

So on this task, nothing separates the passes: no pair I compared is beyond luck. More passes did little harm to the tickets even while the loss showed overfitting. The next slide explains how the loss and the score can disagree like this.

Why the Loss and the Score Can Disagree

The loss and the score measure different things, and the difference is large enough to see in numbers.

The score counts five fields of the ticket, right or wrong. The loss, with the settings the base run used, counts every token of every training example: the one-line prompt, the customer's message, and the ticket. So the loss is only partly about the ticket. The model is also marked on how well it guesses what the customer is about to type, which has nothing to do with writing a good ticket.

An isometric drawing headed what the loss counts, over the 378 steps, titled 127,843 tokens, or only 53,776. On the left, a tall stack of two blocks drawn to scale: a lower block labelled ticket: 53,776 and a taller upper block labelled prompt and message: 74,067. On the right, a single block the same height as the left lower block, labelled ticket: 53,776. Under them: without the mask: 127,843; with --mask-prompt: 53,776. Beneath: heights to scale: the prompt and message block is 1.38 times the ticket block. The ticket is 42.1% of what the unmasked loss counts.

I can measure how much of the loss is about the customer's words. mlx-lm prints how many tokens its loss has counted, and I read its code to be sure that is exactly the number of tokens the loss is averaged over. In the base run, over 378 steps, the loss counted 127,843 tokens. In the masked run, which read exactly the same examples in exactly the same order but counted only the tickets, it counted 53,776. So the tickets were 42.1% of what the base run's loss was measuring. The other 74,067 tokens, well over half, were the prompt, the customer's message and the few tokens of chat formatting.

Now the passes result makes sense. One possible reason the validation loss rose while the score held is that, in later passes, the model got better at repeating the wording of the 500 training messages and so more surprised by the wording of new ones. That raises the loss on more than half of what it counts. The tickets, which are short and follow a strict pattern, can stay just as right. This is a likely explanation, not a proven one: I did not run five passes with the prompt masked, which would have shown the ticket-only loss directly.

The practical lesson is the same one lesson 4 ended on, now with a number behind it. A validation loss is a useful warning that training is going somewhere odd. It is not a measure of the task. Before you stop, start or change anything because of the loss, score the task.

Learning Rate: How Big Each Step Is

The learning rate decides how far the add-on moves at each step. The base run used 0.00001. I tried ten times smaller, 0.000001, and ten times bigger, 0.0001, with everything else the same.

A line chart headed validation loss at three learning rates, titled too small is slow; too big turns up early. Validation loss against training step from 0 to 400, on a scale from 0.8 to 2.4. The rate 0.000001 line, dashed, starts high at about 2.3 at step 20 and falls slowly, ending just under 1.0. The rate 0.00001 line starts at about 1.05 and settles just under 0.9. The rate 0.0001 line starts at about 1.05, dips to about 0.93, then rises after step 250 to about 1.05 to 1.1. Beneath: step 20: 2.316, 1.052, 1.055. Lowest: 0.983 at step 378, 0.870 at step 240, 0.925 at step 120. Step 378: 0.983, 0.904, 1.037.

The loss curves show the usual picture. With the small rate, 0.000001, learning is slow: after 20 steps the validation loss was still 2.316, where the base had already reached 1.052. It kept falling for the whole run, reaching its lowest, 0.983, at the very last step. With the big rate, 0.0001, the first steps looked like the base, but the validation loss reached its lowest, 0.925, early, at step 120, and then rose to 1.037 by the end, while its training loss fell to 0.354, much lower than the base's 0.587. Big steps learn the training examples faster, and start to overfit in the loss sooner.

A table headed learning rate: the loss against the score, titled neither loss showed the gap in the score. Rate 0.000001: validation end 0.983, training end 0.966; whole tickets 58; item 101, wants 113. Rate 0.00001: validation end 0.904, training end 0.587; whole tickets 93; item 123, wants 130. Rate 0.0001: validation end 1.037, training end 0.354; whole tickets 96; item 119, wants 131. Beneath: rate 0.000001 against 0.0001: validation at the end 0.983 against 1.037; at the lowest 0.983 against 0.925. Whole tickets 58 against 96.

The scores again tell a different story from the loss. The small rate got 58 whole tickets right, far below the base's 93, and the sign test puts that beyond luck (43 tickets fixed by the base rate, 8 broken). Its loss had come down a long way, from 4.241 to 0.983, but not far enough to learn the house rules.

I read the item fields the base rate fixed: 7 were describing words the slow run had kept ("office chair", "oak bookshelf") and 7 were longer product names kept whole where the rule keeps only the last word ("camera tripod", "garden bench"). The slow run had learned the shape of a ticket and was mostly copying words from the message. It had not yet learned the rule that shortens them.

The big rate got 96, three more than the base. That sounds like a win, and it is the kind of number that ends up in a blog post. But its final validation loss was the worst of the three, and the sign test says the 96 is not beyond luck: 20 tickets fixed and 17 broken, p 0.7428. It fixed and broke about as many wants and category answers (6 fixed and 5 broken for each), and broke more item fields than it fixed (5 fixed, 9 broken). One of the broken ones is not even a word: for a message about "the pleated skirt" it wrote the item "skel".

Put the two together and the table carries a warning of its own. At the end of training, the small rate had the lower validation loss, 0.983 against 1.037, and scored 58 against 96. At their lowest points the order flips: 0.925 for the big rate against 0.983, and the lowest point is what you would use to pick a checkpoint. So the loss did not clearly prefer the slow rate. But neither loss number shows anything like the 38-ticket gap in the score. Again: judge by the task, not the loss.

Masking the Prompt: Score Only the Answer

The last setting changes what the loss counts. With --mask-prompt, mlx-lm still gives the model the whole example to read, the prompt, the message and the ticket, but it computes the loss only on the ticket. The model is no longer marked on guessing the customer's words.

An editorial frame headed one training row, with --mask-prompt, titled the loss is taken on the ticket only. A zone labelled system, not counted in the loss: Turn the customer message into a support ticket as one line of JSON. A zone labelled user, not counted in the loss: Please refund order no. 812204 before Friday. The sofa was cancelled and I have rent due while that £240 deposit sits with you. A zone labelled assistant, counted in the loss, holding five boxes: category billing; wants refund; order_number 812204; item sofa; urgent true. Beneath: the model still reads the prompt and the message; it is just not scored on guessing them. 500 rows, the same as the base.

The masked run got 97 whole tickets right, the highest in the lesson. Its loss also looks spectacular: the validation loss started at 2.833 and ended at 0.020, where the base ended at 0.904. Those two numbers cannot be compared. A loss is an average over the tokens it counts, and the two runs count different tokens. The masked loss averages over the tickets only, which are short and very predictable: the same five keys in the same order, a few allowed words. The unmasked loss averages over the customers' messages too, which nobody can predict well. Even its starting value differs, 2.833 against 4.241, before a single step of training, only because it is averaging over different text. The masked loss tells you the tickets are being learned; it does not tell you the masked model is 45 times better.

A two-column page headed --mask-prompt against the base, whole tickets, titled 16 fixed, 12 broken. Left, fixed by the mask: te015, Unsubscribe me from the newsletter and delete my account while you ...; urgent: base true, mask false, right false. te138, Please move my delivery slot to Thursday afternoon, nobody is home on ...; urgent: base true, mask false, right false. te095, Order 99014 for the new projector: the invoice shows my old company ...; wants: base change, mask other, right other. te074, My package arrive open and the wireless headphones is missing. I want ...; category: base returns, mask delivery, right delivery. Right, broken by the mask: te081, want refund for the 2 chargers order 70331 box not even opened; category: base returns, mask billing, right returns. te139, I need the refund for order no. 66390 before Friday, my card bill is ...; category: base returns, mask billing, right returns. te114, Hi, please pass a thank you to the driver on order 12876, the big ...; category: base delivery, mask billing, right delivery. te043, Delete my account and refund my remaining gift balance of 15.40.; wants: base other, mask refund, right other. Beneath, left: all 8 shown are among the 26 seed-unstable tickets. Beneath, right: on the tickets all three seeds agreed on: 6 fixed, 6 broken. p 0.5716.

So is 97 really better than 93? Message by message, the masked run fixed 16 tickets and broke 12. On urgent it looks better at first. The base run raised five false alarms, marking messages urgent that had no deadline, and the masked run removed four of them and added none. But look at which four. Each of them was already right in at least one of seed 2 and seed 3, so they are tickets that flip with the shuffle alone.

The one false alarm that all three seeds raised, te119, "please stop emailing me now", the masked run raised too. Elsewhere the changes go both ways. It fixed a customer who wrote that no change was needed, where the right wants is "other" (te095), and a parcel that arrived open with the headphones missing, which belongs to delivery (te074). It broke two refunds for items sent back, writing billing where the rule says returns (te081, te139), and it sent the thank-you to the delivery driver (te114) to billing. All eight of those examples are among the 26 seed-unstable tickets. And once it wrote a strange item: for a "ceramic planter" it answered "plant" followed by a Chinese character. The JSON was valid; the word was not.

Now split the 28 tickets it moved. On the 26 seed-unstable tickets, the mask fixed 10 and broke 6. On the other 105, the tickets all three seeds agreed on, it fixed 6 and broke 6. So its whole gain of four came from tickets that also move when only the seed changes.

The sign test gives p 0.5716 for 16 against 12. So masking scored a little higher here, on one run, and I cannot say it is better. It is not evidence of harm either: it broke 12 and fixed 16, a split as even as the seeds'. What masking does give you for certain is a loss that measures only the answer, which is worth having by itself.

Which Changes Are More Than Luck

With 12 changes against one base, some differences will look big by chance. So I checked them message by message.

A table headed message by message: 11 sign tests, chosen after the totals, titled beyond luck needs p below 0.0045. Seed 1 to seed 2: 93 to 92; fixed 10, broke 11; p 1.0000. Seed 1 to seed 3: 93 to 93; fixed 8, broke 8; p 1.0000. Seed 2 to seed 3: 92 to 93; fixed 8, broke 7; p 1.0000. 50 rows to 100 rows: 26 to 58; fixed 35, broke 3; p below 0.0001, beyond luck. 100 rows to 200 rows: 58 to 78; fixed 29, broke 9; p 0.0017, beyond luck. 200 rows to base: 78 to 93; fixed 22, broke 7; p 0.0081. Pass 1 to pass 3: 82 to 93; fixed 20, broke 9; p 0.0614. Pass 3 to pass 5: 93 to 91; fixed 7, broke 9; p 0.8036. Rate 0.000001 to base: 58 to 93; fixed 43, broke 8; p below 0.0001, beyond luck. Base to rate 0.0001: 93 to 96; fixed 20, broke 17; p 0.7428. Base to mask: 93 to 97; fixed 16, broke 12; p 0.5716. Beneath: 3 of 11 are beyond the line; 0.05 / 11 = 0.0045.

The sign test looks only at the messages where two runs disagree, and asks how likely a split that uneven would be if the two runs were equally good. I ran 11 tests: the three pairs of seeds, and each neighbouring step of each setting. With 11 tests, some will look good by chance, so I use the Bonferroni correction from earlier lessons: a result counts as beyond luck only if p is below 0.05 divided by 11, which is 0.0045.

I have to be honest about when I chose them. The plan for which runs to make was written before any run. But I chose these 11 comparisons after I had seen the whole-ticket totals, though before I counted any of them message by message. Choosing tests after seeing totals can make you test the gaps that look largest, so treat the three results below the line as strong and everything else as unknown.

Three comparisons are beyond the line: 50 to 100 rows, 100 to 200 rows, and the tiny learning rate against the base. Every one of them is about a model that had not learned enough yet. None of the changes to a model that was already trained well enough, more passes, a bigger rate or masking, is beyond luck.

Three panels headed tickets a change moved, and how many were among the 26 the seeds disagree on, titled many tickets that moved also move with the seed alone. --mask-prompt: 28; moved: 16 fixed, 12 broken; 16 of them seed-unstable. Rate 0.0001: 37; moved: 20 fixed, 17 broken; 15 of them seed-unstable. Pass 3 to 5: 16; moved: 7 fixed, 9 broken; 11 of them seed-unstable. Beneath: seed-unstable: right in one or two of the three seeds, not all three.

There is a second check, and it uses the wobble from the start of the lesson. Remember the 26 tickets that were right in some seeds and wrong in others. Of the 28 tickets the masked run moved, 16 were among those 26. Of the 37 the big learning rate moved, 15 were. Of the 16 that changed from pass 3 to pass 5, 11 were. A large share of what these changes "did" happened on tickets that change anyway when you only reshuffle the examples. Of the 16 tickets masking fixed, 10 were such coin tosses; on the tickets all three seeds agreed on, it fixed 6 and broke 6. That does not prove masking did nothing. It does mean its whole four-ticket gain sits inside the wobble.

A two-column table headed answers no rule allows, titled what this 0.5B model wrote in place of a real answer. Seed 2: te009, item: trident for tripod. Seed 2: te011, item: puffets for cushion. Rate 0.0001: te018, item: skel for skirt. 50 rows: te010, category: email, not one of the four; the customer wanted to change their own email address. Beneath: all 140 replies of every run were valid JSON; the words inside were not always real.

Reading the wrong answers also turned up something no total shows. Across these runs, the small model sometimes wrote a word that is not in the message, or not a word at all. Seed 2 wrote "trident" for a camera tripod and "puffets" for velvet cushions. The big learning rate wrote "skel" for a skirt. The 50-row run put "email" in the category field, where only the four team names are allowed. Every reply of every run was valid JSON with the five keys in order, so a check on the format alone would pass all of them. When you serve a small fine-tune, check each value against what the message says and against the list of allowed words.

Try It Yourself

This script has two parts. First it prints the exact mlx-lm command line the lab used for each setting in this lesson: the base, the two seeds, the three sizes, the five-pass run, the two learning rates and the mask. You can read them side by side and see that each one differs from the base in one flag only. Then it runs one of them for real, made small: the --mask-prompt run, on the first 60 of the 500 training examples and the first 8 of the 60 validation examples, for 30 steps instead of 378.

A real screenshot of VS Code with settings_demo.py open, showing the top of the file: the docstring, the imports, the model name, the base settings and the start of the dictionary of runs. The rest of the file is in the code box on this slide. Beneath: copy it from the box on the slide.

Before you run this lab. This one does not use Ollama; it uses Python with mlx-lm (pip install mlx-lm), and the first run downloads the 0.5B base model, about 280 MB. If you have not set up Python for the labs yet, the lab setup guide shows how, on macOS, Windows or Linux. I do not give a training time, because another lab was using the same GPU while mine ran.

If you use Windows or Linux: mlx runs only on Macs with Apple silicon, and this lab used mlx on a Mac, which is the only setup tested here. The same kind of training is commonly done with Hugging Face's library on a GPU, for example a free Google Colab one; I have not run it for this chapter. Your numbers will differ, and nothing in this lesson claims the Mac's numbers hold on other hardware.

"""Every training setting of lesson 5 as an exact mlx-lm command, then one of them run for real, made small.

Lesson 5 of the fine-tuning chapter changed one training setting at a time on a 0.5B model: the seed, the number
of training rows, the number of passes, the learning rate, and whether the loss counts the prompt. This script
prints the exact mlx-lm command line the lab used for each of those runs. Then it runs ONE of them for real,
the one with --mask-prompt, on a small slice: 60 training rows and 30 steps instead of 500 rows and 378 steps,
and prints only the loss lines. It is not the lab's run, and its numbers will not match the lab's.

It needs a Mac with Apple silicon (mlx runs only there) and Python 3 with mlx-lm:
    pip install mlx-lm
    python settings_demo.py
The first run downloads the base model (about 280 MB). The training rows are the first 60 of the lesson's 500
and the validation rows the first 8 of its 60, copied here so the script needs no other file.
"""
import json
import re
import subprocess
import sys
from pathlib import Path

MODEL = "mlx-community/Qwen2.5-0.5B-Instruct-4bit"   # Qwen2.5 0.5B Instruct, its weights stored in 4 bits
PROMPT = "Turn the customer message into a support ticket as one line of JSON."
DEMO_ITERS = 30

# The base run, then each change. Every run changes ONE thing; everything else stays as in the base.
BASE = {"data": "data/ticket2", "iters": 378, "lr": "1e-05", "seed": 1, "extra": []}
RUNS = {
    "base": {},                                                    # 500 rows, 378 steps = 3 passes
    "seed2": {"seed": 2},                                          # the same run, another shuffle
    "seed3": {"seed": 3},
    "n50": {"data": "data/ticket2_n50", "iters": 38},              # the first 50 rows, 3 passes
    "n100": {"data": "data/ticket2_n100", "iters": 75},
    "n200": {"data": "data/ticket2_n200", "iters": 150},
    "epochs5": {"iters": 630, "extra": ["--save-every", "126"]},   # 5 passes, a copy after each one
    "lr1e-06": {"lr": "1e-06"},                                    # nudges 10 times smaller
    "lr0.0001": {"lr": "0.0001"},                                  # nudges 10 times bigger
    "mask": {"extra": ["--mask-prompt"]},                          # the loss counts only the ticket
}


def command(name, data=None, iters=None, eval_every="20", adapter=None):
    """The mlx-lm LoRA command for one run, exactly as the lab ran it (paths aside)."""
    r = {**BASE, **RUNS[name]}
    return ["-m", "mlx_lm", "lora",
            "--model", MODEL,
            "--train",
            "--data", data or r["data"],                 # a folder holding train.jsonl and valid.jsonl
            "--iters", str(iters or r["iters"]),         # steps; each step reads 4 rows
            "--batch-size", "4",
            "--learning-rate", r["lr"],                  # how big each nudge to the add-on is
            "--adapter-path", adapter or f"adapters/{name}",
            "--steps-per-report", "10",                  # print the training loss every 10 steps
            "--steps-per-eval", eval_every,              # print the validation loss this often
            "--seed", str(r["seed"]),                    # which shuffle of the rows
            *r["extra"]]


# (customer message, the ticket we want), the first 60 training rows
TRAIN = [
    ('Please refund order no. 812204 before Friday. The sofa was cancelled and I have rent due while that £240 deposit sits with you.',
     '{"category": "billing", "wants": "refund", "order_number": "812204", "item": "sofa", "urgent": true}'),
    ('Do gift cards expire? I have one from 2024 with £40 left.',
     '{"category": "billing", "wants": "information", "order_number": null, "item": "gift card", "urgent": false}'),
    ('Please apply my 10% loyalty discount to my next order.',
     '{"category": "billing", "wants": "other", "order_number": null, "item": null, "urgent": false}'),
    ("Cancel the wine club payment immediately please, I'm moving abroad and won't be here to receive it.",
     '{"category": "billing", "wants": "cancel", "order_number": null, "item": null, "urgent": true}'),
    ('I need a copy of the receipt for the kettle on order 41190 for my warranty claim, which expires in 2 years.',
     '{"category": "billing", "wants": "other", "order_number": "41190", "item": "kettle", "urgent": false}'),
    ('Someone has used my card on your site to buy perfume. Refund it right now and block the order.',
     '{"category": "billing", "wants": "refund", "order_number": null, "item": "perfume", "urgent": true}'),
    ("Please add our cost centre code to teh invoices from now on, and reissue last month's.",
     '{"category": "billing", "wants": "change", "order_number": null, "item": null, "urgent": false}'),
    ('Are my saved card details shared with anyone?',
     '{"category": "account", "wants": "information", "order_number": null, "item": null, "urgent": false}'),
    ('order no. 41876 was a duplicate, i clicked pay twice for the £120 bookcase when the page froze. please refund one of them',
     '{"category": "billing", "wants": "refund", "order_number": "41876", "item": "bookcase", "urgent": false}'),
    ("Hello, does next day delivery include Sundays for furniture? I'm in postcode LS6 3HN.",
     '{"category": "delivery", "wants": "information", "order_number": null, "item": "furniture", "urgent": false}'),
    ('Can you pleas tell me urgently why my card was declined for the laptop? I have tried three times.',
     '{"category": "billing", "wants": "information", "order_number": null, "item": "laptop", "urgent": true}'),
    ('Stop calling me at work. It must stop today, my manager has noticed.',
     '{"category": "account", "wants": "stop", "order_number": null, "item": null, "urgent": true}'),
    ('charged £189 for the printer and £189 AGAIN for the ink bundle that was meant to be £19. order 57018. fix it and refund me',
     '{"category": "billing", "wants": "refund", "order_number": "57018", "item": "ink", "urgent": false}'),
    ('you chraged me for a 4-slice toaster on 66018 that never shipped, I want my money back asap',
     '{"category": "delivery", "wants": "refund", "order_number": "66018", "item": "toaster", "urgent": true}'),
    ("we're away the week the hot tub is due. please rebook it for the week after",
     '{"category": "delivery", "wants": "change", "order_number": null, "item": "hot tub", "urgent": false}'),
    ('How do I change my password? Or can you jsut reset it for me?',
     '{"category": "account", "wants": "information", "order_number": null, "item": null, "urgent": false}'),
    ("Please delete my child's account, they are under 13.",
     '{"category": "account", "wants": "other", "order_number": null, "item": null, "urgent": false}'),
    ("Hi, what day will the sofa on order 50378 be delivered? I haven't had a date yet.",
     '{"category": "delivery", "wants": "information", "order_number": "50378", "item": "sofa", "urgent": false}'),
    ('The currency converter on your site is wrong, it showed 30 euros for a scarf and charged far more.',
     '{"category": "billing", "wants": "other", "order_number": null, "item": "scarf", "urgent": false}'),
    ("The tracking number you gave me for teh speaker on 40502 doesn't work on the courier's site. Is it the right one?",
     '{"category": "delivery", "wants": "information", "order_number": "40502", "item": "speaker", "urgent": false}'),
    ('Please switch me from monthly to annual billing and tell me what the new price is.',
     '{"category": "billing", "wants": "change", "order_number": null, "item": null, "urgent": false}'),
    ('Is there any delay to deliveries because of the strikes? I need my order by Friday.',
     '{"category": "delivery", "wants": "information", "order_number": null, "item": null, "urgent": true}'),
    ('I need a proper VAT receipt for the dishwasher on order 52480 today for a grant claim, the one you sent has no VAT number.',
     '{"category": "billing", "wants": "other", "order_number": "52480", "item": "dishwasher", "urgent": true}'),
    ('The desk lamp on #20931 was advertised with free delivery but I paid £4.95 postage at checkout. Can I get the postage refunded?',
     '{"category": "billing", "wants": "refund", "order_number": "20931", "item": "lamp", "urgent": false}'),
    ('Sofa delivery estimate is now January. Cancel my order and refund me.',
     '{"category": "delivery", "wants": "cancel", "order_number": null, "item": "sofa", "urgent": false}'),
    ('hoodie from 25003 came with the print peeling off. exchange for a new one please',
     '{"category": "returns", "wants": "replacement", "order_number": "25003", "item": "hoodie", "urgent": false}'),
    ('Thanks for making the return so easy, the drop-off took two minutes.',
     '{"category": "returns", "wants": "other", "order_number": null, "item": null, "urgent": false}'),
    ('can i change the delivery slot for my fridge freezer on 29706 to the evening one',
     '{"category": "delivery", "wants": "change", "order_number": "29706", "item": "fridge freezer", "urgent": false}'),
    ('Do you do returns collection for the armchair on order 39402 or do I have to take it to the post office?',
     '{"category": "returns", "wants": "information", "order_number": "39402", "item": "armchair", "urgent": false}'),
    ('Please unlink my social media login from my account.',
     '{"category": "account", "wants": "other", "order_number": null, "item": null, "urgent": false}'),
    ('Two of the four chairs have wobbly legs. Could you send two new ones and collect the bad ones?',
     '{"category": "returns", "wants": "replacement", "order_number": null, "item": "chair", "urgent": false}'),
    ('what does the charge "SVC FEE" on my receipt for the lawnmower, order 83110, mean',
     '{"category": "billing", "wants": "information", "order_number": "83110", "item": "lawnmower", "urgent": false}'),
    ('#84102 teh mirror is cracked right across. replacement please',
     '{"category": "delivery", "wants": "replacement", "order_number": "84102", "item": "mirror", "urgent": false}'),
    ('The shorts on oder 51408 are too big. Swap for a medium please.',
     '{"category": "returns", "wants": "replacement", "order_number": "51408", "item": "shorts", "urgent": false}'),
    ("The house number on my delivery should be 27 not 72. I've told you twice already. Fix it.",
     '{"category": "delivery", "wants": "change", "order_number": null, "item": null, "urgent": false}'),
    ('order 44720, cancel the telescope before the payment goes through at midnight please',
     '{"category": "billing", "wants": "cancel", "order_number": "44720", "item": "telescope", "urgent": true}'),
    ('Order 18026: is the delivery charge refunded too if I return the chairs?',
     '{"category": "returns", "wants": "information", "order_number": "18026", "item": "chair", "urgent": false}'),
    ('i paid for express shipping on my trainers (order 53308) and they came in 6 days. id like the express fee back',
     '{"category": "billing", "wants": "refund", "order_number": "53308", "item": "trainers", "urgent": false}'),
    ('The artificial tree delivery on #47033 is set for the 2nd. Could you push it back to the 9th?',
     '{"category": "delivery", "wants": "change", "order_number": "47033", "item": "tree", "urgent": false}'),
    ("Where do I find the returns form for the toaster on order 68830? It wasn't in the box.",
     '{"category": "returns", "wants": "information", "order_number": "68830", "item": "toaster", "urgent": false}'),
    ('hi just checking whether you do student discount on laptops, thanks',
     '{"category": "billing", "wants": "information", "order_number": null, "item": "laptop", "urgent": false}'),
    ('Nobody can be home on the booked day for the washing machine delivery. Please change it today.',
     '{"category": "delivery", "wants": "change", "order_number": null, "item": "washing machine", "urgent": true}'),
    ("why is there a pending charge of $1 on my card? i didn't buy anything. this is really dodgy",
     '{"category": "billing", "wants": "information", "order_number": null, "item": null, "urgent": false}'),
    ('order 64419 has been "processing" for 16 days. absolute joke. cancel it',
     '{"category": "delivery", "wants": "cancel", "order_number": "64419", "item": null, "urgent": false}'),
    ('Please cancel the magazine subscription before it renews tonight. I only wanted one issue.',
     '{"category": "billing", "wants": "cancel", "order_number": null, "item": "magazine", "urgent": true}'),
    ("cancel the garden bench before it ships today, I found out I'll be away for a month",
     '{"category": "billing", "wants": "cancel", "order_number": null, "item": "bench", "urgent": true}'),
    ("Please change the address for the cooker on my order right now. It's going to a flat I moved out of last week.",
     '{"category": "delivery", "wants": "change", "order_number": null, "item": "cooker", "urgent": true}'),
    ("Someone opened an account with my email address. It wasn't me. Please shut it down today and tell me what they ordered.",
     '{"category": "account", "wants": "other", "order_number": null, "item": null, "urgent": true}'),
    ('Hello, please cancel the recurring payemnt for the protein powder on order 27781. It was meant to be a one-off purchase.',
     '{"category": "billing", "wants": "cancel", "order_number": "27781", "item": "protein powder", "urgent": false}'),
    ('Do I need an account to place an order, or can I check out as a guest?',
     '{"category": "account", "wants": "information", "order_number": null, "item": null, "urgent": false}'),
    ('Good morning, please amend my profile so my first name reads Sam, not Samuel.',
     '{"category": "account", "wants": "change", "order_number": null, "item": null, "urgent": false}'),
    ("Can I return trainers I've worn outside once? They rub my heel.",
     '{"category": "returns", "wants": "information", "order_number": null, "item": "trainers", "urgent": false}'),
    ('Can you tell me which courier is bringing my trainers? Order 14408.',
     '{"category": "delivery", "wants": "information", "order_number": "14408", "item": "trainers", "urgent": false}'),
    ('Hi, this is about ref 24068. You agreed on the phone to refund the delivery charge on the wardrobe. Please send the refund you promised and confirm by email.',
     '{"category": "billing", "wants": "refund", "order_number": "24068", "item": "wardrobe", "urgent": false}'),
    ('Kindly correct the title on my account from Mr to Dr.',
     '{"category": "account", "wants": "change", "order_number": null, "item": null, "urgent": false}'),
    ('How long does it take for a new email address to show on my account? The receipt for the kettle on order 70211 still went to the old one.',
     '{"category": "account", "wants": "information", "order_number": "70211", "item": "kettle", "urgent": false}'),
    ('My account got locked after one wrong password. That seems excessive and frankly stupid.',
     '{"category": "account", "wants": "other", "order_number": null, "item": null, "urgent": false}'),
    ('i lost the return label for the rug on 44301, can i still send it back? the 30 days run out tomorrow',
     '{"category": "returns", "wants": "information", "order_number": "44301", "item": "rug", "urgent": true}'),
    ('Stop the "back in stock" emails for the laptop, I no longer need it.',
     '{"category": "account", "wants": "stop", "order_number": null, "item": "laptop", "urgent": false}'),
    ('This is a data protection request for my records.',
     '{"category": "account", "wants": "other", "order_number": null, "item": null, "urgent": false}'),
]

# the first 8 validation rows: never trained on, only used to report the loss
VALID = [
    ('The renewal is due in two days and I want it cancelled before then. Please act urgently.',
     '{"category": "billing", "wants": "cancel", "order_number": null, "item": null, "urgent": true}'),
    ("I bought the glass kettle in your sale and it rang up at the full price. I'd like the sale difference back, and please check your pricing.",
     '{"category": "billing", "wants": "refund", "order_number": null, "item": "kettle", "urgent": false}'),
    ('The tent on 15590 was charged once at checkout and again when it shipped. I would like the second charge refunded.',
     '{"category": "billing", "wants": "refund", "order_number": "15590", "item": "tent", "urgent": false}'),
    ('hello, i cancelled the jumper on #19934 within the hour and the paymnet still went through. refund when you can, thanks',
     '{"category": "billing", "wants": "refund", "order_number": "19934", "item": "jumper", "urgent": false}'),
    ('cancel my premium. never used it once',
     '{"category": "billing", "wants": "cancel", "order_number": null, "item": null, "urgent": false}'),
    ('The treadmill on order no. 56679 has to be stopped before teh van leaves this morning. Cancel it asap.',
     '{"category": "billing", "wants": "cancel", "order_number": "56679", "item": "treadmill", "urgent": true}'),
    ('Kindly cancel my pre-order payment for the game, order 16780. The release date slipped and I no longer need it.',
     '{"category": "billing", "wants": "cancel", "order_number": "16780", "item": "game", "urgent": false}'),
    ('Honestly this is ridiculous. Three emails and still no refund for the air purifier on order 30771. Just put the money back on my card.',
     '{"category": "billing", "wants": "refund", "order_number": "30771", "item": "air purifier", "urgent": false}'),
]


def write(path, rows):
    """One chat per line: the one-line prompt, the customer message, and the ticket as the answer."""
    with open(path, "w") as f:
        for message, ticket in rows:
            f.write(json.dumps({"messages": [
                {"role": "system", "content": PROMPT},
                {"role": "user", "content": message},
                {"role": "assistant", "content": ticket}]}) + "\n")


print("The lab's runs, one setting changed at a time:")
for name in RUNS:
    print(f"{name}: python " + " ".join(command(name)))

data = Path("st_demo_data")
data.mkdir(exist_ok=True)
write(data / "train.jsonl", TRAIN)
write(data / "valid.jsonl", VALID)

print(f"\nNow the mask run for real, made small: {len(TRAIN)} rows, {DEMO_ITERS} steps.")
cmd = command("mask", data=str(data), iters=DEMO_ITERS, eval_every="10", adapter="st_demo_adapter")
out = subprocess.run([sys.executable, *cmd], capture_output=True, text=True)
if out.returncode:
    sys.exit(out.stderr[-2000:])
# mlx-lm also prints speeds; only the loss lines are kept here
for step, kind, loss in re.findall(r"Iter (\d+): (Train|Val) loss ([\d.]+)", out.stdout):
    print(f"Iter {step}: {kind} loss {loss}")
counted = re.findall(r"Trained Tokens (\d+)", out.stdout)
print(f"Trained Tokens {counted[-1]}  (the tokens the loss counted: with --mask-prompt, only the tickets')")

This is a real run in VS Code's terminal.

A real screenshot of VS Code's terminal after running python settings_demo.py. First the lab's ten command lines, one per run, from base to mask, each a python -m mlx_lm lora command with its data folder, steps, batch size 4, learning rate, adapter folder and seed; the mask line ends with --mask-prompt. Then the line: now the mask run for real, made small: 60 rows, 30 steps. Then the loss lines: validation loss 2.650 at step 1, 0.119 at step 10, 0.047 at step 20 and 0.021 at step 30; training loss 0.964, 0.116 and 0.052 at steps 10, 20 and 30. Last, trained tokens 4250, the tokens the loss counted.

When I ran it, the validation loss went from 2.650 before training to 0.119 at step 10, 0.047 at step 20 and 0.021 at step 30. The training loss went 0.964, 0.116, 0.052. The loss counted 4,250 tokens, only the tickets' tokens, as the mask intends.

This 30-step demo is not the lab's 378-step run, and it is not meant to match its numbers. With the mask on, the loss falls very fast even on 60 examples, because it only has to learn the short, regular tickets; that is exactly why a masked loss cannot be compared with an unmasked one, as the mask slide explained. The lab's demo mode checked that the script's 60 training rows and 8 validation rows are exactly the first rows of the lab's files, and that each of the 10 printed commands has the same steps, learning rate, seed, batch size and flags as the stored lab run. The output above is from my run; your numbers can differ a little on a different Mac or version of mlx.

The Lab Report

A real terminal recording of python settings_report.py, in eight numbered sections. 1, every run with its rows, steps, passes, learning rate, seed and loss, and the check that the five-pass run repeats the base run and that its pass-3 file is byte for byte the base add-on. 2, the three seeds field by field, 93, 92 and 93 whole tickets, with 79 right in all three, 26 in none and 26 in some. 3, the four sizes, 26, 58, 78 and 93, and what the first 50 rows contain. 4, the five passes, 82, 88, 93, 88 and 91. 5, the three learning rates, 58, 93 and 96. 6, masking, 127843 tokens against 53776, 42.1%, and 97, and the base run's five urgent false alarms, with whether each was right in seed 2, seed 3 and the masked run. 7, the 11 sign tests, three beyond luck, with how many moved tickets were seed-unstable, and one comparison added after review, pass 2 to pass 3, 88 to 93, p 0.3323. 8, what changed field by field, and the item fields fixed by kind. Beneath: the lab's own report. It calls no model.

The report lives in scripts/labs/finetune/settings_report.py. It reads only stored files: each run's training record (the exact command, every loss line and the end of the training output), each add-on's settings file, the stored replies of all 14 scored add-ons on the 140 test messages, the test messages and the gold file. A json mode writes the same numbers to results/st-report.json, which is what the figures read. It calls no model.

Before it prints anything, it checks the files against each other and stops if one disagrees. It grades every stored reply again with the lab's own grader against the current gold, and stops if any stored mark differs. It checks each run's settings against its add-on's settings file, and checks that each run changed only the one setting it was meant to. It checks that the 50, 100 and 200-row data folders hold exactly the first rows of the 500, with the same 60 validation rows. And it checks that the five-pass run repeats the base run: the loss lines they share are equal, and the add-on saved after pass 3 is byte for byte the base run's add-on.

What came after I saw the data: the choice of the 11 sign tests came after I had seen the whole-ticket totals. The "seed-unstable" count, the sorting of fixed item fields by kind, and every example in the figures were chosen after I read the replies. After an independent review, I added one comparison outside the 11, pass 2 against pass 3, and the check of which urgent false alarms were stable across the seeds. The runs, their settings and the data were fixed before any of them started, in the batch plan.

Four brand cards headed the tools, with their logos, titled what ran where. Apple mlx: 10 training runs, 14 add-ons scored. Hugging Face: the 0.5B base model, downloaded once. Python: the batch plan, the report, the demo, the box. asciinema: recorded the lab report in a real terminal.

The lab ran 10 training runs, and scored 14 add-ons: the five-pass run gave five of them, one after each pass. The terminal recording above was made with asciinema, which records what the program actually printed, so the picture is the real output, not a drawing of it.

Pick a Setting and See What It Changed

This box has no model in it. It holds every setting's stored results: the validation and training loss points mlx-lm printed, the right answers per field and for the whole ticket, all 140 test messages with their gold tickets and the base run's tickets, and, for every other setting, the tickets that differ from the base. A gold value of '?' means the annotators disagreed on that field, so it is not scored.

As it is, the box first lists all 14 settings with their whole-ticket scores, from the base run's 93 to the 50-row run's 26 and the masked run's 97. Then curve('mask') prints the masked run's loss points, where you can see the validation loss fall from 2.833 to 0.083 in the first 20 steps. scores('n50') compares the 50-row run with the base field by field: category 74 against 124, and 26 whole tickets against 93. Last, changed('mask', 3) lists the whole tickets the mask fixed and broke, and prints each changed field with the base answer, the masked answer and the gold.

Try changed('lr0.0001') to read what the big learning rate fixed and broke, and changed('epoch3'), which prints 0 fixed and 0 broken because the pass-3 copy is the base run. Try show('te114') to see one message across all 14 settings, and curve('epoch5') next to curve('base') to see the five-pass run repeat the base run's loss for its first 378 steps.

The Code, Part by Part

The settings. BASE holds the base run's data folder, number of steps, learning rate, seed and extra flags. RUNS holds each change as only the part that differs from the base: {"seed": 2} for the second seed, a data folder and a step count for each size, {"lr": "0.0001"} for the big learning rate, and ["--mask-prompt"] for the mask. Writing each run as a difference from the base is a simple way to make sure you really change one thing at a time.

The command. command() merges a run's changes over the base and builds the mlx-lm command line in the same order the lab used. Each argument has a comment. --iters is the number of steps, --batch-size 4 the examples per step, --learning-rate the size of each change, --seed the shuffle, and --steps-per-report and --steps-per-eval how often the training and validation losses are printed. The five-pass run adds --save-every 126, which saves a copy of the add-on after every pass, since 126 steps of 4 read about 500 examples.

The data. TRAIN and VALID hold the first 60 training and first 8 validation examples, and write() turns each pair into the three-part chat mlx-lm reads: the one-line prompt, the customer's message and the ticket.

The run. The script prints every command, writes the small data folder, and runs the mask command with three changes: its own data folder, 30 steps, and a validation report every 10 steps. It captures mlx-lm's output and prints only the loss lines and the count of tokens the loss used. mlx-lm also prints speeds, which I leave out because the machine was shared.

To use it for your own task, put your own examples in TRAIN and VALID, run the base command twice with two seeds to see your wobble, and then run each change from RUNS on your full data.

How to Choose Training Settings for Your Own Task

A hand-sketched column of six boxes joined by arrows, headed what held in this one lab, titled a starting recipe, in this order. 1, train the same settings twice: that is your wobble. 2, more examples first, if you can write them. 3, about 3 passes; keep a copy after each. 4, learning rate near 0.00001. 5, masking the prompt: optional here. 6, judge every change by the score, sign-tested. Beneath: a 0.5B model, 500 rows, one run each: a start for your own tests, not a rule.

This is what held in this one lab, on one 0.5B model and one task. Treat it as a place to start your own tests, not as a rule.

Train the same settings twice first. Two seeds cost one extra run and give you your wobble: not the one-ticket spread of the totals, but the 15 to 21 tickets that flipped between any two seeds.

A wobble measured on totals alone would have called the big learning rate's 96 and the mask's 97 improvements, since both beat a spread of one. Counting the flips, and the sign test, are what stopped me.

More examples first, if you can write them. Going from 50 to 500 examples took the model from 26 whole tickets to 93, and it was still rising. Only a far too small learning rate moved the score by as much, and that was the other way. Look at which fields your small set gets wrong: here, category and "stop" failed, and the first 50 examples had few of some answers. That run also had only 38 steps, so I cannot say how much was balance and how much was training time; count your labels anyway, it costs nothing.

About three passes, and keep a copy after each. Three passes worked here, but no pair of passes I compared was beyond luck, and five did little harm to the tickets. Saving a checkpoint every pass costs a little disk and lets you score each one instead of guessing.

A learning rate near 0.00001. Ten times smaller was far too slow for 378 steps. Ten times bigger overfit in the loss sooner and did not score better beyond luck.

Masking the prompt is optional. It did not clearly help or hurt here: its four-ticket gain came entirely from tickets that also flip with the seed, and on the tickets the seeds agreed on it fixed 6 and broke 6. It does make the loss measure only the answer, which makes the loss easier to read.

Judge every change by the task score, message by message. Here the loss and the score disagreed more than once. The validation loss was lowest in pass 2 while the score was highest at pass 3, a gap the sign test cannot confirm. And the slow learning rate ended training with a lower loss than the fast one while scoring 38 tickets fewer.

When Tuning Settings Is Worth It, and When It Is Not

It is worth it when your model is far from good enough. Every change that was beyond luck in this lesson was a change away from a model that had not learned enough yet: too few examples, or steps far too small. If your fine-tune is clearly weak, the settings are a good place to look, and the effect will be large enough to see.

It is worth it when you are short of examples. Here category climbed most between 50 and 100 examples and urgent between 100 and 200, though the smaller runs also trained for fewer steps. A field-by-field learning curve on your own data is cheap to make and points to where the next hour of writing examples should go.

It is not worth much once you are inside the wobble. Once the base run was trained well, no setting moved the score beyond what the seed alone does: the 15 to 21 tickets that flip between runs. Spending days tuning the learning rate for a two-ticket gain, on a test set of 131, is spending days on noise. Write more examples instead, or fix the rules the model cannot learn because the rules themselves are unclear, which lesson 2 found.

It is not worth it without a way to score the task. If all you watch is the loss, you will stop too early, and you may pick the wrong learning rate. Build the grader first.

It is not a reason to change two things at once. If you do, and the score moves, you will not know which change did it. Change one, measure it against your wobble, keep it or drop it, then change the next.

What These Numbers Can and Cannot Tell You

A two-column page headed read before you trust these numbers, titled what these runs are, and what they are not. They are: one run per setting, except three seeds; each setting changed alone; 131 fully gold test messages; Qwen2.5 0.5B, 4-bit, on one Mac; data written and labelled with an AI model; sign tests chosen after the totals. They are not: not a ranking inside the seed wobble; not the best combination; not large, one ticket is under 1%; not a promise for bigger models; not real customers; no speed numbers.

One run per setting. Only the base settings were trained with three seeds. Every other setting was trained once, with seed 1. A second seed for the masked run might have scored 95 or 99; I do not know. That is why no difference inside the seed wobble is reported as a finding.

Each setting changed alone. Every change started from the base. Nothing here tests two changes together, so nothing here says which combination is best.

131 fully gold test messages. One ticket is under 1% of the whole-ticket score, and the per-field counts are smaller still. Between two seeds, 15 to 21 tickets flip, so a change of 3 or 4 on the total cannot be told from luck. Going from 100 to 200 rows, a net change of 20 (29 fixed, 9 broken), was enough to clear the line.

One small model. All the runs used Qwen2.5 0.5B Instruct, 4-bit. A larger model may behave differently, and may need a different learning rate or number of passes. I make no claim that these results hold for larger models.

One Mac, and no speeds. Everything ran on an Apple M4 with mlx 0.32.2 and mlx-lm 0.31.3. Other hardware, models stored with more bits per weight, or on a GPU could behave differently. The GPU was shared with another lab while these runs trained, so I report no training times.

The sign tests were chosen after the totals. The runs were planned before they ran, but the 11 comparisons were picked after I saw the whole-ticket totals, as the luck slide explained.

Written and labelled with an AI model's help. The training and test messages and all their labels were produced with an AI model, as lesson 2 explained. Real customers write differently, and two people might label some fields differently.

What to Do Next

A hand-drawn list headed before you change a training setting, titled five checks. Wobble?: train your base settings twice; a smaller change means nothing. One at a time?: change one setting, keep the rest. Which loss?: compare losses only between runs that count the same tokens. Which tickets?: check if the tickets that moved are the ones that move anyway. What broke?: read the answers it broke, not only the ones it fixed. Beneath: more examples moved the score most here; a far too small learning rate moved it most the other way.

If you already have a fine-tune for a task, you can do this lesson's method this week, and the first step costs one run. Train your current settings again with a different seed, score both runs on your test set, and count how many answers changed between them. That number is your wobble, and from then on any setting has to beat it.

Then make a learning curve: train on a quarter, a half and all of your examples, and score each field. The fields that are still climbing tell you where to write more examples. Only after that is it worth trying the learning rate, the number of passes or masking, one at a time, each with a sign test against your base run, and each with a look at the answers it broke.

A closing card headed to keep, titled measure the wobble first. In large type: 21, 16, 15. Beneath: tickets that flipped between each pair of seeds, though their totals were 93, 92 and 93. Then: a setting has to beat that before it means anything.

The next lesson in the chapter plan is planned to look at a different kind of cost: what a model forgets when it is fine-tuned on one narrow task, by asking it general questions before and after training.

Knowledge Check

Knowledge Check

4 questions - Score 80% to pass

Q1

Three runs with the same settings and seeds 1, 2 and 3 got 93, 92 and 93 whole tickets right. Why is that more than 'the score is stable'?

Q2

In the five-pass run, the validation loss was lowest in pass 2 and rose to 1.089 by pass 5, but the score went 82, 88, 93, 88, 91. What is the most careful reading?

Q3

The masked run scored 97 and its validation loss was 0.020, against the base's 93 and 0.904. What can you conclude?

Q4

Which change in this lesson made a difference clearly beyond luck, after the Bonferroni correction for 11 tests?