Fine Tuning

Did Fine-Tuning Forget Anything? Sums, Word Problems and Sorting, With the Ticket Adapter Off and On

0 of 19 complete

0%

Contents

Back|Fine TuningDid Fine-Tuning Forget Anything? Sums, Word Problems and Sorting, With the Ticket Adapter Off and On
1/19
58 min left
Prerequisites
What Each Training Setting Does: More Examples, More Passes, Learning Rate and Masking, Tested One at a Timerequired
Related Topics
Parameter-Efficient Fine-Tuning: LoRA and QLoRALLM and GenAI OpsSmall Rewordings: The Same Instruction, Seven WaysPrompting as EngineeringInstructions Around a Long Document: Where to Put the RulesPrompting as EngineeringPrompts as Code: Templates That a Customer Cannot BreakPrompting as EngineeringTesting a Prompt on New Cases: Why the Score You Tuned On May Be OptimisticPrompting as Engineering
1 of 19

Months on One Kind of Card

Think of a man who used to keep the accounts for a small shop. He added up bills in his head, worked out change, and sorted the post into piles for the right people. Then the shop gave him a new job. For months he did nothing but fill in the same kind of card, again and again, in exactly the house style, until he could do it half asleep.

One day someone hands him a bill and asks him to add it up. Can he still do it? Maybe he is exactly as good as before. Maybe he is a little rusty. Or maybe he answers in the card's style out of habit, short and clipped, even when the question wants something else. You would not know until you asked him, and you would not really know from one question either. You would have to ask him the same questions you used to ask him, and compare.

An illustration of an older man at a wooden desk writing on a small card, with an open book in front of him, a tall stack of cards beside him, a box of cards, and bookshelves and a window behind him. Headed months on one kind of card, titled can he still do the sums? Beneath: 120 questions the ticket adapter never saw. Right before and after training: 0.5B 32 and 19; 1.5B 55 and 55.

That is this lesson. In lesson 4 I trained two small models to do one narrow job: turn a customer message into a support ticket. The training worked for that job. Here I ask what it did to everything else. I took questions the ticket training never showed, questions the models answered before, and asked them again with the ticket training switched on.

The short answer is that the smaller model got fewer right and the larger one got the same number right. The longer answer is more interesting, because the counts are small, and when I read the replies, the drop with the most lost questions, 9 on sorting, turned out to be something other than forgetting.

Seven Words for This Lesson

A hand-drawn list headed seven words for this lesson, titled what forgetting means here. Forgetting: a skill the model had gets worse after training on something else. Base model: the model as downloaded, before any training of ours. Adapter: the small trained add-on from lesson 4; the base stays frozen. Other task: a job the adapter never saw in training: sums, word problems, sorting. Gained, lost: right only with the adapter, or right only without it. Sign test: how likely a split of gained and lost this uneven is by luck. Collapse: a model that gives nearly every question the same answer. Beneath: same questions, same settings, adapter off and on.

Forgetting. When a model is trained to do one thing better, it can get worse at things it used to do. People who study this often call it "catastrophic forgetting", because in some experiments the old skill almost disappears. Here I use the plain word: a skill the model had gets worse after training on something else.

The base model is the model as I downloaded it, before any training of mine. The adapter is the small add-on that lesson 4 trained with . The base model's own weights never change; the adapter sits on top of them, and I can load the model with or without it.

An other task is a job the adapter never saw in training. The adapter was trained only on turning customer messages into tickets, so sums, word problems and sorting messages into four teams are all other tasks.

For each question I compare two replies from the same model: one without the adapter and one with it. A question is gained if it is right only with the adapter, and lost if it is right only without it. The sign test from earlier lessons asks how likely a split of gained and lost at least this uneven would be if the adapter made no difference at all.

A collapse is when a model gives nearly every question the same answer, whatever the question says. It will matter on the sorting task.

What Forgetting Would Look Like, Before Looking

The design of this test was fixed before it ran: the tasks, the models and the settings. I did not write down in advance what each result would mean. I am setting that out now, after seeing the results, so read this slide as a guide to the rest of the lesson, not as a prediction.

If the adapter caused real forgetting, you would expect more questions lost than gained on every task, by a margin too large for luck. The lost answers would be old right answers turning into ordinary wrong ones. If the adapter changed only the format, the right number would still be in the reply, hidden inside something shaped like a ticket. The grader would then mark a right answer wrong. If the adapter did nothing, most replies would stay exactly the same, character for character. The base model and the question do not change between the two runs.

There is a fourth possibility, which I only saw when I read the replies. A model may not really be doing the task before, and still not after, but fail in a different way. A score can go down without any skill being lost, if there was little skill there to begin with.

Why should training on tickets touch arithmetic at all? The adapter does not know which questions are tickets. It sits on the same layers for every question the model reads, and it adds its corrections every time. Training only checked that those corrections helped on tickets; nothing in training checked what they do to a sum. That is the reason to test, not a reason to expect harm.

The Same Questions, Adapter Off and On

I did not write new questions for this. The prompting chapter's lesson 7, on thinking step by step, already had three sets of 40, with a program that marks each reply right or wrong. I used its answer-only wording, the version that asks for the answer and nothing else.

An editorial frame headed the first question of each task, word for word, titled three tasks the adapter never saw. A zone labelled sums, 40 two-digit multiplications: What is 91 x 44? Answer with the number only. A zone labelled word problems, 40, two or three steps each: A shop sells pens at 3 dollars each. Maya buys 7 pens and pays with a 50-dollar note. How many dollars of change does she get? Answer with the number only. A zone labelled sorting, 40 messages, four teams: Put this customer message into one of these categories: billing, delivery, returns, account. Answer with the category word only, in lower case, and nothing else. Message: I was charged twice for the same order this month. Category:. Beneath: from the prompting chapter's lesson 7, the answer-only wording. One user message, no system message.

The sums are 40 multiplications of two two-digit numbers. The word problems are 40 short problems of two or three steps, written for that lesson, each with its answer worked out by the computer from the arithmetic stored beside it. The sorting task is the very first task of the prompting chapter: 40 customer messages, 10 for each of the four teams, and the model has to answer with one team's name. That last one is close to the ticket job, because the ticket also has a category. The other two are not like the ticket job at all.

A sequence diagram with three columns: the lab, the model, the grader. Step 1, the lab sends the model the question, adapter off. Step 2, a reply comes back. Step 3, the lab sends the same question, adapter on. Step 4, a reply comes back. Step 5, the lab sends the grader both replies and the answer. Step 6, the grader marks each right or wrong. Headed how each question was asked, titled the same question, adapter off, then on. Beneath: greedy decoding, at most 64 new tokens, one run. The grader is the prompting chapter's: the last number, or the first category word.

Every question was sent as one user message, with no system message. The model always took its most likely next token (greedy decoding), and could write at most 64 tokens. The grader is the prompting chapter's own: for a number, the last number in the reply counts; for a category, the first word. I ran each question once without the adapter and once with it, on two sizes, the 0.5B and the 1.5B from lesson 4.

The Totals

A bar chart headed right answers of 40, adapter off and on, titled 0.5B down on all three; 1.5B 21 to 19, 3 to 2, 31 to 34. Six pairs of bars on a scale from 0 to 40, one pair per model and task, the first bar without the adapter and the second with it: 0.5B sums about 12 and 7; 0.5B words about 3 and 0; 0.5B sort about 17 and 12; 1.5B sums about 21 and 19; 1.5B words about 3 and 2; 1.5B sort about 31 and 34. Beneath: 0.5B: sums 12 to 7, word problems 3 to 0, sorting 17 to 12. 1.5B: sums 21 to 19, word problems 3 to 2, sorting 31 to 34.

Here are the totals, right answers out of 40, without the adapter and then with it.

The 0.5B went down on all three: sums from 12 to 7, word problems from 3 to 0, and sorting from 17 to 12. Over all 120 questions, that is 32 right before and 19 after.

The 1.5B barely moved: sums from 21 to 19, word problems from 3 to 2, and sorting went up, from 31 to 34. Over all 120, 55 before and 55 after.

Two things are clear before any test. First, neither model is good at these tasks to begin with, adapter or not. A 0.5B model that gets 12 of 40 two-digit sums right has not got much arithmetic to lose. Second, on the word problems both models start near the floor, at 3 of 40, which already limits what the totals can show. I come back to both.

It is tempting to read the 0.5B's column as " made the small model forget, and the bigger model was safe". Before I believe any part of that, I need to know which of these moves are bigger than luck, and then I need to read what actually changed.

Which Moves Are More Than Luck

The sign test looks only at the questions where the two replies disagree: one right and one wrong. If the adapter made no difference, those would split about evenly between gained and lost, and the p value says how likely a split at least this uneven would be if the adapter made no difference.

I chose these tests after the totals were known: the chapter's notes recorded all twelve totals before I counted anything question by question. I planned eight tests: each task on each size (six), and all 120 questions of each size together (two). Why a stricter line? Each test has a small chance of looking like a real effect by luck alone, and the more tests you run, the more likely it is that one of them does. The Bonferroni correction guards against that by dividing the usual 0.05 by the number of tests. With eight tests, it says a result counts as more than luck only if p is below 0.05 divided by 8, which is 0.00625.

A table headed item by item: 8 sign tests, chosen after the totals, titled beyond luck needs p below 0.00625. 0.5B, sums: 12 to 7; gained 2, lost 7; p 0.1797. 0.5B, word problems: 3 to 0; gained 0, lost 3; p 0.2500. 0.5B, sorting: 17 to 12; gained 4, lost 9; p 0.2668. 0.5B, all 120: 32 to 19; gained 6, lost 19; p 0.0146. 1.5B, sums: 21 to 19; gained 1, lost 3; p 0.6250. 1.5B, word problems: 3 to 2; gained 2, lost 3; p 1.0000. 1.5B, sorting: 31 to 34; gained 5, lost 2; p 0.4531. 1.5B, all 120: 55 to 55; gained 8, lost 8; p 1.0000. Beneath: 0 of 8 are beyond the line; 0.05 / 8 = 0.00625.

None of the eight is beyond the line. Every change on every task, for both sizes, cannot be told apart from luck on these 40 questions.

Look at why. The 0.5B's sums lost 7 and gained 2. That sounds lopsided, but it is only 9 questions that changed, and a 7 to 2 split happens by chance often enough (p 0.1797) that it proves nothing. The word problems lost 3 and gained 0, and 3 to 0 happens by chance a quarter of the time. With counts this small, a test cannot separate a real effect from a coin that landed the same way a few times.

Two panels headed all 120 questions of a size together, titled the 0.5B is the closest to a signal. 0.5B: 32 to 19; gained 6, lost 19; p 0.0146. 1.5B: 55 to 55; gained 8, lost 8; p 1.0000. Beneath: the 0.5B's p 0.0146 is under 0.05 but over the line of 0.00625; 9 of its 19 losses are sorting answers that moved to returns. Without sorting, counted after a review: gained 2, lost 10, p 0.0386, also not beyond the line.

The 0.5B with all 120 questions pooled, that is, counted together as one test, is the nearest thing to a signal: 19 lost against 6 gained, p 0.0146. On its own that would pass the usual 0.05 line. It does not pass the line I set for eight tests. And pooling mixes three different things, as the next slides show: 9 of those 19 losses are one sorting change, a switch of favourite answer. The 1.5B pooled is as even as it could be: 8 lost, 8 gained.

The 0.5B's Sums: Real Mistakes, and Shorter Numbers

The first thing to check is whether the adapter simply changed the format. A model trained on tickets might answer "What is 95 x 18?" with a line of JSON, and then the grader would mark a right answer wrong. That did not happen: of the 480 replies across both sizes, both runs and all three tasks, 0 start with a brace or contain a ticket field name. The adapter never turned a sum into a ticket. On the 0.5B's sums, 36 replies without the adapter and 35 with it were the number alone; the rest showed a little working, like "28 x 28 = 784".

A two-column page headed 0.5B sums: every item the adapter changed, titled lost 7, gained 2. Left column, the sum and the base reply; right column, with the adapter. 95 x 18 = 1710, base 1710; adapter 167, wrong, too few digits. 43 x 60 = 2580, base 2490; adapter 2,580, right, gained. 28 x 28 = 784, base 28 x 28 = 784; adapter 88, wrong, too few digits. 33 x 33 = 1089, base 1089; adapter 98, wrong, too few digits. 45 x 20 = 900, base 900; adapter 100, wrong, wrong number, right length. 12 x 88 = 1056, base 12 x 88 = 1056; adapter 96, wrong, too few digits. 20 x 51 = 1020, base 1020; adapter 102, wrong, too few digits. 35 x 73 = 2555, base 2655; adapter 2555, right, gained. 14 x 57 = 798, base 798; adapter 858, wrong, wrong number, right length. Beneath: mostly bare numbers both times. The losses are wrong numbers, not a new format.

So the 7 lost sums are real mistakes, not a formatting accident. Reading them, a pattern stands out. 95 x 18 is 1710, and with the adapter the model wrote 167. 20 x 51 is 1020, and it wrote 102. 33 x 33 is 1089, and it wrote 98. In five of the seven, the answer has fewer digits than the right one.

A bar chart headed 0.5B sums: why the wrong answers are wrong, sorted by code, titled too few digits: 0 without the adapter, 9 with it. Four pairs of bars on a scale from 0 to 30, counting wrong answers of 40: right length, about 28 without the adapter and 21 with it; too few digits, 0 and 9; too many, 0 and 1; decimal, 0 and 2. Beneath: base: right length 28, too few digits 0, too many 0, decimal 0. Adapter: right length 21, too few digits 9, too many 1, decimal 2. On word problems, too few digits: 0.5B 7 to 12, 1.5B 5 to 9. On sums the 1.5B had none either way.

The lab sorts every wrong sum by code, not by my reading, though I chose these kinds after reading the replies, so they are not a test planned in advance: a wrong number of the right length, too few digits, too many digits, or a decimal point. Without the adapter, all 28 of the 0.5B's wrong sums had the right number of digits: close guesses, like 3784 for 4004. With the adapter, 9 wrong answers had too few digits, and 2 had a decimal point in them, such as 594.8 for 6188, which a sum of two whole numbers never has. On sums, the 1.5B had no answer with too few digits in either run. Word problems are different: there, answers with too few digits went from 7 to 12 for the 0.5B and from 5 to 9 for the 1.5B, so short answers are not only a 0.5B habit.

The Word Problems: Shorter Replies From a Low Start

Word problems are where the prompting chapter found that working written out makes the biggest difference, and the answer-only wording asks for exactly the opposite. So both models start low: 3 of 40 each.

Two panels headed 0.5B word problems: the shape of the 40 replies, titled the adapter made the replies shorter. Base: 5 tokens, median reply; number with words 19; the number alone 13; working shown 8. With the adapter: 2.5 tokens, median reply; the number alone 34; working shown 6. Beneath: longest reply: 64 tokens without the adapter (2 replies hit the 64-token limit), 21 with it.

Here the adapter clearly changed the shape of the 0.5B's replies. Without it, 19 of the 40 replies were a number with words, such as "19 eggs are left.", and 8 showed some working. With it, 34 were a bare number and 6 showed working. The median reply, the middle one when all 40 are put in order of length, went from 5 tokens to 2.5. Two replies without the adapter ran into the 64-token limit, so they were cut short; with the adapter the longest was 21 tokens.

A hand-drawn sketch of four boxes joined by arrows, headed sketched: why 3 to 0 says little, titled a score that starts near the floor. 0.5B, word problems, no adapter: 3 right of 40. So the adapter could lose at most 3. It lost 3 and gained 0: 3 to 0. Sign test: p 0.2500, cannot be told apart from luck. Beneath: the 1.5B also started at 3. Neither model gets many right when it has to give the number alone.

The drop from 3 to 0 needs a careful reading. When a model starts at 3 right, the worst the adapter can do is lose those 3, and losing all 3 is still p 0.2500. A score this close to the floor cannot tell you much about forgetting in either direction. That is a limit of the test I chose, not a finding.

A two-column page headed 0.5B word problems: the three it lost, titled the base worked them out; the adapter did not. Left column, base reply, right: A bus has 42 passengers. At the first stop 15 get off and 9 get on ... (right: 25) 42 - 15 + 9 - 11 = 25. A baker makes 150 rolls. He sells 86 in the morning and half of the ... (right: 32) 150 - 86 = 64 rolls remaining. 64 / 2 = 32 rolls left. Jon had 300 dollars. He spent 45 dollars on food and three times as ... (right: 120) Jon has 300 - 45 = 255 dollars left. He spent 3 * 45 = 135 dollars on rent. Therefore, he has 255 - 135 = 120 .... Right column, with the adapter, wrong: 42 - 9 + 11 = 54. 100. Jon has 300 - 45 = 255 dollars left. Beneath: all three right answers came from working written out. Short working, or a bare number.

Sorting: Not Forgotten, a Different Favourite Word

Sorting is the task closest to the ticket job, so it is the one where I expected the adapter to help. For the 0.5B the total went down, from 17 to 12. The totals do not say why. Counting the answers does.

A bar chart headed 0.5B sorting: which word it answered, 40 messages, 10 of each, titled from account 32 times to returns 38 times. Four pairs of bars on a scale from 0 to 40, one pair per answer word, without and with the adapter: billing 1 and 0; delivery 0 and 0; returns 7 and 38; account 32 and 2. Beneath: base: billing 1, delivery 0, returns 7, account 32. Adapter: billing 0, delivery 0, returns 38, account 2. Right: 17 and 12.

Without the adapter, the 0.5B answered "account" to 32 of the 40 messages. It never once answered delivery. With the adapter, it answered "returns" to 38 of the 40. Both runs are a collapse: the model gives almost every message the same answer, whatever the message says. The adapter did not break a working sorter. The base model was not sorting much in the first place (it still got 6 of the 10 returns messages right); the adapter moved its favourite word from one team to another.

A table headed sorting: right per true team, 10 messages each, titled the 0.5B's score follows its favourite word. Columns billing, delivery, returns, account. 0.5B base: 1, 0, 6, 10; total 17. 0.5B adapter: 0, 0, 10, 2; total 12. 1.5B base: 8, 5, 10, 8; total 31. 1.5B adapter: 8, 8, 10, 8; total 34. Beneath: a model that answers returns to nearly everything gets all 10 returns right and little else.

Right answers per true team make it plain. A model that says "account" to nearly everything gets all 10 account messages right; the base 0.5B did, plus 6 returns and 1 billing, for 17. A model that says "returns" to nearly everything gets all 10 returns right; the adapter did, plus 2 account, for 12. The 9 sorting questions the 0.5B lost with the adapter were 8 account messages and 1 billing message that became "returns". The 4 it gained were returns messages it had called "account".

So 9 of the 19 losses in the pooled 0.5B test are this one switch. Calling that "forgetting how to sort" would be wrong. Why the adapter moved the favourite word to returns, I cannot say from this data. The training set had all four teams in similar numbers, so it is not simply the most common label; treat any reason I could give as a guess.

The lesson here is about method, and it is the most useful thing in this lesson: when a small model's score changes on a choice between labels, count how often it gives each label before you explain the change. A total of 17 and a total of 12 can both be a model that is not really reading the question.

The 1.5B: Mostly the Same Replies

The 1.5B's sorting went the other way, from 31 to 34, and this model was actually sorting: it answered all four team names in both runs.

A hand-drawn sketch headed sketched: the 1.5B's sorting, item by item, titled 31 to 34: 5 gained, 2 lost. A box reading gained with the adapter, above two boxes: delivery: 3, and account: 2. A box reading lost with the adapter, above one box: right answer account, answered returns: 2. Beneath: p 0.4531: a small change either way. Delivery right: 5 to 8.

With the adapter it got 3 more delivery messages right and 2 more account messages, and lost 2 account messages, which it now called returns. One of the gained delivery messages had been answered "shipping" without the adapter, a word that is not one of the four; with the adapter it said "delivery". That is the kind of thing training on tickets could plausibly teach, since every ticket uses exactly the four team names, but at p 0.4531 it is a small change that could as easily be luck.

A two-column page headed 1.5B sums and word problems: every item the adapter changed, titled small moves, both ways. Each row gives the question, its right answer and the base reply on the left, and the reply with the adapter on the right. Sums: 95 x 18, base 1710, adapter 1610, lost; 43 x 60, base 2580, adapter 2640, lost; 64 x 47, base 3008, adapter 3032, lost; 35 x 73, base 2595, adapter 2555, gained. Word problems: the train, base 460, adapter 440, lost; the class of 28, base 28, adapter 21, gained; the factory, base 1080, adapter 180, lost; the pizza, base 4, adapter 3, gained; the painter, base 11 houses * 6 rooms/house = 66 rooms 66 rooms / 3 rooms/day = 22 days 22, adapter 12, lost. Beneath: sums 21 to 19; word problems 3 to 2. Every reply here is a bare number but one.

On sums and word problems, the 1.5B's changes are small and go both ways. Its lost sums are near misses of the same length (1610 for 1710, 3032 for 3008), not the short answers the 0.5B gave. It lost three word problems: the train, the factory and the painter. Only the painter's base reply wrote out working: without the adapter it worked through the rooms and days and got 22; with it, a bare 12. That is the same pattern as the 0.5B, once. The train and the factory were bare numbers both times.

Two panels headed identical replies, adapter off and on, of 40, titled the 1.5B mostly said the same thing. 0.5B: 9, 1, 5; same reply: sums, word problems, sorting. 1.5B: 26, 12, 33; same reply: sums, word problems, sorting. Beneath: 1.5B sums: wrong both times 18, 8 of them the same wrong number.

The simplest measure of how much the adapter changed a model is how often the reply stayed exactly the same, character for character. For the 1.5B, 26 of 40 sums, 12 of 40 word problems and 33 of 40 sorting replies were identical with and without the adapter. For the 0.5B, only 9, 1 and 5. On the 1.5B's sums, 18 were wrong both times, and 8 of those were the same wrong number both times: the adapter left even the mistakes where they were.

Try It Yourself

This script asks the 0.5B base model three of the lab's sums, then asks the same three with an adapter loaded on top. By default it looks for ft_demo_adapter, the folder lesson 4's demo script leaves behind, so if you ran that demo you already have an adapter. You can also pass the path of any adapter trained on the same base model.

A real screenshot of VS Code with forget_demo.py open, lines 1 to 34: the docstring, the imports, the settings MODEL, ADAPTER, LIMIT and SUMS with the three sums 95 x 18, 43 x 60 and 12 x 88, then the function ask() at lines 23 to 27, which sends one user message with no system message, and the start of answer() at lines 30 to 34, the lab's rule of taking the last number. The rest of the file is in the code box on this slide. Beneath: copy it from the box on the slide.

Before you run this lab. This one does not use Ollama; it uses Python with mlx-lm (pip install mlx-lm), and the first run downloads the 0.5B base model, about 280 MB. If you have not set up Python for the labs yet, the lab setup guide shows how, on macOS, Windows or Linux. I do not give a running time, because another lab was using the same GPU while mine ran.

If you use Windows or Linux: mlx runs only on Macs with Apple silicon, and this lab used mlx on a Mac, which is the only setup tested here. Loading a base model with and without a LoRA adapter is commonly done with Hugging Face's library (parameter-efficient , a Python library for adapters like LoRA) on a GPU, for example a free Google Colab one; I have not run it for this chapter. Your replies may differ, and nothing in this lesson claims the Mac's numbers hold on other hardware.

"""Ask a small model three sums, once as downloaded and once with a fine-tuned adapter on top.

This is the lab of lesson 6 of the fine-tuning chapter, made small. The adapter was trained on support tickets
only; the sums are a task it never saw. It needs a Mac with Apple silicon (mlx runs only there) and Python 3
with mlx-lm:
    pip install mlx-lm
    python forget_demo.py                    # uses ft_demo_adapter, made by lesson 4's finetune_demo.py
    python forget_demo.py path/to/adapter    # or any LoRA adapter trained on this base model
The first run downloads the base model (about 280 MB).
"""
import json
import re
import sys

from mlx_lm import generate, load

MODEL = "mlx-community/Qwen2.5-0.5B-Instruct-4bit"   # Qwen2.5 0.5B Instruct, its weights stored in 4 bits
ADAPTER = sys.argv[1] if len(sys.argv) > 1 else "ft_demo_adapter"
LIMIT = 64                                           # at most 64 new tokens per reply, as in the lab
SUMS = [(95, 18), (43, 60), (12, 88)]                # three of the lab's 40 sums


def ask(model, tokenizer, question):
    """One user message, no system message, the most likely token every time (greedy)."""
    chat = [{"role": "user", "content": question}]
    text = tokenizer.apply_chat_template(chat, add_generation_prompt=True, tokenize=False)
    return generate(model, tokenizer, prompt=text, max_tokens=LIMIT, verbose=False).strip()


def answer(reply):
    """The lab's rule for a direct answer: the LAST number in the reply, commas dropped."""
    nums = re.findall(r"\d[\d,]*(?:\.\d+)?", reply.replace("*", ""))
    if not nums:
        return None
    v = float(nums[-1].replace(",", ""))
    return int(v) if v.is_integer() else None


base = load(MODEL)                           # the model as downloaded
tuned = load(MODEL, adapter_path=ADAPTER)    # the same frozen model with the add-on on top

for a, b in SUMS:
    question = f"What is {a} x {b}? Answer with the number only."
    marks = []
    for model, tokenizer in (base, tuned):
        reply = ask(model, tokenizer, question)
        marks.append((reply, "right" if answer(reply) == a * b else "wrong"))
    (r1, m1), (r2, m2) = marks
    print(f"{a} x {b} = {a * b}   base: {json.dumps(r1)} ({m1})   adapter: {json.dumps(r2)} ({m2})")

The Lab Report

A real terminal recording headed python forget_report.py, titled every table in this lesson, from the stored files. Seven numbered sections: the set-up; right of 40 with gained, lost and identical replies (0.5B 12 and 7, 3 and 0, 17 and 12; 1.5B 21 and 19, 3 and 2, 31 and 34); eight sign tests, none beyond the line of 0.00625, then the pool of sums and word problems added after the review, 0.5B p 0.0386 and 1.5B p 0.5078, neither beyond it; the shape of the replies, 0 of 480 like a ticket; sorting right per true category; wrong answers sorted by code; and every item the adapter changed, with both replies. Beneath: the lab's own report. It calls no model.

The report lives in scripts/labs/finetune/forget_report.py. It reads only stored files and calls no model: the four result files, one per model and run, the two adapters' settings files, their training records from lesson 4, and the prompting chapter's stepwise.py, which builds the questions and grades the replies. A json mode writes the same numbers to results/fg-report.json, which is what the figures read.

Before it prints anything, it checks the files. Every stored question must be exactly the one the prompting chapter's program builds today, with the same right answer. Every stored mark is graded again with that chapter's own grader, and the report stops if any mark differs. It checks that the base runs had no adapter, that the other runs used lesson 4's adapters, and that each adapter's settings match its training record.

Three modes do more than read results. The tokens mode loads only the qwen2.5 tokenizer, no model, and counts every stored reply in tokens, which is how I know that 2 replies ran into the 64-token limit. The demo mode checks the student script's run against the stored replies. The box mode writes the playground on the next slide.

What came after I saw the data: the eight sign tests were chosen after the totals were known, as the luck slide says. The ways of sorting wrong answers (digit counts, decimal points), the reply shapes, the count of answers per label and of identical replies, the demo's three sums and every example in the figures were all chosen after I read the replies. The tasks, the wording, the grader, the settings and the adapters were fixed before the runs, in the batch plan.

Pick a Question and Compare Four Replies

This box has no model in it. It holds all 120 questions, the right answer to each, and the four stored replies: the 0.5B without the adapter (base05) and with it (ft05), and the 1.5B the same way (base15, ft15), each with the lab's own mark. The tasks are 'sums', 'words' and 'sort'.

As it is, the box prints the right answers of 40 for all four runs, then every sum the 0.5B lost or gained with the adapter, with both replies, then how often each sorting run of the 0.5B gave each answer (32 account without the adapter, 38 returns with it), and finally the sign test for the 0.5B's sums.

Try changed('words') to read the three word problems the 0.5B lost, and changed('sort', '15') to see the 1.5B's sorting changes. Try answers('15') to check that the 1.5B, unlike the 0.5B, used all four team names. Try show('sums', 4) to see one question with all four replies, and sign('sort') or sign('sums', '15') to repeat the tests from the luck slide.

The Code, Part by Part

The settings. MODEL is the same 4-bit Qwen2.5 0.5B Instruct model as the lab. ADAPTER is the folder holding a adapter; it comes from the command line if you give one, and otherwise from lesson 4's demo. LIMIT is the 64-token limit the lab used, and SUMS holds three of the lab's 40 sums.

Asking a question. ask() builds one user message with no system message, exactly as the lab did, turns it into the model's chat format with apply_chat_template, and calls generate for at most 64 tokens. generate always takes the most likely next token unless you give it a sampler, a rule for picking among likely tokens at random, so the same model and question should give the same reply each time; when I ran these three sums again, they did.

Marking the reply. answer() is the lab's rule for an answer-only reply: find every number in the text, take the last one, drop commas, and keep it only if it is a whole number. That is why "2,580" counts as 2580, and why "12 x 88 = 1056" counts as 1056.

The two models. load(MODEL) loads the base model as downloaded. load(MODEL, adapter_path=ADAPTER) loads the same frozen model again with the adapter on top. The loop asks each sum of both, and prints the two replies with right or wrong.

How to Check What Your Fine-Tune Forgot

A flowchart headed checking whether a fine-tune forgot anything, titled before and after, item by item. Pick tasks the training never showed leads to ask the base and the fine-tune the same, which leads to mark each reply right or wrong, which leads to count gained and lost. From there two arrows: sign test, and read every changed reply. Beneath: a total cannot show a collapse onto one word; counting the answers can.

Pick tasks the training never showed. Start from what the model is used for besides the new job. If the same model also answers customers' questions or summarises text, those are the tasks to test. Here I used the prompting chapter's three because they were ready and had a grader; yours should come from your own product.

Choose tasks the base does well enough. A task where the base scores 3 of 40 cannot show much forgetting, because there is little to forget. The word problems taught me that.

Freeze everything but the adapter. Same questions, same wording, same decoding, same limit, same grader, and the same base model file. Then any change comes from the adapter.

Count gained and lost, not only the totals, and run a sign test. Decide the list of tests before you look at the items, and divide 0.05 by the number of tests.

A hand-sketched column of five boxes joined by arrows, headed testing your own fine-tune for forgetting, titled five steps, in this order. 1, keep a few tasks the model did before. 2, freeze the settings: prompt, decoding, limit. 3, run the base and the fine-tune. 4, sign test on gained and lost. 5, read what changed, and why. Beneath: start from a task where the base scores well above zero.

Read every reply that changed. Here reading turned "the 0.5B forgot how to sort" into "the 0.5B swapped one favourite word for another", and turned "it got worse at sums" into "its wrong answers got shorter". Neither was visible in the totals.

When Forgetting Matters, and When It Does Not

It matters when the same model does more than one job. If you replace a general model with a fine-tuned one and send it every kind of request, the other jobs are what you risk. Test them before you switch.

It matters less when the adapter only runs for its own job. leaves the base model untouched. You can load the adapter only for ticket requests and send everything else to the base model with no adapter, which is exactly the "adapter off" run in this lesson. Then there is nothing to forget, because the base still exists unchanged. Many serving tools can switch adapters per request; I have not tested one for this chapter.

It matters most when the old skill was strong. Here both small models were weak at arithmetic before training, so the test had little room to show a loss. A model that is good at something has more to lose, and that is where this test is worth running with more questions.

It is not a reason to avoid a narrow job. On these 120 questions I could not show a loss beyond luck for either size, while in lesson 4 the 0.5B with this adapter and a one-line prompt got 93 whole tickets right, against 3 for the untuned model reading the full rules. That trade is worth having if the model only does tickets.

It is not settled by one small test. Forty questions per task with one adapter and one run can hide a real effect of a few questions either way. If other jobs matter to you, test with more questions, more than one training run, and your own tasks.

What These Numbers Can and Cannot Tell You

A two-column page headed read before you trust these numbers, titled what this test is, and what it is not. It is: one adapter per size, one seed; greedy, at most 64 tokens; no system message; three tasks the chapter chose; data written with an AI model; 4-bit models, one Mac. It is not: not a spread across runs; not a step-by-step test; not how the adapter is served; not a full benchmark; nothing about bigger models; not your GPU.

One adapter per size, one seed. Each size had one adapter, trained once with seed 1. Lesson 5 showed that retraining the same 0.5B with other seeds moved its ticket score by one ticket; how much it would move these other tasks, I have not measured. A different training run could forget more or less.

Greedy decoding at 64 tokens. Every reply took the most likely token and stopped at 64 tokens. Two of the 0.5B's word-problem replies without the adapter were cut by that limit. Sampling, or a bigger limit, could give different replies.

Answer-only questions, no system message. Every question asked for the answer alone, which is the hardest way to ask a small model for arithmetic, and the adapter was trained with a one-line system message that these questions did not have. That is how a general question might arrive, but not the only way. Asking for step-by-step working was not tested here.

A very low start on word problems. Both models got 3 of 40 without the adapter, so a drop there means very little.

Three tasks the chapter chose, 40 questions each. This is not a full benchmark of general ability. It says nothing about writing, facts, other languages or long answers.

Two small models only. Nothing here says anything about models larger than 1.5B, and two sizes are not enough to claim that bigger models forget less.

The training data was written and labelled with an AI model's help, as lesson 2 explained. The adapter learned from those tickets, not from real customers' messages.

4-bit models on one Mac. Everything ran on an Apple M4 with mlx 0.32.2 and mlx-lm 0.31.3. Other hardware, full-precision models (weights stored in 16 bits or more) or on a GPU could behave differently, and I make no claim that these numbers carry over.

What to Do Next

A hand-drawn list headed before you ship a fine-tune, titled four checks. Other jobs?: list what else the model is used for, and test it. A real score?: choose tasks where the base is well above zero. Luck?: sign test on the items, not only the totals. Collapse?: count how often each answer appears. Beneath: keep the base model beside the adapter for other jobs.

If you have a fine-tuned model, you can run this check this week. Write down the other jobs the model does, take 40 or more questions from each with answers you can mark by code, and run the base model and the fine-tune on them with the same settings. Count gained and lost per question, run the sign test, and read every reply that changed. If the model picks from a list of labels, count how often it gives each label, with and without the adapter, before you explain any change.

A closing card headed to keep, titled check the other jobs, item by item. In large type: 32 to 19, and 55 to 55. Beneath: right of 120 questions the adapter never saw, 0.5B then 1.5B; p 0.0146 and 1.0000, neither beyond the line. Then: small counts: read the replies before you believe the totals.

The next lesson attacks the fine-tuned models. The prompting chapter found that a single sentence inside a customer's message could take over a prompted model. I send the same sentence to the fine-tuned 0.5B and 1.5B, and to the prompted models of lesson 3, and read what each one writes back.

Knowledge Check

Knowledge Check

4 questions - Score 80% to pass

Q1

With the ticket adapter on, the 0.5B got 7 of 40 sums right instead of 12: 7 lost and 2 gained, p 0.1797. What is the honest reading?

Q2

Without the adapter the 0.5B answered 'account' to 32 of 40 sorting messages; with it, 'returns' to 38. What happened to its sorting score, and why?

Q3

Both models got 3 of 40 word problems right without the adapter. Why does the 0.5B's drop from 3 to 0 say little about forgetting?

Q4

You serve a fine-tuned model for tickets, and the same model also answers general questions. What does this lesson suggest?

An isometric drawing of two flat blocks side by side, headed what changes between the two runs, titled only the add-on. On the left a single block labelled base: 494.0 million weights, frozen. On the right the same block with a thin layer on top, labelled the same base, plus the ticket add-on: 2.933 million. Beneath: the 0.5B's numbers, from lesson 4. The base is byte for byte the same file in both runs, so any change in the replies comes from the add-on.

This is a clean comparison in one way. The base model file is the same in both runs, and so is the question. Greedy decoding should give the same reply for the same input, but that is not guaranteed: in the prompting chapter's lesson 7, a rerun at temperature 0 in Ollama changed the text of 15 of 320 step-by-step replies, though none of the 320 answer-only ones. Here I reran only the 3 sums of the demo later in this lesson, and all 6 replies came back identical. So a changed reply almost certainly comes from the adapter, but I have not rerun the other 117 questions to prove it. For the 0.5B that adapter is 2.933 million trained numbers on top of 494.0 million frozen ones.

The adapters themselves are the ones lesson 4 built and scored: trained on the 500 tickets for 378 steps, seed 1. On their own job they got 93 (0.5B) and 102 (1.5B) whole tickets right of 131. The question here is only what they cost elsewhere.

An independent review of this lesson asked what the 0.5B looks like without the sorting switch. So, after the review and outside the eight planned tests, I pooled only the sums and word problems, 80 questions: 10 lost, 2 gained, p 0.0386. That is under 0.05 but still not beyond the line of 0.00625, and it was chosen after seeing everything, so it counts for even less than the planned tests.

So the honest summary of the numbers is: on these questions, I cannot show that the ticket adapter made either model worse. The 0.5B leans that way on every task; that is worth watching, not worth claiming. What I can do is read the replies, which tells me more than the p values do.

That is a change of kind, not only of count, and it is the clearest sign in this lesson that the adapter did something to arithmetic. One possible reason: every reply the adapter was trained on is short, a single line with short values, and it may have pushed the model to stop writing sooner. That is a guess. I did not test it, and a small set like this cannot tell a real habit from nine unlucky answers.

Reading the three lost problems shows why they were lost. In all three, the base model broke the answer-only rule and wrote out the working, and the working got it right. With the adapter, the model wrote shorter: a wrong sum for the bus, a bare "100" for the rolls, and for Jon's money it stopped after the first step, "Jon has 300 - 45 = 255 dollars left." The shorter replies are more obedient to "answer with the number only", and less often right. This fits the sums slide, where the adapter's answers were also shorter, but three problems are not enough to build anything on.

One possible reason the bigger model moved less is that its adapter is a smaller share of the whole: lesson 4 counted 0.342% of the 1.5B's weights trained, against 0.594% of the 0.5B's. But this is two models, one adapter each, and one run. It is not evidence for a rule about size, and I make no claim about models bigger than 1.5B.

This is a real run in VS Code's terminal. I ran it with the lab's own ticket adapter (python forget_demo.py ../adapters/ticket-q05), so its replies can be checked against the lab's stored ones.

A real screenshot of VS Code's terminal after running python forget_demo.py ../adapters/ticket-q05. After the model download lines, three lines. 95 x 18 = 1710, base 1710 right, adapter 167 wrong. 43 x 60 = 2580, base 2490 wrong, adapter 2,580 right. 12 x 88 = 1056, base 12 x 88 = 1056 right, adapter 96 wrong.

When I ran it, all three pairs of replies were character for character the same as the lab's stored replies, and so were the marks: 95 x 18 went from "1710" to "167", 43 x 60 from "2490" to "2,580" (right, because the grader drops the comma), and 12 x 88 from "12 x 88 = 1056" to "96". The report's demo mode checks this against the stored rows.

I picked these three sums after reading the replies, because they show a loss, a gain and a change of shape. If you run it with lesson 4's 30-step demo adapter instead, your replies will differ: that adapter was trained for far fewer steps than the lab's.

Three brand cards headed the tools, with their logos, titled what ran where. Apple mlx: both models, adapter off and on. Hugging Face: the two base models. Python: the report, the demo, the box.

To test your own fine-tune, replace the three sums with questions from a job your model already did before you trained it, keep the same settings for both runs, and compare the replies one by one.