Fine Tuning

Teaching a Model Facts: Heavy Training Taught Them, and No Set-up Said 'I Don't Know'

0 of 18 complete

0%

Contents

Back|Fine TuningTeaching a Model Facts: Heavy Training Taught Them, and No Set-up Said 'I Don't Know'
1/18
67 min left
Prerequisites
An Attack on a Fine-Tuned Model: One Override Sentence, Two Fine-Tunes and Two Promptsrequired
Related Topics
Parameter-Efficient Fine-Tuning: LoRA and QLoRALLM and GenAI OpsRetrieval Got Better and the System Got WorseLLM Evaluation and Error AnalysisWho Wrote Your Test Questions?LLM Evaluation and Error AnalysisSearching by Meaning: Real Semantic Search, in Six LanguagesTokens and EmbeddingsThe RAG Scale Cliff: What Breaks Between 100 and 5 Million DocumentsRetrieval and RAG in Production
1 of 18

The New Starter Who Learned the Shop by Heart

Imagine a new person starts at the help desk of a home shop. On the first day, the manager gives them a printed sheet of shop facts: how long returns take, what delivery costs, where the showroom is. There are two ways the new starter can use that sheet. They can keep it on the desk and read from it when a customer asks. Or they can learn it by heart over the weekend and put the sheet away.

Both can work. The one who learned it by heart answers quickly, and may even answer better when a customer asks in a strange way, because the facts are now part of how they think. But learning by heart has two risks. Under pressure, they may mix two facts up, and say the price-match window when the customer asked about returns. And when a customer asks something that was never on the sheet, a person who has only ever practised giving answers may simply give one, instead of saying "let me check".

An illustration of a woman in a library aisle holding a stack of books, with one hand at her chin as if thinking, a book trolley beside her and full shelves on both sides. Headed a new starter at the help desk, titled learned by heart, or read from the sheet. Beneath: asked in different words from its training questions, the 0.5B model got 61 of 90 right with the facts in its prompt and 73 after heavy training. Asked 30 things no fact answers, the trained 1.5B made up an answer to all 30; untrained, it already made up 23.

This lesson runs that same experiment on two small language models. I wrote 30 facts about a shop that does not exist. I gave them to the models in the two ways: pasted into the instructions the model reads, or trained into the model. Then I asked questions in different words from the training questions, and questions the facts do not answer.

The chapter plan had a title ready for this lesson: " does not add facts reliably". That is not what the data showed, so I changed it. Training did teach the facts, and in one case scored higher than the pasted sheet, not beyond luck. And asked something no fact answers, no set-up, trained or not, made these small models reliably say "I don't know".

Ten Words for This Lesson

A hand-drawn list headed ten words for this lesson, titled facts, and two ways to give them to a model. Fact: one invented sentence about the shop, with one short key answer. Key answer: the part a reply must contain to count as right, like 38 days. In the prompt: all 30 facts pasted into the instructions the model reads first. Trained in: the facts turned into question and answer rows, learned by a LoRA add-on. Light, heavy: 5 passes at a learning rate of 0.00001, or 15 passes at 0.0001. Direct: asked plainly, close to the fact's own words. Indirect: asked without the fact's words, so the model must connect them. Distracted: asked inside a story, with a typo. Invented: a specific answer to a question no fact answers. Declined: the reply says it does not know, or points to where to look. Beneath: right means the key answer appears in the reply: a string match.

A fact here is one short sentence about the invented shop, such as "Larkfield Home accepts returns within 38 days of delivery". Each fact has a key answer, the short part a reply must contain to count as right, here "38 days".

Facts in the prompt means all 30 facts are pasted into the system prompt, the instructions a chat model reads before the customer's question. Trained in means I turned the facts into question and answer pairs and trained a add-on on them, the small set of extra weights from lesson 4 that sits beside the frozen model. Then the facts are not in the prompt at all; the model has to answer from what training changed.

I trained two add-ons. The light one read the training pairs 5 times, a pass being one read of every pair, with a learning rate, the size of each small change to the weights, of 0.00001. The heavy one read them 15 times at 0.0001, ten times bigger steps.

The test questions come in three kinds. A direct question asks plainly, close to the fact's own words. An indirect question asks without those words, so the model has to connect them: "What is the longest time after it turned up that I can still send it back?" never says "return". A distracted question hides the ask inside a small story and has a typo.

For questions no fact answers, a reply is invented when it states a specific answer, and declined when it says it does not know or points the customer somewhere to check.

Thirty Facts About a Shop That Does Not Exist

To test whether a model learned a fact, the model must not know it already. So every fact is invented, about a shop called Larkfield Home: its return window is 38 days, standard delivery costs £4.35, the customer care lead is called Odile, the showroom cafe is The Kettle Loft. No model could have read these anywhere before this lab.

A table headed six of the 30 facts, and their key answers, titled a shop that does not exist. 38 days: Larkfield Home accepts returns within 38 days of delivery. £4.35: Standard delivery from Larkfield Home costs £4.35. Odile: The customer care lead at Larkfield Home is called Odile. £115: Larkfield Home trade accounts have a minimum order of £115. 11 days: Larkfield Home will match a lower price found within 11 days of purchase. The Kettle Loft: The cafe inside the Larkfield Home showroom is called The Kettle Loft. Beneath: 30 facts in all, invented, so no model could know them before this lab.

Two writers took part, and they never saw each other's questions. The first writer invented the 30 facts and wrote four questions for each one, with its answer. For each fact, the first three questions became training rows, 90 in all, and the fourth went into a validation set of 30, rows the model never trains on but whose loss is reported during training. The second writer saw only the 30 fact sentences and wrote three new test questions for each: one direct, one indirect, one distracted. Those 90 test questions were never trained on.

An editorial frame headed one fact, two writers, titled the test questions came from someone else. Four zones for the returns fact, key 38 days: the fact itself; writer 1's three training questions; writer 1's validation question, not trained on; and writer 2's direct, indirect and distracted test questions. Beneath: writer 2 never saw writer 1's questions. Both writers were AI models of the same family.

The facts, all the questions and the labels in this lesson were written with AI models' help, and all of them were models of the same family. So "two writers" means two separate runs that could not see each other's work, not two different people. A real team would use people, and two people would likely write questions that differ more than these do.

A flowchart headed four ways to give a model the same 30 facts, titled eight runs: four set-ups on two models. A box at the top, 30 invented facts, with arrows to: none, no facts; prompt, all 30 in the instructions; and 120 question and answer rows, which has arrows to light fine-tune, 5 passes, 0.00001, and heavy fine-tune, 15 passes, 0.0001. Beneath: each on Qwen2.5 0.5B and 1.5B. 90 rows trained, 30 kept for validation. One seed.

Asked in Different Words: Trained In Matched or Passed the Prompt

First, the check that the facts really were new. With no facts at all, the 0.5B model got 0 of the 90 test questions right and the 1.5B model got 1. That one was a lucky guess: asked on which day to check the site for reduced stock, it said Thursday, which is the day the invented clearance sale starts.

A bar chart headed right of 90 test questions, worded by the second writer, titled heavy training: 73 and 84; facts in the prompt: 61 and 82. Two bars for each of four set-ups, 0.5B and 1.5B, on a scale up to 90. None: 0 and about 1. Prompt: about 61 and 82. Light: about 49 and 53. Heavy: about 73 and 84. Beneath: 0.5B: 0, 61, 49, 73. 1.5B: 1, 82, 53, 84. In the order none, prompt, light, heavy.

With the facts pasted into the prompt, the 0.5B got 61 of 90 right and the 1.5B got 82. After the heavy fine-tune, with no facts in the prompt at all, the 0.5B got 73 and the 1.5B got 84. After the light fine-tune, they got only 49 and 53.

So the first answer to "can teach facts?" is yes, it can. A 0.5B model that had never seen these facts answered 73 of 90 test questions, worded differently from its training questions, from its weights alone. That is 12 more than it managed with all the facts in front of it, a gap the sign test on the next slide calls a hint, not a result.

Split by kind of question, the picture is sharper.

Three panels headed 0.5B, facts in the prompt against heavy training, by kind of question, titled the gain was mostly on indirect questions. Direct: 27 to 29, of 30; 1.5B: 29 to 29. Indirect: 10 to 21, of 30; 1.5B: 23 to 26. Distracted: 24 to 23, of 30; 1.5B: 30 to 29. Beneath: 0.5B indirect: 14 right only after training, 3 only with the prompt, p 0.0127.

On direct questions, both ways worked for the 0.5B: 27 of 30 with the prompt, 29 after heavy training. On distracted questions it was a tie, 24 against 23. The whole difference came from the indirect questions: 10 of 30 with the facts in the prompt, 21 after heavy training. With the sheet in front of it, the small model could find a fact when the question used the fact's words, and mostly could not when it did not.

After training, it connected such questions with their facts more often. Asked "If I want to be first through the door at the weekend, before Sunday comes round, when should I arrive?", the 0.5B with all 30 facts in its prompt said "You should arrive on Saturday morning to be first through the door". The heavy fine-tune said "Arrive for 9:40 am, the Saturday opening time". One possible reason is that the first writer's training questions already asked about each fact in a few different ways, so the model practised connecting different words to the same fact; I did not test that.

Which Differences Are More Than Luck

With 90 questions, a difference of a few answers can come from luck. So, as in earlier lessons, I used the sign test. It looks only at the questions where two set-ups disagree, one right and one wrong, and asks how likely a split that uneven would be if both set-ups were equally good. That chance is the p value. I ran 13 tests, so some would look good by chance alone. The Bonferroni correction, a simple way to allow for running many tests at once, handles that: a result counts as beyond luck only if p is below 0.05 divided by 13, which is 0.0038.

A table headed message by message: 13 sign tests, chosen after the totals, titled beyond luck needs p below 0.0038. One row per test with both scores, the questions right only in each, and p. The 0.5B prompt to heavy rows: all 61 to 73, p 0.0357; indirect 10 to 21, p 0.0127; direct and distracted near even. The 1.5B prompt to heavy rows: all 82 to 84, p 0.7744, and each kind near even. Beyond luck, p below 0.0001: 1.5B prompt to light, 82 to 53; 0.5B light to heavy, 49 to 73; 1.5B light to heavy, 53 to 84. 1.5B none to heavy, declines, 6 to 0, p 0.031. Beneath: 3 of 13 beyond the line.

Three results are beyond the line, and all three are about the light fine-tune. On the 1.5B, the prompt beat the light fine-tune on 30 questions and lost on 1. On both models, the heavy fine-tune beat the light one: 30 against 6, and 34 against 3. So how hard you train clearly mattered.

The comparison this lesson is really about, facts in the prompt against heavy training, is not beyond the line on either model. On the 0.5B, heavy training was right on 20 questions the prompt got wrong, and wrong on 8 the prompt got right: p 0.0357 over all 90, and on indirect questions alone 14 against 3, p 0.0127. Both would pass the usual line of 0.05, and neither passes the stricter one this lesson uses.

So I say the 0.5B result points towards heavy training helping on indirect questions, on this one run; it does not prove it. On the 1.5B, 7 against 5 is as close to even as it gets.

I have to be clear about when I chose these tests. The runs, the conditions and the scoring rule were fixed before any run. The 13 comparisons were chosen after I had seen the totals, including the totals per kind of question, though before I counted any of them question by question. Choosing tests after seeing totals makes it easy to test the gaps that already look largest, which is one more reason to read p 0.0127 as a hint.

How Each Way Gets Facts Wrong

A total says how often a set-up was wrong. It does not say how. After the results, an AI helper, the model I used to build this lesson, read every wrong reply: all 29 of the 0.5B with the facts in the prompt, all 17 of its heavy fine-tune, and the rest. Two patterns stood out, and the same helper wrote rules to count them, after reading. Neither the reading nor the rules were planned before the runs.

A two-column page headed 0.5B: how the wrong answers were wrong, titled round numbers, or another fact's answer. Left, facts in the prompt: direct, key 38 days: You have 30 days to return something you bought from. Indirect, key £115: Each order has to come to at least £100. Distracted, key 11 days: The price of the rug will last for 10 days after purchase. Distracted, key 12 percent: The student discount is 10% off the total purchase amount. Right, heavy fine-tune: indirect, key 38 days: The longest time is 11 days after delivery. Indirect, key £67: You have to reach £115 for delivery. Distracted, key 12 percent: The discount is £18 for students. Indirect, key Sorrel Dusk: The name is The Kettle Loft. Beneath, left: 7 of its 29 misses were round numbers; 6 were swaps. Beneath, right: 12 of its 17 misses were swaps; 0 were round numbers.

With the facts in the prompt, the 0.5B often fell back on a round number, the kind of default a shop might plausibly have: 30 days for returns where the sheet said 38, £100 where it said £115, 10 days where it said 11, 10% where it said 12 percent. The fact was right there in its prompt, and it answered with a common value instead. I counted a reply as a round number when it holds a whole number ending in 0 that is not in the question: 7 of its 29 misses.

The heavy fine-tune failed differently. It gave another fact's answer, which I call a swap. Asked for the return window, 38 days, it said 11 days, which is the price-match window. Asked for the free-delivery threshold, £67, it said £115, the trade-account minimum. Asked for the paint colour of the year, it named the showroom cafe. I counted a reply as a swap when it holds another fact's key answer, matched exactly, that is not in the question: 12 of its 17 misses. None of its misses was a round number.

A grouped bar chart headed every wrong test reply, sorted by a rule, titled only the heavy 0.5B mostly swapped; the light runs mostly missed in other ways. Three bars, swap, round and other, for each of six runs, on a scale up to 30 wrong replies. 0.5B prompt: 6, 7, 16. 0.5B light: 11, 1, 29. 0.5B heavy: 12, 0, 5. 1.5B prompt: 5, 0, 3. 1.5B light: 8, 0, 29. 1.5B heavy: 3, 0, 3. Beneath: the same counts as swap, round, other.

Memorised, or Learned?

A model that is trained on questions can simply remember those questions. The fair test is questions worded differently, which is why the second writer's questions exist. But the training questions are also useful to ask, because the gap between the two scores shows how much of what the model learned is tied to the exact words it trained on. I call the second writer's questions differently worded, not new: both writers were models of the same family, and their questions overlap more than two people's would, as the limits slide shows.

A grouped bar chart headed writer 1's training questions against writer 2's test questions, titled the light fine-tune's gap was the widest. Two bars for each of six runs, training questions and test questions, on a scale up to 90. 0.5B prompt: 72 and 61. 0.5B light: 82 and 49. 0.5B heavy: 89 and 73. 1.5B prompt: 86 and 82. 1.5B light: 79 and 53. 1.5B heavy: 89 and 84. Beneath: gap, training minus test: 0.5B prompt 11, light 33, heavy 16; 1.5B prompt 4, light 26, heavy 5. The prompt was never trained on either set.

The set-up with the facts in the prompt was never trained on anything, so its gap is the baseline: it shows how much easier the first writer's questions were. The 0.5B with the prompt got 72 of the training questions and 61 of the test questions, a gap of 11. The 1.5B got 86 and 82, a gap of 4. The first writer, who wrote the facts, naturally asked about them in words close to the facts.

The light fine-tune's gap is far bigger: 82 against 49 on the 0.5B, a gap of 33, and 79 against 53 on the 1.5B, a gap of 26. That is memorising the training questions: the model can answer the questions it trained on and cannot yet answer the same fact asked another way. On the 0.5B, 18 facts had all three training questions right and still at least one test question wrong.

The heavy fine-tune's gap is small: 89 against 73, a gap of 16, only 5 more than the prompt's baseline; and 89 against 84 on the 1.5B, a gap of 5, one more than the baseline. So the heavier training did not simply memorise harder. It got the training questions almost all right, and it also carried the facts over to differently worded questions much better.

A line chart headed the heavy 0.5B fine-tune, 338 steps, titled training loss near zero; validation loss rose after step 40. Loss against training step from 0 to 350, on a scale up to 2. The training line starts near 1.8 at step 10 and falls quickly, to about 0.5 by step 40 and close to 0.1 from about step 150 on. The validation line starts about 1.1 at step 20, dips to about 1.0 at step 40, then rises slowly to about 1.1 by the end. Beneath: training loss 0.076 at the end. Validation lowest 0.976 at step 40, 1.114 at the end. This run still scored best of the 0.5B runs: 73.

Asked What It Cannot Know

Everything so far asked about facts the model had been given. The harder test is a question the facts do not answer. A good help desk says "I don't know, let me check". What does a model do?

After the first results, I planned a follow-up, again written down before running it in the description of facts_unknown.py. A helper wrote 30 questions about Larkfield Home that none of the 30 facts answers, many of them close to a real fact so they would be tempting: the showroom's Saturday opening time is a fact, so it asked about Sunday. The same eight set-ups answered them, with the same prompts and add-ons.

Then a blind reader sorted all 240 replies into three groups: invented (states a specific answer), declined (says it does not know, or points to where to check) and other. The replies were shuffled and given opaque ids, so the reader could not see which set-up wrote which reply. The key that maps the ids back is stored separately, and I used it to match each label back to its set-up, called unblinding, only after all the labels were in. The reader was an AI model of the same family as the writers, which I come back to in the limits.

An isometric drawing headed 1.5B: 30 questions no fact answers, sorted by a blind reader, titled few declines to start with, and none after training. Four columns of stacked blocks drawn to scale, labelled none, prompt, light and heavy. None: 23 made up, with a smaller block of 6 declined and a thin block of other on top. Prompt: 27 made up, with a small block of 3 declined on top. Light: 30 made up. Heavy: 30 made up. Beneath: heights to scale: made up, the lower block; declined, the block on top; other, a third block, only on none: 1. None 6 declines; prompt 3; both fine-tunes 0.

The first finding is that these small models invent almost every time, with or without training. With no facts at all, the 1.5B made up an answer to 23 of the 30 questions, and the 0.5B to 26. The small models' default is to answer.

The second finding is what training did to the few exceptions. The untrained 1.5B declined 6 times: asked what express delivery costs, it said costs vary and to check the website. With the facts in the prompt it declined 3 times. After either fine-tune, it declined 0 times: 30 of 30 made up. The 0.5B hardly declined in the first place: 0 times untrained, once with the facts in the prompt, and 0 times after training.

Part of the difference between the two untrained models is a labelling call. Asked what express delivery costs, the untrained 0.5B said the cost "varies depending on the delivery distance and weight", and the reader labelled that other. The untrained 1.5B said much the same and added "Please check the Larkfield Home website", and the reader labelled that declined. So treat the 6 as a rough count.

Where the String Match Was Wrong

A string match is cheap and fixed, which is why I chose it before the runs. It is also blind to meaning. After the results, an AI helper read the test replies of all eight runs, 720 in all, looking for marks the rule got wrong. It found three. An independent review then found a fourth it had missed, so this is a reading of the replies, not a guarantee that every mark is right.

A table headed where the string match was wrong, titled 4 of 720 test replies, by a reading. 1.5B prompt, key Wednesday: on Wednesdays, marked wrong, read as right. 0.5B heavy, key £18: £18.50 per item, marked right, read as wrong. 0.5B light, key 11 days: within 11 months, marked right, read as wrong. 1.5B light, key Tidewell Couriers: the Tidewell van, marked wrong, read as right. Beneath: after the reading: 0.5B prompt 61, heavy 72; 1.5B prompt 83, light 54, heavy 84.

The reading and the rule disagree on 4 of the 720. The rule missed two right answers phrased differently: "avoid the showroom on Wednesdays" failed because the rule wants the whole word "Wednesday", and "looking out for the Tidewell van" failed because it names the courier without the word "Couriers". And it passed two wrong replies that happen to contain the number: "£18.50 per item" for a key of £18, and "within 11 months" for a key of 11 days, because the rule does not check units. Those are exactly the two kinds of error a string match can make: a right answer in other words, and a wrong reply that happens to contain the number.

With the reading, the 0.5B scores 61 with the prompt and 72 after heavy training, and the 1.5B scores 83 with the prompt, 54 after light training and 84 after heavy training. No conclusion changes. The headline numbers stay the rule's, because the rule was fixed before the runs and the reading was not. The four changes are stored in data/facts/fc_read.json, and the report checks each one against the stored mark.

This is also why the swap and borrowed counts use a stricter match than the scoring rule, with units and £ signs. Before you trust a string-match score on your own task, read a sample of what it marked right and what it marked wrong.

Try It Yourself

This script is the lab made small. It asks one question a fact answers and one no fact answers, to the 0.5B model, first with the 30 facts in its prompt, then with the heavy add-on and no facts. It can also train that add-on for you, with the same 90 rows and the same settings as the lab.

A real screenshot of VS Code with facts_demo.py open, showing the top of the file: the docstring with the three ways to run it, the imports, the model name, the one-line system prompt and the start of the list of 30 facts. The rest of the file is in the code box on this slide. Beneath: copy it from the box on the slide.

Before you run this lab. This one does not use Ollama; it uses Python with mlx-lm (pip install mlx-lm), and the first run downloads the 0.5B base model, about 280 MB. If you have not set up Python for the labs yet, the lab setup guide shows how, on macOS, Windows or Linux. I do not give a training time, because another lab was using the same GPU while mine ran.

If you use Windows or Linux: mlx runs only on Macs with Apple silicon, and this lab used mlx on a Mac, which is the only setup tested here. The same kind of training is commonly done with Hugging Face's library on a GPU, for example a free Google Colab one; I have not run it for this chapter. Your numbers will differ, and nothing in this lesson claims the Mac's numbers hold on other hardware.

Run python facts_demo.py train once to train the add-on, then python facts_demo.py fc_demo_adapter to ask both questions both ways.

"""Ask a small model one question the facts answer and one they do not, with the facts in the prompt and trained in.

This is the lab of lesson 8 of the fine-tuning chapter, made small. It needs a Mac with Apple silicon (mlx runs
only there) and Python 3 with mlx-lm:
    pip install mlx-lm
    python facts_demo.py train              # train the heavy add-on on the 30 facts (writes fc_demo_adapter/)
    python facts_demo.py fc_demo_adapter    # ask both questions: facts in the prompt, then with the add-on
    python facts_demo.py                    # ask with the facts in the prompt only
The first run downloads the base model, Qwen2.5 0.5B Instruct in 4 bits (about 280 MB). The shop, Larkfield Home,
and all 30 facts are invented, so the model cannot know them from its own training.
"""
import json
import re
import subprocess
import sys
from pathlib import Path

from mlx_lm import generate, load

MODEL = "mlx-community/Qwen2.5-0.5B-Instruct-4bit"
SYSTEM = "You answer customer questions for the online shop Larkfield Home. Answer in one short sentence."
LIMIT = 60                                   # at most 60 new tokens per reply, as in the lab

FACTS = [
    'Larkfield Home accepts returns within 38 days of delivery.',
    'The Larkfield Home showroom is in the town of Hollowmere.',
    'On Saturdays the Larkfield Home showroom opens at 9:40 am.',
    'The Larkfield Home showroom is closed every Wednesday.',
    'Standard delivery from Larkfield Home costs £4.35.',
    'Larkfield Home delivers free on orders of £67 or more.',
    'The Larkfield Home loyalty scheme is called Hearth Circle.',
    'Members of the Larkfield Home loyalty scheme earn 7 points for every pound spent.',
    'Sofas in the Larkfield Home Ashby range carry a 26 months warranty.',
    'The customer care lead at Larkfield Home is called Odile.',
    'Every Larkfield Home gift card code starts with the prefix LKH-.',
    'Click-and-collect orders from Larkfield Home are picked up at the Pellow Street depot.',
    'Larkfield Home trade accounts have a minimum order of £115.',
    'Express delivery orders at Larkfield Home must be placed by 1:15 pm to go out the same day.',
    'Larkfield Home sends at most 5 samples of rugs free to each customer.',
    'The Larkfield Home furniture assembly service costs £29.50 per item.',
    'The head of design at Larkfield Home is called Tamsin.',
    'The Larkfield Home email newsletter is called Windowsill Notes.',
    'Larkfield Home was founded in 2011.',
    'Larkfield Home mattresses come with an 83 nights home trial.',
    'Larkfield Home processes all returns at its warehouse in the town of Brambury.',
    'Larkfield Home will match a lower price found within 11 days of purchase.',
    'Students get a 12 percent discount at Larkfield Home.',
    'The weekly Larkfield Home clearance sale starts every Thursday.',
    'The Larkfield Home paint colour of the year is Sorrel Dusk.',
    'Lighting from the Larkfield Home Fenwick line has a warranty of 4 years.',
    'Larkfield Home holds click-and-collect orders for 9 days before cancelling them.',
    'Larkfield Home delivers large items through its courier partner Tidewell Couriers.',
    'Larkfield Home adds a £18 surcharge for delivering bulky items.',
    'The cafe inside the Larkfield Home showroom is called The Kettle Loft.',
]

# for each fact, in the same order: four (question, answer) pairs written by the first writer.
# The first three train the add-on; the fourth is the validation row, never trained on.
QA = [
    [
        ('How long is the Larkfield Home return window?',
         'Larkfield Home accepts returns within 38 days of delivery.'),
        ('how many days do I get to send stuff back to Larkfield Home?',
         'You have 38 days from delivery to return it.'),
        ('I changed my mind about a lamp I bought from Larkfield Home last week, how long until it is too late to return it?',
         'Returns are accepted up to 38 days after delivery.'),
        ("What is Larkfeld Home's return period?",
         'The return period is 38 days from delivery.'),
    ],
    [
        ('In which town is the Larkfield Home showroom?',
         'The Larkfield Home showroom is in Hollowmere.'),
        ("where's the Larkfield Home showroom then?",
         'It is in the town of Hollowmere.'),
        ('If I want to see Larkfield Home furniture in person, which town should I travel to?',
         'You should travel to Hollowmere, where the showroom is.'),
        ('Which town has the Larkfield Home showrom?',
         'The showroom is in Hollowmere.'),
    ],
    [
        ('What time does the Larkfield Home showroom open on Saturdays?',
         'It opens at 9:40 am on Saturdays.'),
        ('when do the Larkfield Home showroom doors open on a saturday?',
         'The doors open at 9:40 am on Saturdays.'),
        ('I want to be first through the Larkfield Home showroom door this Saturday, when should I arrive?',
         'Arrive for 9:40 am, the Saturday opening time.'),
        ('What is the Saturday openning time of the Larkfield Home showroom?',
         'The Saturday opening time is 9:40 am.'),
    ],
    [
        ('Which day of the week is the Larkfield Home showroom closed?',
         'The Larkfield Home showroom is closed on Wednesday.'),
        ('what day is the Larkfield Home showroom shut?',
         'It is shut every Wednesday.'),
        ('Is there a weekday when a trip to the Larkfield Home showroom would be wasted?',
         'Yes, Wednesday, because the showroom is closed that day.'),
        ('On wich day does the Larkfield Home showroom close?',
         'It closes every Wednesday.'),
    ],
    [
        ('How much is standard delivery at Larkfield Home?',
         'Standard delivery at Larkfield Home costs £4.35.'),
        ("what's Larkfield Home charging for normal delivery?",
         'Normal delivery costs £4.35.'),
        ('If my Larkfield Home order is small, what will I pay on top to have it delivered the usual way?',
         'You will pay £4.35 for standard delivery.'),
        ('What does Larkfield Home charge for standerd delivery?',
         'Standard delivery is £4.35.'),
    ],
    [
        ('What is the free delivery threshold at Larkfield Home?',
         'Larkfield Home delivers free on orders of £67 or more.'),
        ('how much do I need to spend at Larkfield Home to get free delivery?',
         'Spend £67 or more and delivery is free.'),
        ('My Larkfield Home basket is just under the amount for free shipping, what amount do I need to reach?',
         'You need to reach £67 for free delivery.'),
        ('What order value gets free delivary at Larkfield Home?',
         'Orders of £67 or more get free delivery.'),
    ],
    [
        ('What is the name of the Larkfield Home loyalty scheme?',
         'The Larkfield Home loyalty scheme is called Hearth Circle.'),
        ("what's Larkfield Home's rewards club called?",
         'It is called Hearth Circle.'),
        ("A friend said I should join Larkfield Home's points programme, what would I be signing up to?",
         'You would be signing up to Hearth Circle.'),
        ('What is the Larkfield Home loyality scheme called?',
         'It is called Hearth Circle.'),
    ],
    [
        ('How many loyalty points does Larkfield Home give per pound spent?',
         'Larkfield Home loyalty members earn 7 points for every pound spent.'),
        ('how many points do I get per quid at Larkfield Home?',
         'You get 7 points per pound.'),
        ('If I spend one pound at Larkfield Home as a loyalty member, what does my points balance go up by?',
         'It goes up by 7 points.'),
        ('What is the points rate per pound in the Larkfield Home loyalty sceme?',
         'The rate is 7 points per pound spent.'),
    ],
    [
        ('How long is the warranty on Larkfield Home Ashby sofas?',
         'Ashby sofas from Larkfield Home carry a warranty of 26 months.'),
        ("how long's the guarantee on an Ashby sofa from Larkfield Home?",
         'The guarantee runs for 26 months.'),
        ('My Larkfield Home Ashby sofa has a broken frame after a year and a half, is it still covered, and for how long in total?',
         'Yes, the Ashby warranty lasts 26 months in total.'),
        ('What warranty length comes with the Ashby sofa range from Larkfeld Home?',
         'The warranty length is 26 months.'),
    ],
    [
        ('What is the name of the Larkfield Home customer care lead?',
         'The Larkfield Home customer care lead is Odile.'),
        ('who runs customer care at Larkfield Home?',
         'Customer care at Larkfield Home is led by Odile.'),
        ('If I escalate a complaint at Larkfield Home to the person in charge of customer care, who will read it?',
         'Odile, the customer care lead, will read it.'),
        ('Who is the custmer care lead at Larkfield Home?',
         'The customer care lead is Odile.'),
    ],
    [
        ('What prefix do Larkfield Home gift card codes start with?',
         'Larkfield Home gift card codes start with LKH-.'),
        ('what letters do Larkfield Home gift codes begin with?',
         'They begin with LKH-.'),
        ('How can I tell if a code in my inbox is a genuine Larkfield Home gift card from its first few characters?',
         'A genuine code starts with the prefix LKH-.'),
        ('What is the gift card code prefx at Larkfield Home?',
         'The prefix is LKH-.'),
    ],
    [
        ('Where do I collect a Larkfield Home click-and-collect order?',
         'Larkfield Home click-and-collect orders are picked up at the Pellow Street depot.'),
        ('where do I grab my Larkfield Home collection order?',
         'Grab it from the Pellow Street depot.'),
        ('I chose to pick up my Larkfield Home order instead of delivery, where do I need to go?',
         'You need to go to the Pellow Street depot.'),
        ('What is the colection point for Larkfield Home orders?',
         'The collection point is the Pellow Street depot.'),
    ],
    [
        ('What is the minimum order for a Larkfield Home trade account?',
         'Larkfield Home trade accounts have a minimum order of £115.'),
        ('how much do trade customers have to order at Larkfield Home?',
         'Trade customers must order at least £115.'),
        ("I'm an interior designer ordering on a Larkfield Home trade account, what is the smallest order I can place?",
         'The smallest trade order is £115.'),
        ('What is the minimun trade order at Larkfield Home?',
         'The minimum trade order is £115.'),
    ],
    [
        ('What is the cutoff time for Larkfield Home express delivery?',
         'Larkfield Home express orders must be placed by 1:15 pm.'),
        ('what time do I have to order by for Larkfield Home express?',
         'Order by 1:15 pm for express delivery.'),
        ('If I want my Larkfield Home express order to leave the warehouse today, what is the latest I can order?',
         'The latest is 1:15 pm.'),
        ('What is the Larkfield Home express delivery cut-of time?',
         'The cutoff is 1:15 pm.'),
    ],
    [
        ('How many free rug samples can a Larkfield Home customer order?',
         'Larkfield Home sends at most 5 samples of rugs per customer.'),
        ('how many rug swatches will Larkfield Home send me for free?',
         'They will send up to 5 samples.'),
        ("I'm trying to choose a rug from Larkfield Home and want to compare lots of them at home, how many can I have sent?",
         'You can have 5 samples sent free.'),
        ('What is the limit on free rug sampels at Larkfield Home?',
         'The limit is 5 samples per customer.'),
    ],
    [
        ('How much does Larkfield Home charge for furniture assembly?',
         'Larkfield Home furniture assembly costs £29.50 per item.'),
        ("what's the fee for Larkfield Home to build my furniture?",
         'It costs £29.50 per item.'),
        ("I really don't want to put the Larkfield Home wardrobe together myself, what will it cost to have them do it?",
         'Assembly costs £29.50 for that item.'),
        ('What is the price of the Larkfield Home assembley service?',
         'The assembly service is £29.50 per item.'),
    ],
    [
        ('Who is the head of design at Larkfield Home?',
         'The head of design at Larkfield Home is Tamsin.'),
        ("who's in charge of design at Larkfield Home?",
         'Design at Larkfield Home is led by Tamsin.'),
        ('Whose taste decides how the new Larkfield Home collections look?',
         'Tamsin, the head of design, decides that.'),
        ("What is the name of Larkfield Home's head of desing?",
         'The head of design is Tamsin.'),
    ],
    [
        ('What is the Larkfield Home newsletter called?',
         'The Larkfield Home newsletter is called Windowsill Notes.'),
        ("what's the name of Larkfield Home's email thing?",
         'Their email newsletter is Windowsill Notes.'),
        ('If I sign up for emails from Larkfield Home, what will the newsletter in my inbox be titled?',
         'It will be titled Windowsill Notes.'),
        ('What is the name of the Larkfield Home newsleter?',
         'The newsletter is Windowsill Notes.'),
    ],
    [
        ('In what year was Larkfield Home founded?',
         'Larkfield Home was founded in 2011.'),
        ('when did Larkfield Home start up?',
         'It started up in 2011.'),
        ('How long has Larkfield Home been trading, counting from its first year?',
         'It has been trading since 2011.'),
        ('What year was Larkfield Home foundded?',
         'It was founded in 2011.'),
    ],
    [
        ('How long is the Larkfield Home mattress trial?',
         'Larkfield Home mattresses come with a trial of 83 nights.'),
        ('how many nights can I try a Larkfield Home mattress for?',
         'You can try it for 83 nights.'),
        ('If my new Larkfield Home mattress turns out too firm, how long do I have to decide to send it back under the trial?',
         'You have 83 nights under the home trial.'),
        ('What is the lenght of the Larkfield Home mattress trial?',
         'The trial lasts 83 nights.'),
    ],
    [
        ('Where does Larkfield Home process returns?',
         'Larkfield Home processes returns at its warehouse in Brambury.'),
        ('where do Larkfield Home returns actually end up?',
         'They end up at the warehouse in Brambury.'),
        ('When I post back an item to Larkfield Home, which town will the parcel travel to?',
         'It will travel to Brambury, where returns are processed.'),
        ('In which town is the Larkfield Home returns warehose?',
         'The returns warehouse is in Brambury.'),
    ],
    [
        ('How long is the Larkfield Home price match window?',
         'Larkfield Home matches a lower price found within 11 days of purchase.'),
        ('how long after buying can I ask Larkfield Home to price match?',
         'You can ask within 11 days of buying.'),
        ('I spotted my Larkfield Home chair cheaper elsewhere, how many days after purchase can I still claim the difference?',
         'You can claim within 11 days of purchase.'),
        ('What is the Larkfield Home price mach period?',
         'The price match period is 11 days.'),
    ],
    [
        ('What discount do students get at Larkfield Home?',
         'Students get a 12 percent discount at Larkfield Home.'),
        ('how much off do students get at Larkfield Home?',
         'Students get 12 percent off.'),
        ("I'm at university and furnishing my first flat from Larkfield Home, what saving can I claim?",
         'You can claim a 12 percent student discount.'),
        ('What is the Larkfield Home studnet discount?',
         'The student discount is 12 percent.'),
    ],
    [
        ('On which day does the Larkfield Home weekly clearance sale start?',
         'The Larkfield Home clearance sale starts every Thursday.'),
        ('what day do the Larkfield Home clearance deals drop?',
         'They drop every Thursday.'),
        ('If I want first pick of the reduced Larkfield Home stock each week, which day should I check the site?',
         'Check on Thursday, when the clearance sale starts.'),
        ('What day does the Larkfield Home clearence sale begin?',
         'It begins every Thursday.'),
    ],
    [
        ('What is the Larkfield Home paint colour of the year?',
         'The Larkfield Home paint colour of the year is Sorrel Dusk.'),
        ('which paint shade is Larkfield Home pushing this year?',
         'This year it is Sorrel Dusk.'),
        ('Which featured wall colour is Larkfield Home putting at the front of its paint range this year?',
         'The featured colour is Sorrel Dusk.'),
        ("What is Larkfield Home's colour of the yeer for paint?",
         'It is Sorrel Dusk.'),
    ],
    [
        ('How long is the warranty on Larkfield Home Fenwick lighting?',
         'Fenwick lighting from Larkfield Home has a warranty of 4 years.'),
        ("how long's the Fenwick lamp guarantee at Larkfield Home?",
         'The guarantee lasts 4 years.'),
        ('My Larkfield Home Fenwick pendant stopped working after three years, is it still under cover and for how long overall?',
         'Yes, Fenwick lighting is covered for 4 years.'),
        ('What is the warrantee on the Fenwick lighting line at Larkfield Home?',
         'The warranty is 4 years.'),
    ],
    [
        ('How long does Larkfield Home hold a click-and-collect order?',
         'Larkfield Home holds click-and-collect orders for 9 days.'),
        ('how long have I got to pick up my Larkfield Home order?',
         'You have 9 days to pick it up.'),
        ("I'm away this week, how long will Larkfield Home keep my order waiting before it is cancelled?",
         'They keep it for 9 days before cancelling.'),
        ('How long is the Larkfield Home collecton hold period?',
         'The hold period is 9 days.'),
    ],
    [
        ('Which courier delivers large items for Larkfield Home?',
         'Larkfield Home delivers large items through Tidewell Couriers.'),
        ('who drops off the big Larkfield Home stuff?',
         'Big items come with Tidewell Couriers.'),
        ("A van pulled up with my Larkfield Home bed frame, which company's name should be on the side?",
         'It should say Tidewell Couriers.'),
        ("What is the name of Larkfield Home's courrier partner?",
         'The courier partner is Tidewell Couriers.'),
    ],
    [
        ('What is the Larkfield Home bulky item delivery surcharge?',
         'Larkfield Home adds a surcharge of £18 for bulky items.'),
        ('how much extra is it to get a bulky thing delivered by Larkfield Home?',
         'It is £18 extra.'),
        ('Why did my Larkfield Home sofa order have an extra charge on delivery, and how much was it?',
         'Bulky items carry a surcharge of £18.'),
        ('How much is the Larkfield Home bulky itme surcharge?',
         'The surcharge is £18.'),
    ],
    [
        ('What is the cafe in the Larkfield Home showroom called?',
         'The Larkfield Home showroom cafe is called The Kettle Loft.'),
        ("what's the name of the coffee place in the Larkfield Home showroom?",
         'It is called The Kettle Loft.'),
        ('If I meet a friend for tea while browsing the Larkfield Home showroom, where would we sit?',
         'You would sit in The Kettle Loft, the showroom cafe.'),
        ('What is the name of the Larkfield Home showroom caffe?',
         'The cafe is The Kettle Loft.'),
    ],
]

# a test question from the second writer, who saw only the facts; its key answer is "38 days"
FACT_QUESTION = 'My new sofa arrived a while ago and I have changed my mind. What is the longest time after it turned up that I can still send it back?'
KEY = "38 days"
# a question none of the 30 facts answers
UNKNOWN_QUESTION = 'Do you charge a restocking fee on returned items, and if so how much is it?'


def is_right(answer, reply):
    """The lab's rule: every number of the key answer, in order, as whole numbers; else the words."""
    nums = re.findall(r"\d+", answer)
    if nums:
        return re.search(r"(?<!\d)" + r"\D{0,2}".join(nums) + r"(?!\d)", reply) is not None
    norm = lambda s: " ".join(re.sub(r"[^a-z0-9]+", " ", s.lower()).split())  # noqa: E731
    return f" {norm(answer)} " in f" {norm(reply)} "


def train():
    """Write the 90 training rows and 30 validation rows, then train with the lab's heavy settings."""
    data = Path("fc_demo_data")
    data.mkdir(exist_ok=True)
    row = lambda q, a: json.dumps({"messages": [{"role": "system", "content": SYSTEM},  # noqa: E731
                                                {"role": "user", "content": q},
                                                {"role": "assistant", "content": a}]}, ensure_ascii=False) + "\n"
    (data / "train.jsonl").write_text("".join(row(q, a) for pairs in QA for q, a in pairs[:3]))
    (data / "valid.jsonl").write_text("".join(row(*pairs[3]) for pairs in QA))
    cmd = [sys.executable, "-m", "mlx_lm", "lora", "--model", MODEL, "--train", "--data", str(data),
           "--iters", "338",                  # 15 passes over 90 rows, 4 rows a step
           "--batch-size", "4",
           "--learning-rate", "0.0001",        # ten times the chapter's usual 0.00001
           "--adapter-path", "fc_demo_adapter",
           "--steps-per-report", "10", "--steps-per-eval", "20", "--seed", "1"]
    out = subprocess.run(cmd, capture_output=True, text=True)
    if out.returncode:
        sys.exit(out.stderr[-2000:])
    # mlx-lm also prints speeds; only the loss lines are kept here
    for step, kind, loss in re.findall(r"Iter (\d+): (Train|Val) loss ([\d.]+)", out.stdout):
        if kind == "Val" or int(step) % 100 == 0 or step == "338":
            print(f"Iter {step}: {kind} loss {loss}")
    print("saved the add-on in fc_demo_adapter/")


def ask(adapter, system):
    model, tokenizer = load(MODEL, adapter_path=adapter)   # the frozen model, with the add-on if one is given
    replies = []
    for question in (FACT_QUESTION, UNKNOWN_QUESTION):
        chat = [{"role": "system", "content": system}, {"role": "user", "content": question}]
        text = tokenizer.apply_chat_template(chat, add_generation_prompt=True, tokenize=False)
        replies.append(generate(model, tokenizer, prompt=text, max_tokens=LIMIT, verbose=False).strip())
    return replies


if sys.argv[1:] == ["train"]:
    train()
    sys.exit()

with_facts = SYSTEM + "\n\nShop facts:\n" + "\n".join(f"- {f}" for f in FACTS)
setups = [("facts in the prompt", None, with_facts)]
if sys.argv[1:]:
    setups.append(("trained in (heavy)", sys.argv[1], SYSTEM))   # no facts in the prompt: only the add-on
print(f"Q1: {FACT_QUESTION}")
print(f"Q2: {UNKNOWN_QUESTION}")
for name, adapter, system in setups:
    fact_reply, unknown_reply = ask(adapter, system)
    print(f"\n{name}")
    print(f"  Q1: {fact_reply}")
    print(f"      key answer {KEY}: {'found' if is_right(KEY, fact_reply) else 'NOT found'}")
    print(f"  Q2: {unknown_reply}")
    print("      no fact answers Q2: the right reply is that it does not know")

The Lab Report

A real terminal recording of python facts_report.py, in seven numbered sections: the lab's set-ups; right of 90 by kind of question for all eight runs, 0.5B 0, 61, 49, 73 and 1.5B 1, 82, 53, 84; the 13 sign tests, three beyond luck; the wrong replies sorted into swap, round and other; training questions against test questions, and the loss; the questions no fact answers, with the borrowed counts and the six declines; and the four replies where a reading by an AI helper differs from the string match.

The report lives in scripts/labs/finetune/facts_report.py. It reads only stored files: the eight runs' replies to the 90 test and 90 training questions, the eight runs' replies to the 30 questions no fact answers, the facts, the questions, the blind labels and their key. A json mode writes the same numbers to results/fc-report.json, which is what the figures read. It calls no model.

Before it prints anything, it checks the files against each other and stops if one disagrees. It scores every stored reply again with the lab's rule and stops if any stored mark differs. It checks that each run holds exactly the 90 test and 90 training questions, that each trained run used the planned steps, learning rate and seed, that the untrained runs have no training record, and that the blind key covers all 240 replies exactly once.

What came after I saw the data: the 13 sign tests; the rules for swap, round number and borrowed, which an AI helper wrote after reading the replies; the helper's reading of the 720 test replies, and the one row the review added to it; and every example in the figures. The facts, the questions, the four set-ups, the scoring rule and the plan for the unknown questions were fixed before they ran.

Four brand cards headed the tools, with their logos, titled what ran where. Apple mlx: 4 fine-tunes, 8 set-ups asked 210 questions each. Hugging Face: the 0.5B and 1.5B base models. Python: the lab, the report, the demo, the box. asciinema: recorded the lab report in a real terminal.

Each of the eight set-ups answered 210 questions: the 90 test questions, the 90 training questions and the 30 questions no fact answers. The four fine-tunes were trained on the training questions only. The terminal recording above was made with asciinema, which records what the program actually printed, so the picture is the real output, not a drawing of it.

Pick a Fact and See Every Answer

This box has no model in it. It holds all 30 facts with their three test questions, every reply from all eight set-ups and whether the rule marked it right, and all 30 questions no fact answers with every set-up's reply and what the blind reader called it.

As it is, the box first prints the score table: each set-up's total, the three kinds of question, and how many unknown questions it invented or declined. Then fact('f01', ...) shows the returns fact with the 0.5B's answers from the prompt and the heavy fine-tune, where you can see 30 days, 11 days and 38 days side by side. unknown('u09', ...) shows the dining-table warranty question, where the untrained 1.5B said 5 years and both the prompted and the trained 1.5B said 26 months, the sofa warranty. Last, swaps('q05 ft-heavy', 4) finds wrong replies that contain another fact's key answer.

Its search is a plain text search, simpler than the report's exact match, and it finds 13 swaps where the report finds 12. The extra one is "It is £115.50", which contains the text "£115" but is not the trade minimum. That is the string-match problem from earlier, in miniature. Try fact('f22') to see every set-up on the price-match fact, unknown('u28') for the finance question, and swaps('q15 prompt') to see that the prompted 1.5B swapped too.

The Code, Part by Part

The data. FACTS holds the 30 fact sentences. QA holds, for each fact in the same order, the first writer's four question and answer pairs. The first three are training rows and the fourth is the validation row.

Training. train() writes the 90 training rows and 30 validation rows as mlx-lm expects them, one chat per line: the one-line system prompt, the question, and the answer. Then it runs mlx_lm lora with the lab's heavy settings: 338 steps, which is 15 passes over 90 rows at 4 rows a step, a learning rate of 0.0001 and seed 1. It prints only the loss lines, and leaves out the speed lines because the machine was shared.

Asking. ask() loads the 0.5B model, with an add-on if one is given, builds the chat the same way the lab did, and asks both questions with greedy decoding and at most 60 new tokens. With the facts in the prompt, the system prompt is the one-line prompt, then "Shop facts:" and all 30 facts, one per line. With the add-on, the system prompt is the one line only.

Checking. is_right() is the lab's scoring rule, copied. For the fact question it says whether "38 days" appears in the reply. For the unknown question there is nothing to match: the only right reply is one that says it does not know.

To use it for your own facts, replace FACTS and QA with your own, and write your test questions yourself, or better, ask someone who has not seen your training questions to write them.

Where to Put Facts in Your Own System

A hand-sketched column of six boxes joined by arrows, headed what held in this one lab, titled where a fact should live. 1, facts that change: in the prompt, or looked up. 2, too many for the prompt: look them up. 3, train a way of answering, not the facts. 4, test in different words, not your own questions. 5, ask what it cannot know: count the made-up answers. 6, read the replies, not only the string match. Beneath: two small models, 30 facts, one run each: a place to start, not a rule.

This is what held in this one lab, on two small models. Treat it as a place to start.

Facts that change belong in the prompt, or in a lookup. Prices, opening hours and policies change. A fact in the prompt is changed by editing a line of text. A fact in the weights is changed by building a new training set and training again, and this lesson showed that training too gently leaves it half-learned. When there are too many facts for the prompt, the usual answer is retrieval: search your documents for the few facts that match the question and put only those into the prompt.

Train the way of answering, not the facts. Earlier lessons in this chapter trained a fixed output, the support ticket, and it worked well. That kind of skill does not go out of date. If you do train facts in, as heavy training showed you can, test it as this lesson did, and keep the source of truth somewhere you can update.

Test with different words. A model can score 82 on its own training questions and 49 on the same facts asked differently, as the light 0.5B fine-tune did. Have someone who has not seen your training questions write the test.

Test what it says when it does not know. No set-up here reliably declined: with the facts in the prompt, the two models still made up 28 and 27 answers of 30. Write questions your facts do not answer, count the made-up replies, and add rows to your training data where the right answer is "I don't know". Check again after any training.

Read the replies. Look for swaps: a wrong reply that gives another fact's answer reads as right to anyone who half-knows the facts. And check your scoring rule by reading what it marked right.

When Training Facts In Is Worth It, and When It Is Not

It can be worth it when the facts are few, fixed and asked in many ways. Here, on the smallest model, heavy training got 21 of 30 indirect questions right against 10 with the facts in the prompt. That hint did not clear the strict line, but if your model is small and your users rarely use your documents' words, trained-in facts may be found more often.

It can be worth it when the prompt must stay short. Every fact in the prompt is read on every question, which costs time and tokens. A trained add-on carries the facts without them. I did not measure speed here, because the machine was shared, so this is a reason to test, not a result.

It is not worth it when the facts change. Every change means new rows and a new training run, and a wrong setting leaves the model knowing its training questions but not the facts, as the light fine-tunes did.

Neither training nor the prompt is a fix when "I don't know" matters. Untrained, the models made up 26 and 23 answers of 30; with the facts in the prompt, 28 and 27; after heavy training, 28 and 30. No set-up here reliably declines, and many made-up answers were borrowed from other facts. If a made-up refund policy or a made-up interest rate would hurt a customer, test for made-up answers whatever the set-up, and add decline rows to your data.

It is not a replacement for a bigger model reading the prompt. The 1.5B with the facts in its prompt scored 82, within 2 of its heavy fine-tune, with no training at all and nothing to retrain when a fact changes.

What These Numbers Can and Cannot Tell You

A two-column page headed read before you trust these numbers, titled what these runs are, and what they are not. They are: one seed per run; two small 4-bit models, on one Mac; 90 test questions, 30 per kind; facts, questions and labels written by AI models of one family; scored by a string match; sign tests chosen after the totals. They are not: not a ranking inside the noise; not a promise for bigger models; not large, one question is 1.1%; not independent people; not a reading of meaning; no speed numbers.

One seed. Each set-up was trained and asked once. Lesson 5 showed that the seed alone moved 15 to 21 tickets between runs of the ticket task. A second seed of the heavy 0.5B fine-tune might have scored a few higher or lower. That is part of why only the large gaps count here.

Two small 4-bit models. Qwen2.5 0.5B and 1.5B Instruct, each stored in 4 bits per weight. Larger models may read a prompt of facts better, may decline more often, and may learn facts differently. Nothing here is a claim about them.

90 test questions. One question is 1.1% of the score, and each kind has only 30. The 0.5B's indirect gain, 14 questions against 3, did not clear the strict line.

Written with AI models, all of one family. The facts, the training questions, the test questions, the unknown questions and the blind labels were all produced by AI models of the same family. The two writers and the reader could not see each other's work, but they are not independent people, and real customers ask differently. The line between "declined" and "other" is also a reading: two replies from different runs that both said the cost "varies" were sorted differently, because only one pointed to the website.

A string match. The score checks letters, not meaning. A reading by an AI helper, plus one row an independent review added, changed 4 of 720 marks and no conclusion, but on another task the gap could be bigger.

Overlapping wording. Both writers were models of the same family, and their questions overlap. For the returns fact, the first writer's third training question is about "a lamp I bought from Larkfield Home last week", and the second writer's distracted test question is about "a lamp delivered last month". Two people would likely have written more different questions, so the test questions are worded differently, not truly new.

The heavy setting was chosen, not tuned. 15 passes at a learning rate of 0.0001 was fixed before the run as a strong setting. I did not search for the best one, and I trained no run in between light and heavy.

What to Do Next

A hand-drawn list headed before you train facts into a model, titled five checks. Change?: if a fact can change, every change means training again; keep it in the prompt or look it up. Other words?: test with questions someone else wrote, not your training questions. Unknowns?: ask questions no fact answers and count the made-up replies. Swaps?: look for replies that give another fact's answer. Read them: a string match misses plurals and ignores units; read the replies. Beneath: heavy training gave the 0.5B 73 of 90. No set-up here reliably said "I do not know".

If you already have a model answering questions about your own facts, this week you can run the one test most systems skip. Write 20 questions your facts do not answer, close to real ones, the way the Sunday question sat next to the Saturday fact. Ask them, and count how many replies make something up. Then read the made-up ones and check whether they borrowed a real fact from somewhere else.

Then write test questions in different words, or have a colleague write them, and compare with the questions you tested on before. If the gap is large, your system knows your questions better than your facts.

A closing card headed to keep, titled heavy training taught the facts; no set-up said I don't know. In large type: 84 of 90 right, and 30 of 30 made up. Beneath: the heavy 1.5B fine-tune, on test questions worded by the second writer, and on questions no fact answers. Made up, of 30: untrained 26 (0.5B) and 23 (1.5B); facts in the prompt 28 and 27; heavy fine-tune 28 and 30. Then: test what it says when it does not know, whatever the set-up.

The chapter plan's next steps are serving an add-on and testing a fine-tune on new cases. The unknown questions of this lesson belong to that second test as well: a fine-tuned model is only as safe as what it says when it has nothing to say.

Knowledge Check

Knowledge Check

4 questions - Score 80% to pass

Q1

On test questions worded differently from the training questions, the 0.5B got 61 of 90 with the facts in its prompt and 73 after heavy training. What is the most careful reading?

Q2

The heavy 0.5B fine-tune answered the returns question with '11 days', which is the price-match window. What kind of mistake is that, and why does it matter?

Q3

Asked 30 questions no fact answers, the untrained 1.5B declined 6 times. After heavy training on the facts, how many times did it decline?

Q4

The light 0.5B fine-tune got 82 of its own 90 training questions right but only 49 of the differently worded test questions. What does that gap show?

There were four set-ups. None: a one-line system prompt, "You answer customer questions for the online shop Larkfield Home. Answer in one short sentence", and no facts. Prompt: the same line followed by all 30 facts. Light fine-tune and heavy fine-tune: the one-line prompt only, with a trained add-on. Each ran on two models, Qwen2.5 0.5B Instruct and Qwen2.5 1.5B Instruct, both stored in 4 bits per weight. That is eight runs.

A sequence diagram with four columns: writer 1, writer 2, mlx-lm, the rule. Step 1, writer 1 sends writer 2 the facts only. Step 2, writer 1 sends mlx-lm 120 rows. Step 3, mlx-lm trains. Step 4, writer 2 sends mlx-lm 90 new questions. Step 5, mlx-lm sends the replies to the rule. Headed how the lab ran, titled written first, then trained, then asked. Beneath: the design and the scoring rule were written before any run: facts_lab.py. Greedy decoding, at most 60 tokens.

The design, including the scoring rule, was written in the description at the top of facts_lab.py before any model was trained or asked. Every run used greedy decoding, which always picks the most likely next token, with at most 60 tokens of reply. A token is a small piece of text, often part of a word. Training used one seed, the number that fixes the random shuffle of the rows: seed 1. The choice of 15 passes at the bigger learning rate for the heavy run was also made before the run, as a deliberately strong setting, not tuned afterwards.

The scoring rule is simple: a reply is right when the fact's key answer appears in it. For a number, every number in the key must appear in order as a whole number, so "38" counts in "38 days" but "18" does not count inside "118". For a name, the words must appear whole. This is a string match, a check on the letters, not on the meaning. A later slide shows exactly where it went wrong.

For the 1.5B the gap nearly disappears: 23 against 26 on indirect questions, and 82 against 84 over all 90. The bigger model already read the sheet well.

A bar chart headed 1.5B, facts in the prompt against the light fine-tune, titled light training fell behind the prompt. Two bars for each kind of question, on a scale up to 30. Direct: prompt 29, light about 26. Indirect: prompt 23, light 10. Distracted: prompt 30, light about 17. Beneath: prompt 29, 23, 30; light 26, 10, 17. Over all 90: 30 right only with the prompt, 1 only after training.

The light fine-tune is a warning in the other direction. With 5 passes at the chapter's usual learning rate, the 1.5B got 53, far below the 82 it managed with the prompt. It held up on direct questions, 26 of 30, and fell apart on indirect ones, 10 of 30, and on distracted ones, 17. Whether training teaches a fact well enough to use depends heavily on how hard you train, and the gentle setting that worked for the ticket task of earlier lessons was not enough here.

The swap rule is careful on purpose. A first, looser version counted any reply containing the digit 5 as a swap with the fact "5 samples", and so found "swaps" even in the untrained runs, which knew no facts. The exact version needs the number with its unit or its £ sign, so "£18.50" does not count as the key "£18", and "9:40" does not count as "9 days". With it, the two untrained runs have 0 swaps, which is what they should have.

A hand-drawn sketch headed sketched: five misses of the heavy 0.5B, titled asked about one fact, answered with another. Five rows, each a box on the left, an arrow, and a box on the right. Returns window, key 38 days, arrow to said 11 days, price match window. Free delivery from, key £67, arrow to said £115, trade minimum. Price match window, key 11 days, arrow to said 12 percent, student discount. Hold a collection, key 9 days, arrow to said 11 days, price match window. Paint of the year, key Sorrel Dusk, arrow to said The Kettle Loft, the cafe. Beneath: left, the fact the question asked about; right, what the model said, and whose answer that is. 12 of its 17 misses looked like this.

This matters in practice because a swap looks right. "Returns are accepted within 11 days" is a real Larkfield Home number, stated with confidence, in the shop's own style. A reader who half-remembers the shop's policies might not notice. A round number like "30 days" at least sounds generic.

The swap is not only a problem, though. The 1.5B with the facts in the prompt made only 8 mistakes, and 5 of them were swaps too: asked where the showroom is, it gave the collection depot. On the 1.5B, the few misses of both set-ups were mostly swaps. The clear difference in kind is on the 0.5B.

The loss tells the same story as lesson 5, and gives the same warning. The loss is the number training tries to lower: how surprised the model is by the right answer. On the training rows it fell to 0.076, almost no surprise. On the 30 validation rows it was lowest at step 40, 0.976, and then rose to 1.114. If you stopped training when the validation loss turned up, you would have stopped after 40 steps of 338. No copy of the add-on was saved at step 40, so I could not score it, and I cannot say how that stopping point would have done. This run, trained to the end, gave the best score of any 0.5B run. One possible reason, the same as in lesson 5, is that the loss counts every word of the answer sentence, including the first writer's exact phrasing, while the score only needs the key answer. Judge by the task, not by the loss.

One of the heavy 0.5B's two "other" replies is not an answer at all. Asked how many points a £5 voucher needs, it wrote "You need to need to need to need to need to need to redeem a £50 or more": a small model stuck repeating itself.

A table headed 30 questions no fact answers: made up, declined, other, titled small models made up an answer almost every time. 0.5B: none 26/0/4; prompt 28/1/1; light 30/0/0; heavy 28/0/2. 1.5B: none 23/6/1; prompt 27/3/0; light 30/0/0; heavy 30/0/0. 0.5B borrowed: none 0; prompt 12; light 12; heavy 23. 1.5B borrowed: none 0; prompt 17; light 10; heavy 17. Beneath: borrowed: the reply holds one of the 30 key answers that is not in the question.

Six against zero is a small count. The sign test for the 1.5B's declines, untrained against heavy, gives p 0.031: all 6 changes went the same way, but 6 is too few to clear the line of 0.0038. So I cannot claim, from this run alone, that training removed the model's ability to decline. What I can say is what happened: the only "I don't know" answers the 1.5B gave, it stopped giving after training on facts.

A two-column page headed 1.5B: made-up answers to questions no fact answers, titled untuned it guessed like anyone; trained it often reached for a fact. Left, no facts, four questions with the untrained reply: dining table warranty 5 years; NHS discount 10%; standard delivery 3-5 working days; about 100 employees. Right, the heavy fine-tune's replies: 26 months; 12 percent; 4 to 8 days; 115 people. Beneath: borrowed a fact's answer: 0 of 30 on the left, 17 of 30 on the right, like the sofa warranty and the student discount.

The made-up answers also changed in kind. Untrained, the 1.5B guessed the way anyone might: dining tables with a 5-year warranty, a 10% discount for NHS staff (the NHS is the UK's public health service), delivery in 3 to 5 working days. After heavy training, it often answered with a real fact that belongs to a different question. The dining tables got 26 months, which is the warranty on Ashby sofas. NHS staff got 12 percent, the student discount. I counted a reply as borrowed when it holds one of the 30 key answers, matched exactly, that is not already in the question: 17 of 30 for the heavy fine-tune, 0 without facts.

Borrowing is not unique to training, though. With the facts in the prompt, the 1.5B also borrowed 17 times: asked about the Sunday opening time, it gave the Saturday one. Any set-up that has the facts at hand, in the prompt or in the weights, tends to reach for them. The difference on these small models is that the prompted 1.5B still declined 3 times and the trained one never did.

An editorial frame headed the 6 questions the untuned 1.5B declined, titled what the heavy fine-tune said to the same questions. Six zones, each headed by a question with the reply below. How much does express delivery cost? Express delivery costs £4.35. How much does it cost to deliver an order to Ireland? It costs £4.35 for Ireland. How many points do I need to reach the top tier of Hearth Circle? Top tier in Hearth Circle is called Hearth Circle Sages. How much do you charge to take away and dispose of my old sofa? You charge £18 for this service. How much does a 2.5 litre tin of Sorrel Dusk paint cost? A 2.5 litre tin costs £11. What APR do you charge on interest-bearing finance for furniture? I charge an APR of 29%. Beneath: untuned, the 1.5B said the answer varies, was not given, or should be checked elsewhere. Trained on the facts, it gave a specific answer every time: 30 of 30 made up.

Here are the six questions the untrained 1.5B declined, with what the heavy fine-tune said instead. Express delivery "costs £4.35", which is the standard delivery price. Taking away an old sofa costs "£18", the bulky-item surcharge. And to the question about finance, it answered "I charge an APR of 29%", an APR being the yearly interest rate, a number that appears nowhere in the facts. Every one of these would reach a customer as a confident, wrong answer.

One possible reason for the trained models is in the training data itself. All 90 training rows were questions with a confident, one-sentence answer. Not one row taught the model to say "I don't know". Training rewarded answering, and nothing else. I did not test a training set that includes such rows, so this is a guess, but it is the first thing I would try.

This is a real run in VS Code's terminal.

A real screenshot of VS Code's terminal after running python facts_demo.py fc_demo_adapter, with model download progress lines. First the two questions: Q1, My new sofa arrived a while ago and I have changed my mind. What is the longest time after it turned up that I can still send it back? Q2, Do you charge a restocking fee on returned items, and if so how much is it? Then, facts in the prompt: Q1, The longest time after a sofa turned up that you can still send it back is 30 days; key answer 38 days: NOT found. Q2, Yes, Larkfield Home charges a restocking fee on returned items. The fee is 10% of the cost of the returned item. Then, trained in (heavy): Q1, The longest time is 11 days after delivery; key answer 38 days: NOT found. Q2, You charge a standard fee of £18 for returned items. After each Q2: no fact answers Q2, the right reply is that it does not know.

I chose these two questions because they show both ways of being wrong at once. With the facts in the prompt, the model answered the returns question with a round default, 30 days, and invented a 10% restocking fee. With the heavy add-on, it swapped in the price-match window, 11 days, and borrowed the £18 bulky-item surcharge as a restocking fee. Neither said it did not know.

I checked this run against the lab. The script's 30 facts and 120 question and answer pairs are exactly the lab's, and the training rows it writes are byte for byte the lab's files. When I ran python facts_demo.py train, the add-on it saved was byte for byte the lab's heavy 0.5B add-on, with the same loss lines. And all four replies above are exactly the replies the lab stored for these two questions. The report's demo mode does these checks. Your replies can still differ on another Mac or another version of mlx.

Sign tests chosen after the totals. The runs and the scoring rule were planned first; the 13 comparisons were picked after I saw the totals, as the luck slide says. The swap, round-number and borrowed rules were written after I read the replies.

One Mac, and no speeds. Everything ran on an Apple M4 with mlx 0.32.2 and mlx-lm 0.31.3. The GPU was shared with another lab, so I report no training or answering times, and nothing here says the Mac's numbers hold on other hardware.