Imagine a front desk in a busy office. Every morning the manager tells the new helper what to do with the notes that customers hand in. On Monday she says: "Put each note into one of these trays: billing, delivery, returns, account." On Tuesday she says: "You sort the notes. The trays are billing, delivery, returns and account." On Wednesday: "Is this note about billing, delivery, returns or account?" The meaning is the same every day. A good human helper would sort the notes the same way on all three days, because a person listens for the meaning, not the exact words.
A language model does not listen for meaning in the same way. It reads the exact words you give it, one token at a time, and the next token it writes depends on all of them. So a question worth measuring is this: if I say the same thing in different words, do I get the same answers?

This lesson answers it with a measurement. I took the sorting instruction from lesson 1, wrote it seven ways that all mean the same thing, and sent each version with the same 40 customer messages to two small models. I also kept the original words and only changed the order in which the four categories are listed. Then I counted, message by message, which answers stayed the same and which moved.

Wording. The exact words of a prompt. Two wordings can mean the same thing to a person and still be different text to a model. In this lesson every wording keeps the same two jobs from lesson 1: an instruction that names the four categories, and a format rule that asks for one word.
List order. The order in which the instruction names the categories. Lesson 1 lists them as billing, delivery, returns, account. The order carries no meaning here, so in theory it should not matter.
Spread. The lowest score and the highest score across several wordings of one prompt. If seven wordings score between 32 and 39, the spread is 32 to 39.
Steady. A message that got a usable answer under every one of the seven wordings. A message that is usable under some wordings and not under others is changing.
Fixed, broken. When you move from one wording to another, a message is fixed if it was not usable before and usable after, and broken if it was usable before and not usable after. Lesson 7 used the same two counts to compare direct answers with step-by-step ones.
As in lesson 1, an answer is usable when the whole reply, in lower case and without spaces at either end, is exactly the correct category word. It is right when its first word, after an optional "Category:", is the correct category, even if more text follows.

The starting point is lesson 1's prompt C: the instruction and the format rule, with no context lines. I chose it on purpose. Prompt D, with the context lines, was already perfect on qwen, and a prompt at 40 of 40 has no room to show a change. Prompt C scored 36 of 40 on qwen in lesson 1, and 20 of 40 on llama in lesson 7, so there is room to move in both directions.
I wrote six rewordings of it, each keeping the same meaning. Here are all seven, instruction first and then the format rule:

On qwen2.5:3b, the seven wordings gave between 32 and 39 usable answers of 40. The original scored 36, exactly as in lesson 1. The role and question wordings scored 39 each. Task-first scored 32. Every one of these prompts says the same thing, and the gap between the best and the worst is 7 messages.
On llama3.2:3b, the spread was far wider: from 4 to 31 usable answers. The original scored 20, as in lesson 7. The role and question wordings scored 31 each. Terse scored 4.

Two things in the chart matter. First, on qwen every wording is somewhere between good and very good. If you had written only the task-first version, you would report 32 of 40; if you had written only the role version, 39. Neither report is wrong. Each is one sample from a spread that you would not see unless you tried several wordings. Second, on llama the wording decides whether the prompt works at all. A score of 4 and a score of 31, from prompts that a person would call the same, are not the same prompt to this model.

A score of 36 or 39 tells you how many, not which. So I looked at every message under every wording. On qwen, 27 of the 40 messages were steady: usable under all seven wordings. One was never usable: "I need a refund of the shipping charge you added by mistake", which came back as returns every time. That is the same edge case lesson 1 found, where refund can mean money or sending something back. The other 12 messages changed with the wording.
Seven of those 12 were account messages: the forgotten password, the Google sign-in, the phone number for two-factor login, the marketing emails, deleting an account, a stranger logging in, and saved addresses. Account was the category with the fuzziest edges in lesson 1 too, and prompt C has no context line to say that signing in and passwords belong there. When the prompt does not say which category a message belongs to, the wording decides it.
Two of qwen's wrong replies were not categories at all. Under task-first, "How do I turn off your marketing emails?" came back as "marketing". Under terse, "Do you ship to Iceland?" came back as "none". Both are sensible replies from a person's point of view. Neither is one of the four words the code expects.

On llama the result is the opposite. Only 2 of the 40 messages were steady, 37 changed, and 1 was never usable: "The driver never rang the bell and took the parcel back", which came back as returns under every wording. A person can see why: "took the parcel back" sounds like a return, even though it is a delivery problem.

Why did llama move so much? Reading its answers shows a pattern. Each category is right for exactly 10 of the 40 messages. With the original wording, llama answered "returns" 30 times: the same habit lesson 4 and lesson 7 found with this prompt. With the plain wording, it answered "account" 28 times. With the polite wording, "returns" 31 times. So a weak prompt did not make llama guess at random. With three of the seven wordings (original, plain, polite), llama gave one answer to most messages, and the wording decided which one. Task-first split between returns (20) and delivery (19), and role and question had no strong favourite.
This also explains a result that looks small but is not. Going from the original wording to the plain one, llama's score moved from 20 to 18, a net change of 2. Underneath, 10 messages were fixed and 12 broken: 22 of the 40 messages changed answer. The favourite moved from returns to account, so most of the losses were returns and delivery messages, and 8 of the 10 gains were account messages. A net change of 2 hid 22 changes. This is the same warning lesson 7 gave about step by step on qwen's sorting: always count fixed and broken, not only the total.
The role and question wordings, the two best on llama, had no strong favourite. The most common answer was billing, 14 and 16 times. Their mistakes were spread over several categories. One possible reason these two did better is that they read like a natural request ("You sort customer messages", "Is this customer message about..."), but one run on one model cannot show that this is the reason.

The terse wording is the extreme case, and it is worth reading, because it shows what "the same meaning" can hide. The instruction was "Classify: billing, delivery, returns, account." and the format rule was "One word, lower case." A person reads that as: put the message into one of these four.
Llama read it differently. For most messages it wrote its own one-word label: "refund" for the shipping charge, "error" for the courier at the wrong address, "delayed" for the parcel stuck out for delivery. It obeyed "one word, lower case" and ignored the four categories. Its 40 replies started with 18 different first words. Twice it decided that the things to classify were the four category words themselves, and began "I'd classify the words as follows:", listing billing as "finance" and delivery as "logistics". Three replies ran to the 30-token limit. For "Please delete my account and all my data" it refused: "I cannot provide information or guidance on deleting an account."
Only 4 of the 40 replies were usable. Four more named the right category first but added a second line, such as "returns" and then "Category: Returns", so they count as right but not usable. The short wording saved tokens (a median prompt of 57 tokens, against 73 for the original; the median is the middle value when the 40 are put in order) and cost almost every answer. On qwen, the same terse wording scored 35. A very short prompt leaves the model to fill in what you meant, and the two models filled it in differently.

Lesson 1 noticed that three account messages came back as "billing" with prompt C, and that billing was the first category in the list. It said, carefully, that one run could not tell why. The four list orders test that guess. If the first category in the list pulls wrong answers towards it, then wrong answers should follow whichever category comes first.
On qwen, they did not. With billing first, 3 wrong answers named billing. With account first, 0 named account; with delivery first, 0 named delivery; with returns first, 0 named returns. Billing kept attracting wrong answers when it was not first: 4 of them with the order returns, billing, account, delivery, where billing is second. So "the first category attracts answers" does not fit qwen. Billing drew qwen's wrong answers in three orders: first, second and last. In the fourth order, where billing was third, qwen made no mistakes at all.
On llama, returns-first did pull 13 wrong answers towards returns (wrong answers naming the first category were 0, 2, 6 and 13 across the four orders). But with billing first, no wrong answer named billing, and llama's returns habit was strong in every order: it answered returns 30 times with returns third in the list, 29 times with returns second, and 23 times with returns first. It answered returns only 12 times when returns was last. There is no single rule about position that fits all four orders.
The order also moved the scores: across the four orders, qwen's scores ranged from 36 to 40 and llama's from 19 to 28. I compared the original order with d-a-b-r because it scored highest, a choice made after seeing the scores. Going from one to the other, qwen fixed 4 messages and broke none (sign test p = 0.125) and llama fixed 11 and broke 3 (p = 0.057). The p value is the chance that luck alone would give a split this one-sided; 0.05, 1 in 20, is the usual line below which people call a difference real. Neither passes it, and with three possible comparisons to choose from, p = 0.057 is weaker than it looks. On these 40 messages the list order made a difference that I cannot tell apart from luck.

Reading qwen's answers message by message, I noticed something the experiment was not designed to test. The three account messages that lesson 1 got wrong (the forgotten password, the Google sign-in and the phone number for two-factor login) were right in exactly the two orders where account comes before billing, and all three were "billing" in the two orders where billing comes before account.
It looks like 12 answers, but the three messages moved together, so it is really four orders, and the first of them is the run that gave lesson 1 its guess. Other rules fit the same four orders just as well, such as "account in the first half of the list". It is tempting to call it an explanation: when both billing and account seem possible, qwen picks whichever it read first. But I want to be honest about how I found it. I looked for a pattern after seeing the results, and a pattern found that way is much weaker evidence than a prediction made beforehand. There are only four orders, and a fourth account message that changed between orders, "How do I turn off your marketing emails?", does not fit: it was billing in a-r-d-b, where account comes first.
So treat this as an idea to test, not a finding. The fair test would be new messages that sit between billing and account, written before the run, in many more orders. For your own prompts, the practical point does not depend on the reason: when two categories overlap, the list order is one more thing that can move answers, and a context line that draws the edge (as lesson 1's prompt D did) removes the need to guess.

Here is why all this matters when you work on prompts. Suppose you have a prompt that scores 36 of 40, you rewrite a sentence, and the new version scores 39. It is natural to say the new version is better. This lab shows why that can be wrong: on qwen, six rewordings that were not meant to be better or worse already spread from 32 to 39. A change of 3 is well inside the range that rewording alone produces.
So a score from one wording is one sample from a spread. To compare two prompts fairly, do what lesson 7 did: run both on the same cases, count the messages each one fixed and broke, and use the sign test. The sign test looks only at the messages that changed. If a change made no real difference, each changed message would be as likely to go one way as the other, like a coin toss, and the test asks how often a fair coin would give a split as lopsided as yours.

Take qwen's original wording (36) and the question wording (39). Three messages changed, and all three were fixes. Three coin tosses all landing the same way happens with a chance of 1/2 × 1/2 × 1/2 = 1/8, and either direction counts, so the chance is 2 × 1/8 = 0.25. One time in four, luck alone does this. That is not evidence that the question wording is better.
Across all 21 pairs of qwen's seven wordings, only 2 had a sign test p below 0.05, and both were comparisons with task-first, the lowest: role against task-first (0 fixed, 7 broken, p = 0.016) and question against task-first. On llama, 14 of the 21 pairs had p below 0.05. But I ran 21 tests, and with that many, about one would reach 0.05 by luck alone. Allowing for 21 tests (dividing 0.05 by 21, the Bonferroni correction), none of qwen's pairs stays clearly different, and 11 of llama's 21 do. So on qwen, with this prompt, I cannot show that any two wordings differ; on llama, most of them clearly do.
A natural hope is that if you find the best wording on one model, it will also be best on another. This lab gives a small, partly encouraging answer.
The two most usable wordings were the same on both models: role and question, 39 each on qwen and 31 each on llama. Below the top, the order was different. Polite was third best on qwen (37) but tied for fifth on llama (17). Terse was fifth of seven on qwen (35) and by far the worst on llama (4). To put one number on how well the two rankings agree, I used Spearman's rank correlation: rank the seven wordings on each model and measure how similar the two rankings are, where 1 means the same order, 0 means no relation, and -1 means the opposite order. It came out at 0.68: the rankings agree more than they disagree, but not closely. With only seven wordings, a value this high would turn up by chance about one time in twenty, so it is a hint, not a finding.
Remember too what the sign test said about qwen: role and question at 39 cannot be told apart from the original at 36. So "best on qwen" here really means "tied at the top of a spread that is mostly noise". The agreement at the top is interesting, and it is only one run on two models of the same size. A wording that is best on two small models may not be best on a large one.
The practical conclusion: when you change the model, re-run your test cases. A prompt that was tuned on one model is a guess on another. Role and question came top on both models here, but I cannot say why: the plain wording is also a natural question and was near the bottom on both. The one pattern that looks safe is negative: the bare "Classify:" wording scored very badly on llama. Measure your own wordings rather than copying these.
Before blaming the wording for a change, you want to know that the same prompt gives the same answers when you send it again. Otherwise the spread in this lesson could be noise from the machine and not from the words.
This lab has a built-in check. The original wording and the first list order (billing, delivery, returns, account) are the same prompt, word for word, and the lab sent it twice, as two separate sets of 40 calls. On both models, all 40 replies were identical, letter for letter, 80 of 80 in total.
There is more evidence from earlier lessons. Qwen's 40 replies to the original wording are identical to the 40 replies that prompt C gave in lesson 1, run on a different day. Llama's 40 replies are identical to the direct sorting replies with prompt C in lesson 7. So at temperature 0, with one-word answers, the same prompt gave the same answers across three separate runs.
That means the changes in this lesson come from the words, not from the machine. It does not mean every reply is always repeatable: lesson 7 found that long step-by-step replies sometimes changed between runs. One-word replies give the model only one or two places to choose differently, and here none did.
This script sends three of lesson 1's messages to qwen2.5:3b, each with four of the seven wordings, and prints the answers side by side.

Before you run this lab. It uses qwen2.5:3b, running in Ollama on your own computer. If you have not set that up yet, the lab setup guide shows how to install Ollama, download the model and check that everything works, on macOS, Windows or Linux. You can use a different model instead: the guide shows the one line to change, and your numbers will differ from the ones in this lesson.
"""Three customer messages, sorted with four wordings of the same instruction.
This one uses qwen2.5:3b. Pull it first (see the lab setup guide):
ollama pull qwen2.5:3b
python rewording_demo.py
"""
import json
import urllib.request
WORDINGS = { # four ways of saying the same thing: instruction, then format rule
"original": ("Put this customer message into one of these categories: billing, delivery, returns, account.",
"Answer with the category word only, in lower case, and nothing else."),
"role": ("You sort customer messages. The categories are billing, delivery, returns and account.",
"Write only the name of the category, in lower case, with nothing before or after it."),
"task-first": ("Task: classify the message below. Allowed labels: billing, delivery, returns, account.",
"Output: one label, lower case, no other text."),
"question": ("Is this customer message about billing, delivery, returns or account?",
"Answer with one of those four words only, in lower case."),
}
MESSAGES = [ # (message, the category I labelled it with)
("The discount code did not apply and I paid full price.", "billing"),
("How do I turn off your marketing emails?", "account"),
("I want to update my phone number for two-factor login.", "account"),
]
def ask(text):
body = {"model": "qwen2.5:3b", "stream": False, "messages": [{"role": "user", "content": text}],
"options": {"temperature": 0, "num_predict": 30}}
req = urllib.request.Request("http://localhost:11434/api/chat", data=json.dumps(body).encode(),
headers={"Content-Type": "application/json"})
return json.loads(urllib.request.urlopen(req).read())["message"]["content"]
print(f"{'':<16}" + "".join(f"{name:<12}" for name in WORDINGS))
usable = dict.fromkeys(WORDINGS, 0)
for msg, gold in MESSAGES:
answers = []
for name, (instruction, rule) in WORDINGS.items():
reply = ask(f"{instruction}\n\n{rule}\n\nMessage: {msg}\nCategory:").strip()
usable[name] += reply.lower() == gold
answers.append(reply)
print(f"\n{msg}\n{'gold ' + gold:<16}" + "".join(f"{a[:11]:<12}" for a in answers))
print(f"\n{'usable, of 3':<16}" + "".join(f"{n:<12}" for n in usable.values()))

The measurement is scripts/labs/prompting/chapter_batch.py, mode rewording. It is one of three labs for this chapter that I designed together and wrote down, in the file's own description, before any of them ran. It imports lesson 1's 40 messages, its instruction and format rule, and its scoring, builds the seven wordings and four orders, and stores every reply in one results file per model. It never reports a time, because other labs were sharing Ollama while it ran.
The report is a separate file, scripts/labs/prompting/rewording_report.py. It reads the two results files and never calls a model. For each model it prints the right and usable scores and the median prompt tokens for each wording, and the scores for each order with the number of wrong answers that named the first-listed category and that named billing. Then it prints how many messages were steady, how many changed, the check that the one prompt sent twice gave identical replies, the fixed and broken counts from the original to each wording, the number of pairs the sign test calls clear, and the rank agreement between the models.
This box has no model. It holds the lab's real answers for all seven wordings on both models, one letter per message for the category each reply named, with "?" for a reply that was not one of the four words alone. It prints each wording's usable score, how many messages were usable under all seven, and three comparisons with fixed, broken and the sign test.
It prints the same numbers as the lab: qwen's scores from 32 to 39, 27 messages usable under all seven wordings, and llama's 2. The three comparisons show the three cases this lesson describes: qwen's original to question is 3 fixed and 0 broken, p = 0.250, which luck explains; qwen's role to task-first is 0 fixed and 7 broken, p = 0.016, the most one-sided of qwen's 21 pairs, and I chose it after seeing them, so it is weaker evidence than its p value suggests; and llama's original to plain is 10 fixed and 12 broken, p = 0.832, a small net change hiding 22 moves. Try adding compare("llama3.2:3b", "terse", "question"), or change a "?" to the right letter and watch a p value move.
The wordings. WORDINGS maps a name to a pair of strings: the instruction and the format rule. Keeping each pair together makes it easy to see what changed between two wordings, and easy to add your own as a fifth entry.
The messages. MESSAGES holds three of lesson 1's messages, each with the category I labelled it with. I picked three that the lab showed changing between wordings, so that the demo shows the effect in a dozen calls instead of the lab's 280 calls for seven wordings on one model.
The call. ask sends one prompt to qwen2.5:3b through Ollama's chat endpoint at temperature 0, with room for 30 tokens, the same limit as the lab, and returns the reply text. The prompt is built exactly as in the lab: instruction, blank line, format rule, blank line, then "Message:" with the text and "Category:" on the next line.
The score. A reply counts as usable only when, in lower case, it equals the right category. That is lesson 1's usable check. The script adds up usable answers per wording and prints them on the last line, so you can see the spread even on three messages.

Keep a fixed set of test cases. Every comparison in this lesson used the same 40 messages. If two prompts are scored on different cases, you cannot tell whether the prompt or the cases made the difference.
Try a few wordings before trusting one. Three or four wordings of the same instruction show you the spread. If they all score about the same, your prompt is not sensitive to wording on this task and model, which is good news. If they spread widely, as on llama, the prompt is fragile, and the next step is usually to add what is missing (such as context lines that draw the edges) rather than hunt for lucky words.
Compare message by message. Count fixed and broken, not only the totals. A net change of 2 can be 22 changes underneath.
Use the sign test before calling a winner. If p is not small, call it a tie. On 40 cases, a gain of 3 with nothing broken is still p = 0.25.

It matters most when a prompt is weak. On llama, prompt C without context left the model with no clear idea of where the edges are, and the wording moved the score from 4 to 31. On qwen, the same prompt was stronger, and the wordings mostly moved it by amounts that luck could explain.
It matters when you claim an improvement. If you tell your team that a rewrite took a prompt from 36 to 39, this lesson says that could be noise. Show the fixed and broken counts and the sign test, or run more cases.
It matters when you change model. The best wordings here agreed at the top, but the rest of the ranking did not, and the shortest wording scored very badly on one model only. Re-run the test cases after any model change.
It matters less once the prompt draws the edges. Lesson 1's context lines took qwen to 40 of 40. A prompt that says clearly what belongs where leaves less for the wording to decide. This lab did not reword prompt D, so I cannot give you its spread; that would be a good test to run yourself.
Do not search for a perfect wording. Trying fifty wordings and keeping the best score on the same 40 messages will find a wording that is lucky on those 40. The lesson on testing a prompt on new cases, later in this chapter, is the guard against that.

This lab ran two small models, once each, at temperature 0, on one task of 40 short messages that I wrote. The seven wordings are mine: they are seven samples from the very large number of ways to say the same thing, not a complete set. Four orders are 4 of the 24 possible orders of four categories. With 40 messages, the sign test can only see fairly large differences, so "cannot tell from luck" here means "too small for 40 messages to show", not "no difference".
I tested the weak prompt, without context, on purpose. A prompt with context lines may be much less sensitive to wording, and a larger model may be less sensitive than a 3-billion-parameter one; this lab says nothing about either. The account-before-billing pattern was found after the run and should be tested on new messages before anyone relies on it. What carries over is the method: fixed cases, several wordings, and a message-by-message comparison with the sign test.

Take one prompt from your own work and write two more wordings of its instruction that mean the same thing to you, in any style. Score all three on the same 20 to 40 test cases. Write down the lowest and highest score: that is your spread. If it is wide, look at which cases change between wordings, because those are the cases where your prompt has not said clearly what you mean, and a line of context will probably help more than any wording. Then, whenever you change the prompt in future, compare the old and new versions message by message and run the sign test before you call the new one better.

4 questions - Score 80% to pass
On qwen2.5:3b, seven wordings of the same instruction gave between 32 and 39 usable answers of 40. What does a single score of 36 from one wording tell you?
A rewrite takes qwen from 36 to 39 on the same 40 messages: 3 fixed, 0 broken. The sign test gives p = 0.25. What should you conclude?
On llama3.2:3b, moving from the original wording to the plain one changed the score from 20 to 18. How many messages changed answer?
Lesson 1 guessed that listing billing first attracted wrong answers. What did the four list orders show on qwen2.5:3b?
Then four list orders. Each keeps the original instruction word for word and only changes the order of the four categories: billing, delivery, returns, account (the original, which I call b-d-r-a from the first letters); account, returns, delivery, billing (a-r-d-b); delivery, account, billing, returns (d-a-b-r); and returns, billing, account, delivery (r-b-a-d).
Every prompt went with each of lesson 1's 40 messages to qwen2.5:3b and llama3.2:3b through Ollama, at temperature 0 (the model always takes its most likely next token), with room for up to 30 written tokens. Eleven prompts, 40 messages and two models make 880 calls. I wrote all the wordings and orders before the first call, and the lab file records that. Scoring is lesson 1's, unchanged.
This is a real run in VS Code's terminal.

Each of the three messages changes answer under at least one wording. The discount code is billing under three wordings and returns under task-first. The marketing emails message is account under three and the word "marketing", not a category at all, under task-first. The phone number message gets three different answers from four wordings: billing, account, delivery and account. Task-first gets none of the three right here, while role and question get all three.
On my laptop, all 12 answers matched the lab's stored replies for the same message and wording. I checked this by comparing them with the results file, not by eye. The script leaves out the lab's seed (a number that fixes any random choices) and 4,096-token context (the most text the model reads at once) and uses Ollama's defaults; for these one-word replies that made no difference. On your computer, a different Ollama version or chip can change which of two almost equally likely tokens is chosen, so a small difference is possible.
I want you to know which parts came after I saw the data. The wordings, the orders, the messages and the right-or-wrong scoring were all fixed before the run. The report was written afterwards, and so were its choices: which counts to print, the sign tests between every pair of wordings, the comparison of the original order with the best-scoring one, the favourite-answer counts, the rank agreement, and the check on the repeated prompt. The account-before-billing pattern on the previous slides was also found by reading the results, not predicted.

