At school, a maths teacher often says: "Show your working." There are two good reasons. When a student writes each step on the page, they are less likely to lose track of a long calculation, because they do not have to hold every number in their head at once. And when the answer is wrong, the teacher can see exactly where it went wrong.
But the teacher does not ask for working on every question. If the question is "Which of these four boxes does this letter belong in?", a paragraph of working before the one-word answer makes the page longer and the marking slower, and it may not make the answer any better.

You can ask a language model to show its working too. This lesson measures when that helps and what it costs. I asked the same questions in two ways, once for the answer alone and once for the working first and then the answer, on two kinds of task: sums and word problems, which need several steps, and lesson 1's sorting task, which needs one word. Then I counted the right answers and the length of every reply.

Step by step. Asking the model to write out its working before it gives the answer. In this lesson the request is one sentence: "Think step by step, then give the final answer on a last line starting with 'Answer:'." You will often see this called chain of thought: the chain is the list of steps the model writes.
Direct. Asking for the answer alone: "Answer with the number only", or, for the sorting task, lesson 1's format rule, "Answer with the category word only, in lower case, and nothing else."
Answer line. The last line of a step-by-step reply, starting with "Answer:". A person can read the working; a program needs to find the one answer inside a long reply, and this line is where it looks.
Tokens written. How many tokens the model wrote in its reply. A token is a piece of a word, as the previous chapter showed. Ollama reports this count for every call as eval_count. More tokens written means a longer wait and, on a paid service, a larger bill.
Cut. A reply that the token limit stopped before the model had finished. Ollama marks it with done_reason set to "length". A cut step-by-step reply may never reach its answer line.

The lab uses four tasks, 40 questions each.
For the sorting tasks the step-by-step sentence takes the place of the format rule, because the two contradict each other: you cannot write "the category word only" and also write your working.
Each question went to qwen2.5:3b and llama3.2:3b at temperature 0 (the model always takes its most likely next token), with up to 512 tokens for every reply, direct ones too. A generous limit matters: in the previous chapter, a model that reasoned at length scored 0 of 40 simply because 32 tokens were not enough room. Here I counted the replies that hit the limit, so a cut reply cannot hide. Four tasks, two ways, 40 questions and two models make 640 calls. I then ran all 640 a second time.

Start with the word problems, the task that most needs several steps. Asked for the number only, qwen2.5:3b got 2 of the 40 right. Asked to work step by step, it got 37. On llama3.2:3b, the same change took the score from 7 to 37.
A score of 2 of 40 looked too bad to be true, so before believing it I read all 40 replies. They were bare numbers, and wrong ones: 31 for the pens (the right answer is 29), 46 for a bus question whose answer is 25, 159 for a reading question whose answer is 113. The answer check was not the problem. The model, asked to jump straight to a number, simply wrote a wrong one.
Llama's 7 right answers have a detail worth noticing. It was told "Answer with the number only", yet 8 of its 40 replies included working, such as "85 * 4 = 340, 60 * 2 = 120, 340 + 120 = 460". Six of its 7 right answers were among those 8. Of the 32 replies that obeyed and gave a bare number, only 1 was right. So even inside the direct test, the answers that came with working were the ones that were right.
Question by question, step by step fixed 35 of qwen's problems and broke none. For llama it fixed 31 and broke 1. With a split that one-sided, the sign test from the previous chapter's lesson 9 (it asks how often luck alone would split the changed questions this unevenly) gives p below 0.001. The p value is the chance of a split this lopsided if step by step made no real difference. Below 1 in 1,000, luck is very unlikely to explain it.

Here is every task on both models, right answers of 40, direct then step by step.
| Task | qwen2.5:3b | llama3.2:3b |
|---|---|---|
| Sums | 30 to 32 | 23 to 37 |
| Words | 2 to 37 | 7 to 37 |
| Sort C | 36 to 38 | 20 to 38 |
| Sort D | 40 to 40 | 36 to 39 |
The pattern is not "step by step helps arithmetic and not sorting". It is closer to this: in this lab, step by step helped most where the direct answer was weak, and it did little where the direct answer was already good. Llama's sorting with prompt C was weak (20 of 40) and step by step took it to 38. Qwen's sorting with prompt D was already perfect, and there was nothing left to gain.
Two results need care. On sums, llama went from 23 to 37, 15 questions fixed and 1 broken (sign test p = 0.0005, clearly more than luck), but qwen went only from 30 to 32: 6 fixed and 4 broken, p = 0.754, which is well within luck. And on sort D, llama's 36 to 39 is 3 fixed and 0 broken, p = 0.25: it may be real, but 40 messages cannot show it. The sign test is the fair way to read a small change here, because a net "plus two" can hide very different things underneath, as the sorting slide shows.
The sums are the most interesting case, because both models were asked exactly the same thing and only one of them gained.
The direct scores repeat the previous chapter. Lesson 11 measured the same 40 sums with the same wording and got 30 for qwen and 23 for llama; this lab got 30 and 23 again. Most of llama's direct mistakes were one digit off: 5425 for 57 x 95, whose answer is 5415. Step by step, llama usually split the sum into tens and ones ("50 x 95 = 4750, 7 x 95 = 665, 4750 + 665 = 5415") and got 37 right, in a median of 78 tokens (the median is the middle value when the 40 are put in order).
Qwen wrote more, a median of 180 tokens per sum, and it often chose a cleverer route than tens and ones. Three of its 8 wrong answers used an algebra trick that does not apply. For 98 x 92 it wrote that 98 x 92 equals (98 - 92) x (98 + 92), which is not true, and answered 1140; the right answer is 9016. Asked directly, it had answered this sum correctly, 9016. One possible reason qwen gained less is that its longer, fancier working gave it more places to go wrong; one run on 40 sums cannot show that it is the reason.
So for sums, in short: step by step clearly helped llama, and it did not clearly help qwen. A result on one model does not transfer automatically to another.

On qwen, sorting with prompt C went from 36 to 38. That small net change hides movement in both directions. Step by step fixed the four mistakes lesson 1 had already found (three account messages answered "billing", and the shipping refund answered "returns"). But it also broke two billing messages that the direct answer had right. For "My card was declined but the money still left my bank", the working decided the problem was "related to the customer's account and payment settings", and the answer became "account".
On llama, prompt C went from 20 to 38, the biggest change on any sorting run. With the direct prompt, llama answered "returns" 30 times, the same habit lesson 4 found with this prompt. Step by step, it answered "returns" 12 times and each of the other categories 9 or 10 times. One possible reason: when llama had to say what the message was about first ("The message mentions an invoice, which is typically related to financial transactions"), the category it named after that was more often the right one. The lab cannot show the reason, only the change.
In the second run, one of qwen's step-by-step final answers changed and its score was 39; the second-run slide explains. With prompt D, where the context lines already say what belongs in each box, qwen was 40 of 40 both ways, and llama went from 36 to 39, which the sign test cannot tell from luck. So on sorting, step by step was worth trying only where the prompt left the model confused; once the prompt said clearly where the lines are, one word was enough.

Because the working is written down, you can read where an answer went wrong, just as the teacher does. For "What is 91 x 44?", qwen's working found both parts correctly, 3640 and 364. Then it added them in pieces, 3640 + 300 + 64, and wrote that this was 3900 + 64. It is 3940 + 64. Every step was right except that one, and the final answer, 3964, was wrong.
To count cases like this and not only find them by eye, I added a check to the lab after reading the replies: it finds every written calculation in the working, such as "42 - 15 + 9 = 36 + 9", and tests whether it is true. It found a false calculation in the working of 11 of the 17 wrong step-by-step answers across both models, and in none of the 143 right ones. The first versions of this check flagged true lines as false, for example "400 metres x 13 = 5200", where the word "metres" split the calculation; I fixed it and read every line it flagged before trusting the counts.

The checker flags a false calculation, not a kind of mistake. Some of the 11 are wrong plans that happen to also write a false line, such as qwen's misused algebra on 49 x 52 and llama's 58 x 65, which counted the parts twice like 43 x 60 below. The other 6 wrong answers had no false line at all. For 43 x 60, llama wrote "3 x 60 = 180" twice and added both. Every line is true; the plan is wrong. Llama answered this sum correctly when asked directly, so this is the one sum that step by step broke for it. In a word problem about a car's fuel, llama worked out 900 ÷ 100 = 9 and called that the litres, forgetting to multiply by 7.
The most surprising replies I found, by reading, are the ones where the working reaches the right answer and the final line says something else.
A word problem said: "A pizza is cut into 8 slices. 12 friends each eat 2 slices. How many whole pizzas do they eat?" Llama's working was right: 12 x 2 = 24 slices, and "24 slices ÷ 8 slices/pizza = 3 pizzas". Then it added a sentence that makes no sense ("24 slices x 3 pizzas = 72 slices"), decided they had not eaten a full third pizza, and wrote "Answer: 2 whole pizzas."
The same thing happened on the sorting task. For "How do I turn off your marketing emails?", llama's working said the customer "is inquiring about their account settings, specifically how to opt-out of marketing emails". Its last line was "Answer: Returns."
These two cases matter for one practical reason: the program that reads the reply only sees the answer line. A person reading the working would be sure the model meant "3" and "account", and would be wrong about what the program stored. When you judge a step-by-step prompt by reading a few replies, read the answer lines, not only the working.
I also looked for the reverse, a right final answer after wrong working. The calculation check found none: no right answer had a false calculation in its working. That does not prove the working behind every right answer was sound, because the check cannot see a wrong plan, only a wrong sum.

The working is not free. A direct answer was usually a number or a word plus the end-of-turn marker: a median of 2 to 5 tokens. A step-by-step answer on qwen was a median of 180 tokens for a sum, 204 for a word problem, and 78.5 to 129 for a sorting message. Llama was shorter but still far longer than a direct answer: a median of 78 for a sum, 112 for a word problem and 86.5 to 87.5 for a sorting message.
So for a qwen word problem, the working added 204 - 4 = 200 tokens to the median reply, about 51 times as many tokens (204 ÷ 4). The model writes one token at a time, so the wait for the answer grows with the length of the reply. This lab does not report times, because other work was running on the same Ollama during the runs and would have made any time wrong. Unlike time, the token count is not stretched by other work on the machine, and it is what a paid service charges for. It is not perfectly fixed: in the second run, 15 step-by-step replies came back with a different length.
No reply was cut. The longest step-by-step reply in the whole run was 352 tokens, under the 512 limit. With a limit of 100, 121 of qwen's 160 step-by-step replies and 53 of llama's 160 were longer than that and would have been cut before their answer line, and the score would have measured the limit instead of the model, as happened to a thinking model in the previous chapter.


A fair way to judge the cost is to ask what each extra right answer cost. For qwen on the word problems, the 40 direct replies wrote 145 tokens in total and got 2 right. The 40 step-by-step replies wrote 8,194 tokens and got 37 right. That is 8,194 - 145 = 8,049 more tokens for 37 - 2 = 35 more right answers, or 8,049 ÷ 35 = 230 tokens per extra right answer. In most uses I can think of, a wrong answer costs more than 230 tokens of writing, so on this task it is a good trade.
The same sum for llama's word problems: (4,630 - 216) ÷ (37 - 7) = 4,414 ÷ 30 = 147 tokens per extra right answer. For llama's sums: (3,619 - 110) ÷ (37 - 23) = 3,509 ÷ 14 = 251 tokens.
Now qwen's sums: (7,981 - 190) ÷ (32 - 30) = 7,791 ÷ 2 = 3,896 tokens per extra right answer, and since a gain of 2 is within luck, it may be a lot of tokens for nothing. And qwen's sort D: 3,799 tokens against 80 for the same 40 right answers, which gains no right answers. The same sentence in the prompt was cheap for what it gained on one task and only a cost on another.

One run on small models is one data point, so I sent all 640 prompts a second time with the same settings. Every score came out the same except one: qwen's sorting with prompt C, step by step, was 38 in the first run and 39 in the second.
Looking closer shows something useful about long replies. All 320 direct replies were identical in both runs, letter for letter. Of the 320 step-by-step replies, 15 came back with different text: 11 of qwen's prompt C replies, 2 of llama's sums, 1 of llama's word problems and 1 of llama's prompt C replies. In 14 of those 15 the wording changed but the final answer did not. The one final answer that changed was "My card was declined but the money still left my bank": "account" in 176 tokens in the first run, "billing" in 103 tokens in the second.
Temperature 0 means the model always takes its most likely token, so why would a reply change? I did not test the reason. One possible reason: other labs were sending requests to the same Ollama while both runs ran, and when a server works on several requests at once, its arithmetic can differ in the last decimal places. That is enough to change which of two almost equally likely tokens is chosen. A one-word reply has one or two chances for that to happen; a reply of 150 tokens has 150, and once one token differs, everything after it can differ too. Whatever the reason, the practical point stands: a long reply was less repeatable than a short one in this lab, and a test of a step-by-step prompt should allow for that.
This script asks qwen2.5:3b three of the lab's questions, each two ways, and prints the final answer and the tokens written.

Before you run this lab. It uses qwen2.5:3b, running in Ollama on your own computer. If you have not set that up yet, the lab setup guide shows how to install Ollama, download the model and check that everything works, on macOS, Windows or Linux. You can use a different model instead: the guide shows the one line to change, and your numbers will differ from the ones in this lesson.
"""The same three questions, answered directly and then step by step, with the tokens each way cost.
This one uses qwen2.5:3b. Pull it first (see the lab setup guide):
ollama pull qwen2.5:3b
python stepwise_demo.py
"""
import json
import re
import urllib.request
ONLY = "Answer with the number only."
STEP = "Think step by step, then give the final answer on a last line starting with 'Answer:'."
RULES = "Put this customer message into one of these categories: billing, delivery, returns, account."
FORMAT = "Answer with the category word only, in lower case, and nothing else."
PEN = ("A shop sells pens at 3 dollars each. Maya buys 7 pens and pays with a 50-dollar note. "
"How many dollars of change does she get?") # right answer: 29
SUM = "What is 35 x 61?" # right answer: 2135
RESET = "Message: I forgot my password and the reset email never comes." # right answer: account
QUESTIONS = [
("pens", f"{PEN} {ONLY}", f"{PEN} {STEP}"),
("35 x 61", f"{SUM} {ONLY}", f"{SUM} {STEP}"),
("reset", f"{RULES}\n\n{FORMAT}\n\n{RESET}\nCategory:", f"{RULES}\n\n{STEP}\n\n{RESET}"),
]
def ask(text):
body = {"model": "qwen2.5:3b", "stream": False, "messages": [{"role": "user", "content": text}],
"options": {"temperature": 0, "num_predict": 512}}
req = urllib.request.Request("http://localhost:11434/api/chat", data=json.dumps(body).encode(),
headers={"Content-Type": "application/json"})
d = json.loads(urllib.request.urlopen(req).read())
return d["message"]["content"], d["eval_count"]
def final(reply):
"""The text after the last 'Answer:', or the whole reply if there is none."""
parts = re.split(r"answer\s*:", reply.replace("*", ""), flags=re.I)
return parts[-1].strip().splitlines()[0] if parts[-1].strip() else ""
for name, direct, steps in QUESTIONS:
d_reply, d_tokens = ask(direct)
s_reply, s_tokens = ask(steps)
print(f"\n{name}")
print(f" direct: {final(d_reply)!r:<12} {d_tokens:>4} tokens written")
print(f" step by step: {final(s_reply)!r:<12} {s_tokens:>4} tokens written")

The lab is scripts/labs/prompting/stepwise.py. It imports lesson 1's messages, rules and scoring for the sorting tasks, rebuilds the previous chapter's 40 sums from the same random seed, and holds the 40 word problems with the arithmetic for each answer. It stores one results file per model. The repeat mode sends every stored prompt a second time; the rescore mode re-applies the answer check to every stored reply without calling the model, and the report mode prints both models.
The report shows, for each task: right answers both ways, how many questions step by step fixed and broke, the sign test, the median tokens written, and how many replies were cut. Then the second run, how many direct replies contained more than the answer, and how many step-by-step answers had a false calculation in their working.
Three parts of the lab were written or changed after I saw results, and I want you to know which. The check for false calculations was added after reading the replies, and fixed several times as described on the inside slide. The count of direct replies with more than the answer was added after I saw llama writing working when told not to; my first version called every wrong one-word category "not bare", because it reused lesson 1's usable check, which also requires the word to be right, and I fixed it. The mode was added after the demo script disagreed with the lab on one message. The check for false calculations also skips lines that only repeat the question (such as "91 * 44 = 3964"), a rule I added after reading replies, which changes its counts. The right-or-wrong verdicts themselves never changed: the answer check was the same from the first run.
This box has no model. It holds the lab's real results for the word problems and the sums on both models: for each of the 40 questions, whether the final answer was right (1 or 0) and how many tokens the reply wrote, both ways. It prints the right answers, how many questions step by step fixed and broke, the sign test, and the tokens per extra right answer.
It prints the same numbers as the lab: for qwen's word problems, 2 to 37 right, 35 fixed and 0 broken, and 230 tokens per extra right answer; for qwen's sums, 30 to 32, 6 fixed and 4 broken, p = 0.754, and 3,896 tokens per extra right answer. Try changing a "0" to a "1" in one of the "steps" strings and watch the price of an extra right answer fall, or add a task of your own with its own right-and-wrong string.
The two sentences. ONLY and STEP are the only difference between the two ways of asking. Keeping them as named strings makes it easy to see that nothing else changed.
The questions. QUESTIONS holds a name and the two full prompts for each question. For the sorting message, the step-by-step prompt drops the format rule, because "the category word only" and "think step by step" cannot both be followed.
The call. ask sends one prompt to qwen2.5:3b at temperature 0 with room for 512 tokens, and returns the reply and eval_count, the number of tokens written. The lab also sets seed 1 and a 4,096-token context; the demo leaves Ollama's defaults.
The final answer. final splits the reply at every "Answer:" (ignoring upper or lower case and the stars some replies use for bold) and keeps the first line after the last one. A direct reply has no "Answer:", so the whole reply is the answer. The lab's own parser does the same and then pulls out the first number or the first category word from that line.

Always ask for an answer line. A step-by-step reply is many lines long. Without an agreed last line, your code has to guess which number in the working is the answer, and the working is full of numbers. "Give the final answer on a last line starting with 'Answer:'" is a format rule for the end of the reply, and all 320 step-by-step replies in the first run followed it.
Parse only the answer line, and take the last one. Six of qwen's word-problem replies wrote "Final answer:". The lab also ignores upper and lower case, and the stars that Markdown (a common text format) uses for bold, in case a reply writes the word as **Answer:**. Take the text after the last "Answer:", and from it the first number or category word. Do not take "the last number in the reply" if you can avoid it: the working can end with a number that is not the answer.
Give it room, and count cut replies. Set the token limit well above the longest reply you expect and count done_reason equal to "length". A cut step-by-step reply has no answer at all.
Score both ways on the same cases. Run the same 40 or so cases direct and step by step, and compare question by question: how many it fixed, how many it broke, and the sign test. A net gain of 2 can be 6 fixed and 4 broken.


Use it when the answer needs several steps and the direct answer is often wrong. On the word problems, both models went from a handful right to 37 of 40. That was the largest change in this lab.
Use it when a weak prompt leaves the model with one favourite answer. Llama's "returns" habit on prompt C mostly went away when it had to describe the message first. But fixing the prompt, as lesson 1's context lines did, is cheaper: one word per answer instead of about 87 tokens.
Do not use it when the direct answer is already right. On qwen with prompt D, 40 of 40 both ways, step by step only added a median of 78.5 tokens to every reply.
Do not assume it transfers between models. On the same 40 sums, it clearly helped llama and did not clearly help qwen.
Do not trust the working as an explanation. I found two replies where the working reached one answer and the final line gave another. The answer line is what your program uses; judge the prompt by that.

This lab ran two small models, twice, at temperature 0, on four tasks of 40 questions. I wrote all 40 word problems for this test, and after the runs I found that two of them are ambiguous. One is: "A delivery van makes 14 stops a day, 5 days a week. It misses 9 stops for a holiday. How many stops does it make in 4 weeks?" Both models read the 9 missed stops as happening every week and answered 244; I meant once, 271. A clearer wording might have scored one more right on each model. The other says "his mother's age is 3 times his age minus 5 years", which can also be read as 3 × (14 − 5), giving 33 in six years instead of 43; both models read it as I meant (3 × 14 − 5 = 37 now, 43 in six years), so no score changes.
It tried one wording of the step-by-step request. Other wordings, or a worked example of the steps (the few-shot idea from lesson 4 applied to working), may do better or worse. It did not test larger models, or "thinking" models such as qwen3, which reason at length whether you ask or not. It measured tokens, not time, because other work shared the machine. And the word problems are short, two or three steps: on a problem with ten steps, the chance of a slip somewhere in the working grows, and this lab says nothing about that.

Take a prompt of yours whose answers need some working out: a calculation, a date, a rule applied to a case. Collect 40 cases with the answers you expect. Run them direct and step by step, with an answer line and a generous token limit. Count right answers both ways, question by question, and count the tokens written. Work out the price of each extra right answer, the way the math slide does. Then read the working of five wrong answers: you will see whether the model slipped on a calculation, followed a wrong plan, or reasoned correctly and then wrote a different final line.

4 questions - Score 80% to pass
On 40 word problems, qwen2.5:3b got 2 right when asked for the number only and 37 when asked to think step by step. What did the extra right answers cost?
On the 40 sums, qwen2.5:3b went from 30 to 32 right: 6 sums fixed and 4 broken. What can you conclude?
Llama's working for the pizza problem reached 3 pizzas, but its last line said 'Answer: 2 whole pizzas.' What does your program store?
With lesson 1's prompt D, qwen2.5:3b sorted 40 of 40 messages right both ways. Should that prompt ask for step-by-step working?
This is a real run in VS Code's terminal.

All three questions change answer, and all three changes are fixes. The pens problem is 31 directly, which is wrong, and 29 step by step. Notice that the step-by-step answer line was a sentence, "Maya gets $29 in change.", not a bare number: that is why the lab takes the first number from the answer line instead of comparing the whole line. 35 x 61 is one of the three sums that none of the three direct models got right in the previous chapter's lesson 11: 2165 directly, 2135 (right) step by step, at 200 tokens against 5. The password message is "billing" directly and "account" step by step, one of the four prompt C messages that step by step fixed for qwen.
On my laptop, all six answers and all six token counts matched both of the lab's runs. The first version of this script used the card message instead of the password message, and its step-by-step answer disagreed with the lab's first run. That disagreement is why the lab has a second run. I changed the script to a message whose reply was the same in both runs, so what you see here is repeatable.
repeat
