Prompting

Thinking Step by Step: When It Helps, and What It Costs

0 of 21 complete

0%

Contents

Back|PromptingThinking Step by Step: When It Helps, and What It Costs
1/21
55 min left
Prerequisites
Quoting Untrusted Input: Tags, a Reminder, and a Fake Closing Tagrequired
Related Topics
Chat Templates: The Text a Conversation BecomesHow Models GenerateStructured Output Costs Right Answers: One JSON Box, MeasuredAgents in ProductionWhat a Token Is: How a Language Model Reads TextTokens and EmbeddingsHow a Tokenizer Learns Its Pieces: Byte Pair Encoding, Built From ScratchTokens and EmbeddingsWhat a Language Model Actually Outputs: Odds for Every Next TokenHow Models Generate
1 of 21

Show Your Working

At school, a maths teacher often says: "Show your working." There are two good reasons. When a student writes each step on the page, they are less likely to lose track of a long calculation, because they do not have to hold every number in their head at once. And when the answer is wrong, the teacher can see exactly where it went wrong.

But the teacher does not ask for working on every question. If the question is "Which of these four boxes does this letter belong in?", a paragraph of working before the one-word answer makes the page longer and the marking slower, and it may not make the answer any better.

An illustration of a teacher at a kitchen table in the evening, marking sheets of numbers with a pen, a tall pile of sheets beside her and a lamp and a mug of tea on the table. Under the heading asking for the working, not only the answer, titled step by step helped on word problems, and cost tokens. Beneath: qwen2.5:3b, 40 word problems: 2 right when asked for the number only, 37 when asked to work step by step, at a median 204 tokens written instead of 4.

You can ask a language model to show its working too. This lesson measures when that helps and what it costs. I asked the same questions in two ways, once for the answer alone and once for the working first and then the answer, on two kinds of task: sums and word problems, which need several steps, and lesson 1's sorting task, which needs one word. Then I counted the right answers and the length of every reply.

Five Words for This Lesson

A hand-drawn list headed five words for this lesson, titled thinking step by step. Step by step: the model writes its working first, then the answer. Direct: the model writes the answer only. Answer line: a last line starting with Answer:, so code can find it. Tokens written: the length of the reply, counted by Ollama. Cut: a reply stopped by the token limit before it ended. Beneath: step-by-step prompting is often called chain of thought.

Step by step. Asking the model to write out its working before it gives the answer. In this lesson the request is one sentence: "Think step by step, then give the final answer on a last line starting with 'Answer:'." You will often see this called chain of thought: the chain is the list of steps the model writes.

Direct. Asking for the answer alone: "Answer with the number only", or, for the sorting task, lesson 1's format rule, "Answer with the category word only, in lower case, and nothing else."

Answer line. The last line of a step-by-step reply, starting with "Answer:". A person can read the working; a program needs to find the one answer inside a long reply, and this line is where it looks.

Tokens written. How many tokens the model wrote in its reply. A token is a piece of a word, as the previous chapter showed. Ollama reports this count for every call as eval_count. More tokens written means a longer wait and, on a paid service, a larger bill.

Cut. A reply that the token limit stopped before the model had finished. Ollama marks it with done_reason set to "length". A cut step-by-step reply may never reach its answer line.

The Lab: Four Tasks, Two Ways of Asking

An editorial frame labelled one user message each, headed the same question, asked two ways, titled only the last sentence changes. The question: A shop sells pens at 3 dollars each. Maya buys 7 pens and pays with a 50-dollar note. How many dollars of change does she get? Direct: Answer with the number only. Step by step: Think step by step, then give the final answer on a last line starting with 'Answer:'. Beneath: every call, temperature 0, up to 512 tokens, both ways.

The lab uses four tasks, 40 questions each.

  • Sums: the 40 two-digit multiplications from lesson 11 of the previous chapter, such as "What is 57 x 95?". In this lesson "sums" means these two-digit multiplications. The direct wording is exactly that lesson's: "Answer with the number only."
  • Words: 40 short word problems that I wrote for this test, each needing two or three steps, such as the pens question above. The lab stores the arithmetic beside each problem (for the pens, 50 - 3*7) and computes the right answer from it, so a slip in my own arithmetic cannot become a wrong answer key.
  • Sort C: lesson 1's 40 customer messages and its prompt C, the instruction naming four categories plus the format rule.
  • Sort D: the same, with lesson 1's context lines, which made qwen2.5:3b perfect in lesson 1.

For the sorting tasks the step-by-step sentence takes the place of the format rule, because the two contradict each other: you cannot write "the category word only" and also write your working.

Each question went to qwen2.5:3b and llama3.2:3b at temperature 0 (the model always takes its most likely next token), with up to 512 tokens for every reply, direct ones too. A generous limit matters: in the previous chapter, a model that reasoned at length scored 0 of 40 simply because 32 tokens were not enough room. Here I counted the replies that hit the limit, so a cut reply cannot hide. Four tasks, two ways, 40 questions and two models make 640 calls. I then ran all 640 a second time.

Word Problems: 2 Right, Then 37

Two panels headed qwen2.5:3b, 40 word problems of two or three steps, titled asked for the working, it got 37 of 40 right. Left, number only: 2 of 40, median 4 tokens written. Right, step by step: 37 of 40, median 204 tokens written. Beneath: on llama3.2:3b, 7 and 37 of 40.

Start with the word problems, the task that most needs several steps. Asked for the number only, qwen2.5:3b got 2 of the 40 right. Asked to work step by step, it got 37. On llama3.2:3b, the same change took the score from 7 to 37.

A score of 2 of 40 looked too bad to be true, so before believing it I read all 40 replies. They were bare numbers, and wrong ones: 31 for the pens (the right answer is 29), 46 for a bus question whose answer is 25, 159 for a reading question whose answer is 113. The answer check was not the problem. The model, asked to jump straight to a number, simply wrote a wrong one.

Llama's 7 right answers have a detail worth noticing. It was told "Answer with the number only", yet 8 of its 40 replies included working, such as "85 * 4 = 340, 60 * 2 = 120, 340 + 120 = 460". Six of its 7 right answers were among those 8. Of the 32 replies that obeyed and gave a bare number, only 1 was right. So even inside the direct test, the answers that came with working were the ones that were right.

Question by question, step by step fixed 35 of qwen's problems and broke none. For llama it fixed 31 and broke 1. With a split that one-sided, the sign test from the previous chapter's lesson 9 (it asks how often luck alone would split the changed questions this unevenly) gives p below 0.001. The p value is the chance of a split this lopsided if step by step made no real difference. Below 1 in 1,000, luck is very unlikely to explain it.

Four Tasks, Two Models

A bar chart headed right answers of 40, four tasks, two ways, two models, titled it helped most where the direct answer was weak. For sums, words, sort C and sort D, four bars each: qwen2.5 direct 30, 2, 36, 40; qwen2.5 steps 32, 37, 38, 40; llama3.2 direct 23, 7, 20, 36; llama3.2 steps 37, 37, 38, 39. Beneath: the same numbers, direct to steps.

Here is every task on both models, right answers of 40, direct then step by step.

Taskqwen2.5:3bllama3.2:3b
Sums30 to 3223 to 37
Words2 to 377 to 37
Sort C36 to 3820 to 38
Sort D40 to 4036 to 39

The pattern is not "step by step helps arithmetic and not sorting". It is closer to this: in this lab, step by step helped most where the direct answer was weak, and it did little where the direct answer was already good. Llama's sorting with prompt C was weak (20 of 40) and step by step took it to 38. Qwen's sorting with prompt D was already perfect, and there was nothing left to gain.

Two results need care. On sums, llama went from 23 to 37, 15 questions fixed and 1 broken (sign test p = 0.0005, clearly more than luck), but qwen went only from 30 to 32: 6 fixed and 4 broken, p = 0.754, which is well within luck. And on sort D, llama's 36 to 39 is 3 fixed and 0 broken, p = 0.25: it may be real, but 40 messages cannot show it. The sign test is the fair way to read a small change here, because a net "plus two" can hide very different things underneath, as the sorting slide shows.

Sums: One Model Gained, One Did Not

The sums are the most interesting case, because both models were asked exactly the same thing and only one of them gained.

The direct scores repeat the previous chapter. Lesson 11 measured the same 40 sums with the same wording and got 30 for qwen and 23 for llama; this lab got 30 and 23 again. Most of llama's direct mistakes were one digit off: 5425 for 57 x 95, whose answer is 5415. Step by step, llama usually split the sum into tens and ones ("50 x 95 = 4750, 7 x 95 = 665, 4750 + 665 = 5415") and got 37 right, in a median of 78 tokens (the median is the middle value when the 40 are put in order).

Qwen wrote more, a median of 180 tokens per sum, and it often chose a cleverer route than tens and ones. Three of its 8 wrong answers used an algebra trick that does not apply. For 98 x 92 it wrote that 98 x 92 equals (98 - 92) x (98 + 92), which is not true, and answered 1140; the right answer is 9016. Asked directly, it had answered this sum correctly, 9016. One possible reason qwen gained less is that its longer, fancier working gave it more places to go wrong; one run on 40 sums cannot show that it is the reason.

So for sums, in short: step by step clearly helped llama, and it did not clearly help qwen. A result on one model does not transfer automatically to another.

Sorting: Where One Word Is Enough, Mostly

A two-column page headed qwen2.5:3b, sorting with lesson 1's prompt C, titled plus two was four fixed and two broken. Fixed by step by step: I need a refund of the shipping charge you added by mistake, returns to billing; I forgot my password and the reset email never comes, billing to account; I cannot sign in with Google any more, billing to account; I want to update my phone number for two-factor login, billing to account. Broken by step by step: My card was declined but the money still left my bank, billing to account; I cancelled my subscription but you took another payment, billing to account. Beneath: 4 messages, 2 messages.

On qwen, sorting with prompt C went from 36 to 38. That small net change hides movement in both directions. Step by step fixed the four mistakes lesson 1 had already found (three account messages answered "billing", and the shipping refund answered "returns"). But it also broke two billing messages that the direct answer had right. For "My card was declined but the money still left my bank", the working decided the problem was "related to the customer's account and payment settings", and the answer became "account".

On llama, prompt C went from 20 to 38, the biggest change on any sorting run. With the direct prompt, llama answered "returns" 30 times, the same habit lesson 4 found with this prompt. Step by step, it answered "returns" 12 times and each of the other categories 9 or 10 times. One possible reason: when llama had to say what the message was about first ("The message mentions an invoice, which is typically related to financial transactions"), the category it named after that was more often the right one. The lab cannot show the reason, only the change.

In the second run, one of qwen's step-by-step final answers changed and its score was 39; the second-run slide explains. With prompt D, where the context lines already say what belongs in each box, qwen was 40 of 40 both ways, and llama went from 36 to 39, which the sign test cannot tell from luck. So on sorting, step by step was worth trying only where the prompt left the model confused; once the prompt said clearly where the lines are, one word was enough.

Inside the Working: Right Steps, Wrong Answer

A table headed qwen2.5:3b, What is 91 x 44? Right answer 4004, titled right steps, wrong answer. Direct: 3924, wrong. The parts: 91 x 40 = 364 x 10 = 3640 and 91 x 4 = 364. The sum: 3640 + 364 = 3640 + 300 + 64 = 3900 + 64 = 3964. Final line: Answer: 3964, wrong. Beneath: 3640 + 300 is 3940, not 3900. One slip while adding, and the working was otherwise right.

Because the working is written down, you can read where an answer went wrong, just as the teacher does. For "What is 91 x 44?", qwen's working found both parts correctly, 3640 and 364. Then it added them in pieces, 3640 + 300 + 64, and wrote that this was 3900 + 64. It is 3940 + 64. Every step was right except that one, and the final answer, 3964, was wrong.

To count cases like this and not only find them by eye, I added a check to the lab after reading the replies: it finds every written calculation in the working, such as "42 - 15 + 9 = 36 + 9", and tests whether it is true. It found a false calculation in the working of 11 of the 17 wrong step-by-step answers across both models, and in none of the 143 right ones. The first versions of this check flagged true lines as false, for example "400 metres x 13 = 5200", where the word "metres" split the calculation; I fixed it and read every line it flagged before trusting the counts.

A table headed llama3.2:3b, What is 43 x 60? Right answer 2580, titled every line true, the plan wrong. Direct: 2580, right. Steps 1 to 3: 40 x 60 = 2400, 3 x 60 = 180, 3 x 60 = 180. Step 4: 2400 + 180 + 180 = 2760. Final line: Answer: 2760, wrong. Beneath: each calculation is correct. The step 3 x 60 was done twice, so a checker of each line alone finds nothing.

The checker flags a false calculation, not a kind of mistake. Some of the 11 are wrong plans that happen to also write a false line, such as qwen's misused algebra on 49 x 52 and llama's 58 x 65, which counted the parts twice like 43 x 60 below. The other 6 wrong answers had no false line at all. For 43 x 60, llama wrote "3 x 60 = 180" twice and added both. Every line is true; the plan is wrong. Llama answered this sum correctly when asked directly, so this is the one sum that step by step broke for it. In a word problem about a car's fuel, llama worked out 900 ÷ 100 = 9 and called that the litres, forgetting to multiply by 7.

When the Working and the Answer Disagree

The most surprising replies I found, by reading, are the ones where the working reaches the right answer and the final line says something else.

A word problem said: "A pizza is cut into 8 slices. 12 friends each eat 2 slices. How many whole pizzas do they eat?" Llama's working was right: 12 x 2 = 24 slices, and "24 slices ÷ 8 slices/pizza = 3 pizzas". Then it added a sentence that makes no sense ("24 slices x 3 pizzas = 72 slices"), decided they had not eaten a full third pizza, and wrote "Answer: 2 whole pizzas."

The same thing happened on the sorting task. For "How do I turn off your marketing emails?", llama's working said the customer "is inquiring about their account settings, specifically how to opt-out of marketing emails". Its last line was "Answer: Returns."

These two cases matter for one practical reason: the program that reads the reply only sees the answer line. A person reading the working would be sure the model meant "3" and "account", and would be wrong about what the program stored. When you judge a step-by-step prompt by reading a few replies, read the answer lines, not only the working.

I also looked for the reverse, a right final answer after wrong working. The calculation check found none: no right answer had a false calculation in its working. That does not prove the working behind every right answer was sound, because the check cannot see a wrong plan, only a wrong sum.

What the Working Costs

A sketched bar chart headed qwen2.5:3b, median tokens written per answer, titled the working is most of the reply. Two-digit sums: 5 and 180. Word problems: 4 and 204, the longest bar. Sorting, prompt C: 2 and 129. Sorting, prompt D: 2 and 78.5. Beneath: bar length is tokens written step by step; the first number is the direct answer's tokens, the second the step-by-step answer's.

The working is not free. A direct answer was usually a number or a word plus the end-of-turn marker: a median of 2 to 5 tokens. A step-by-step answer on qwen was a median of 180 tokens for a sum, 204 for a word problem, and 78.5 to 129 for a sorting message. Llama was shorter but still far longer than a direct answer: a median of 78 for a sum, 112 for a word problem and 86.5 to 87.5 for a sorting message.

So for a qwen word problem, the working added 204 - 4 = 200 tokens to the median reply, about 51 times as many tokens (204 ÷ 4). The model writes one token at a time, so the wait for the answer grows with the length of the reply. This lab does not report times, because other work was running on the same Ollama during the runs and would have made any time wrong. Unlike time, the token count is not stretched by other work on the machine, and it is what a paid service charges for. It is not perfectly fixed: in the second run, 15 step-by-step replies came back with a different length.

No reply was cut. The longest step-by-step reply in the whole run was 352 tokens, under the 512 limit. With a limit of 100, 121 of qwen's 160 step-by-step replies and 53 of llama's 160 were longer than that and would have been cut before their answer line, and the score would have measured the limit instead of the model, as happened to a thinking model in the previous chapter.

Four isometric blocks headed tokens written for all 40 word problems, both models, titled forty answers, two ways. qwen2.5, direct: 145 tokens, a thin slab. qwen2.5, steps: 8,194 tokens, the tallest. llama3.2, direct: 216 tokens. llama3.2, steps: 4,630 tokens. Beneath: height is the tokens written for all 40, tallest 8,194. Right answers: qwen2.5 2 and 37, llama3.2 7 and 37.

The Price of Each Extra Right Answer

A two-column page headed qwen2.5:3b, word problems, all 40 answers, titled what each extra right answer cost. Direct: 145 tokens, 2 right. Step by step: 8,194 tokens, 37 right. Difference: 8,049 tokens, 35 more right. Per extra right answer: 8,049 ÷ 35 = 230 tokens. Beneath: tokens written, counted by Ollama; what 35 more right answers cost.

A fair way to judge the cost is to ask what each extra right answer cost. For qwen on the word problems, the 40 direct replies wrote 145 tokens in total and got 2 right. The 40 step-by-step replies wrote 8,194 tokens and got 37 right. That is 8,194 - 145 = 8,049 more tokens for 37 - 2 = 35 more right answers, or 8,049 ÷ 35 = 230 tokens per extra right answer. In most uses I can think of, a wrong answer costs more than 230 tokens of writing, so on this task it is a good trade.

The same sum for llama's word problems: (4,630 - 216) ÷ (37 - 7) = 4,414 ÷ 30 = 147 tokens per extra right answer. For llama's sums: (3,619 - 110) ÷ (37 - 23) = 3,509 ÷ 14 = 251 tokens.

Now qwen's sums: (7,981 - 190) ÷ (32 - 30) = 7,791 ÷ 2 = 3,896 tokens per extra right answer, and since a gain of 2 is within luck, it may be a lot of tokens for nothing. And qwen's sort D: 3,799 tokens against 80 for the same 40 right answers, which gains no right answers. The same sentence in the prompt was cheap for what it gained on one task and only a cost on another.

The Second Run

A table headed every prompt sent a second time, same settings, titled the second run. qwen2.5, direct: final answers changed: 0 of 160; reply text changed: 0. qwen2.5, steps: final answers changed: 1 of 160; reply text changed: 11. llama3.2, direct: final answers changed: 0 of 160; reply text changed: 0. llama3.2, steps: final answers changed: 0 of 160; reply text changed: 4. Beneath: temperature 0, but Ollama was shared with other work while both runs ran.

One run on small models is one data point, so I sent all 640 prompts a second time with the same settings. Every score came out the same except one: qwen's sorting with prompt C, step by step, was 38 in the first run and 39 in the second.

Looking closer shows something useful about long replies. All 320 direct replies were identical in both runs, letter for letter. Of the 320 step-by-step replies, 15 came back with different text: 11 of qwen's prompt C replies, 2 of llama's sums, 1 of llama's word problems and 1 of llama's prompt C replies. In 14 of those 15 the wording changed but the final answer did not. The one final answer that changed was "My card was declined but the money still left my bank": "account" in 176 tokens in the first run, "billing" in 103 tokens in the second.

Temperature 0 means the model always takes its most likely token, so why would a reply change? I did not test the reason. One possible reason: other labs were sending requests to the same Ollama while both runs ran, and when a server works on several requests at once, its arithmetic can differ in the last decimal places. That is enough to change which of two almost equally likely tokens is chosen. A one-word reply has one or two chances for that to happen; a reply of 150 tokens has 150, and once one token differs, everything after it can differ too. Whatever the reason, the practical point stands: a long reply was less repeatable than a short one in this lab, and a test of a step-by-step prompt should allow for that.

Try It Yourself

This script asks qwen2.5:3b three of the lab's questions, each two ways, and prints the final answer and the tokens written.

A real screenshot of VS Code with stepwise_demo.py open, lines 1 to 34 visible. It defines ONLY and STEP, the two ways of asking, and lesson 1's RULES and FORMAT; three questions: the pens word problem, What is 35 x 61?, and the message I forgot my password and the reset email never comes, each written both ways; a function ask that sends a prompt to qwen2.5:3b at temperature 0 with up to 512 tokens and returns the reply and the tokens written. The function that reads the final answer comes after line 34, in the box on the slide. Beneath: copy it from the box on the slide.

Before you run this lab. It uses qwen2.5:3b, running in Ollama on your own computer. If you have not set that up yet, the lab setup guide shows how to install Ollama, download the model and check that everything works, on macOS, Windows or Linux. You can use a different model instead: the guide shows the one line to change, and your numbers will differ from the ones in this lesson.

"""The same three questions, answered directly and then step by step, with the tokens each way cost.

This one uses qwen2.5:3b. Pull it first (see the lab setup guide):
    ollama pull qwen2.5:3b
    python stepwise_demo.py
"""
import json
import re
import urllib.request

ONLY = "Answer with the number only."
STEP = "Think step by step, then give the final answer on a last line starting with 'Answer:'."
RULES = "Put this customer message into one of these categories: billing, delivery, returns, account."
FORMAT = "Answer with the category word only, in lower case, and nothing else."

PEN = ("A shop sells pens at 3 dollars each. Maya buys 7 pens and pays with a 50-dollar note. "
       "How many dollars of change does she get?")                             # right answer: 29
SUM = "What is 35 x 61?"                                                       # right answer: 2135
RESET = "Message: I forgot my password and the reset email never comes."       # right answer: account

QUESTIONS = [
    ("pens", f"{PEN} {ONLY}", f"{PEN} {STEP}"),
    ("35 x 61", f"{SUM} {ONLY}", f"{SUM} {STEP}"),
    ("reset", f"{RULES}\n\n{FORMAT}\n\n{RESET}\nCategory:", f"{RULES}\n\n{STEP}\n\n{RESET}"),
]


def ask(text):
    body = {"model": "qwen2.5:3b", "stream": False, "messages": [{"role": "user", "content": text}],
            "options": {"temperature": 0, "num_predict": 512}}
    req = urllib.request.Request("http://localhost:11434/api/chat", data=json.dumps(body).encode(),
                                 headers={"Content-Type": "application/json"})
    d = json.loads(urllib.request.urlopen(req).read())
    return d["message"]["content"], d["eval_count"]


def final(reply):
    """The text after the last 'Answer:', or the whole reply if there is none."""
    parts = re.split(r"answer\s*:", reply.replace("*", ""), flags=re.I)
    return parts[-1].strip().splitlines()[0] if parts[-1].strip() else ""


for name, direct, steps in QUESTIONS:
    d_reply, d_tokens = ask(direct)
    s_reply, s_tokens = ask(steps)
    print(f"\n{name}")
    print(f"  direct:       {final(d_reply)!r:<12} {d_tokens:>4} tokens written")
    print(f"  step by step: {final(s_reply)!r:<12} {s_tokens:>4} tokens written")

The Lab Report

A real terminal recording of python stepwise.py report. For llama3.2:3b on Apple M4, 24 GB, temperature 0, up to 512 tokens a call: task, right of 40 direct and steps, fixed, broke, sign p, median tokens direct and steps, and cut. Sums 23, 37, 15, 1, below 0.001, 3, 78, 0, 0. Words 7, 37, 31, 1, below 0.001, 2, 112, 0, 0. Sort-c 20, 38, 18, 0, below 0.001, 2, 86.5, 0, 0. Sort-d 36, 39, 3, 0, 0.250, 2, 87.5, 0, 0. Run 2: the same right answers, 0 final answers changed. Direct replies with more than the answer: words 8, the rest 0. Sums, steps with a false equation: 0 of 37 right, 2 of 3 wrong; words: 0 of 37 right, 0 of 3 wrong. For qwen2.5:3b: sums 30, 32, 6, 4, 0.754, 5, 180; words 2, 37, 35, 0, below 0.001, 4, 204; sort-c 36, 38, 4, 2, 0.688, 2, 129; sort-d 40, 40, 0, 0, 1.000, 2, 78.5; no reply cut. Run 2: sort-c steps 39, with 1 final answer changed, the rest unchanged. Direct replies with more than the answer: 0 everywhere. Sums, steps with a false equation: 0 of 32 right, 7 of 8 wrong; words: 0 of 37 right, 2 of 3 wrong. Beneath: the lab's own report, you do not need to run it.

The lab is scripts/labs/prompting/stepwise.py. It imports lesson 1's messages, rules and scoring for the sorting tasks, rebuilds the previous chapter's 40 sums from the same random seed, and holds the 40 word problems with the arithmetic for each answer. It stores one results file per model. The repeat mode sends every stored prompt a second time; the rescore mode re-applies the answer check to every stored reply without calling the model, and the report mode prints both models.

The report shows, for each task: right answers both ways, how many questions step by step fixed and broke, the sign test, the median tokens written, and how many replies were cut. Then the second run, how many direct replies contained more than the answer, and how many step-by-step answers had a false calculation in their working.

Three parts of the lab were written or changed after I saw results, and I want you to know which. The check for false calculations was added after reading the replies, and fixed several times as described on the inside slide. The count of direct replies with more than the answer was added after I saw llama writing working when told not to; my first version called every wrong one-word category "not bare", because it reused lesson 1's usable check, which also requires the word to be right, and I fixed it. The mode was added after the demo script disagreed with the lab on one message. The check for false calculations also skips lines that only repeat the question (such as "91 * 44 = 3964"), a rule I added after reading replies, which changes its counts. The right-or-wrong verdicts themselves never changed: the answer check was the same from the first run.

Price the Working in Your Browser

This box has no model. It holds the lab's real results for the word problems and the sums on both models: for each of the 40 questions, whether the final answer was right (1 or 0) and how many tokens the reply wrote, both ways. It prints the right answers, how many questions step by step fixed and broke, the sign test, and the tokens per extra right answer.

It prints the same numbers as the lab: for qwen's word problems, 2 to 37 right, 35 fixed and 0 broken, and 230 tokens per extra right answer; for qwen's sums, 30 to 32, 6 fixed and 4 broken, p = 0.754, and 3,896 tokens per extra right answer. Try changing a "0" to a "1" in one of the "steps" strings and watch the price of an extra right answer fall, or add a task of your own with its own right-and-wrong string.

The Code, Part by Part

The two sentences. ONLY and STEP are the only difference between the two ways of asking. Keeping them as named strings makes it easy to see that nothing else changed.

The questions. QUESTIONS holds a name and the two full prompts for each question. For the sorting message, the step-by-step prompt drops the format rule, because "the category word only" and "think step by step" cannot both be followed.

The call. ask sends one prompt to qwen2.5:3b at temperature 0 with room for 512 tokens, and returns the reply and eval_count, the number of tokens written. The lab also sets seed 1 and a 4,096-token context; the demo leaves Ollama's defaults.

The final answer. final splits the reply at every "Answer:" (ignoring upper or lower case and the stars some replies use for bold) and keeps the first line after the last one. A direct reply has no "Answer:", so the whole reply is the answer. The lab's own parser does the same and then pulls out the first number or the first category word from that line.

How to Use Step by Step

A sketch of four stacked boxes joined by arrows, headed sketched, reading a step-by-step reply, titled working, then one line for code. The question + Think step by step. The working: many lines, for people. Answer: 29. Code keeps the text after the last Answer:. Beneath: the answer line is a format rule for the end of the reply.

Always ask for an answer line. A step-by-step reply is many lines long. Without an agreed last line, your code has to guess which number in the working is the answer, and the working is full of numbers. "Give the final answer on a last line starting with 'Answer:'" is a format rule for the end of the reply, and all 320 step-by-step replies in the first run followed it.

Parse only the answer line, and take the last one. Six of qwen's word-problem replies wrote "Final answer:". The lab also ignores upper and lower case, and the stars that Markdown (a common text format) uses for bold, in case a reply writes the word as **Answer:**. Take the text after the last "Answer:", and from it the first number or category word. Do not take "the last number in the reply" if you can avoid it: the working can end with a number that is not the answer.

Give it room, and count cut replies. Set the token limit well above the longest reply you expect and count done_reason equal to "length". A cut step-by-step reply has no answer at all.

Score both ways on the same cases. Run the same 40 or so cases direct and step by step, and compare question by question: how many it fixed, how many it broke, and the sign test. A net gain of 2 can be 6 fixed and 4 broken.

A table headed what step by step did in this lab, titled steps, measured. Word problems: qwen2.5 2 to 37, llama3.2 7 to 37. Two-digit sums: qwen2.5 30 to 32, llama3.2 23 to 37. Sorting, C: qwen2.5 36 to 38, llama3.2 20 to 38. Sorting, D: qwen2.5 40 to 40, llama3.2 36 to 39. Beneath: right of 40, direct to step by step. One run each, then a second run.

When to Use It, and When Not To

A flowchart headed should this prompt ask for the working?, titled measure it on your task. Is the direct answer often wrong? No: ask directly; test steps only if cheap. Yes: ask for steps and an Answer: line, then score both ways on 40 cases, then: clearly more right? Yes: keep it; count the extra tokens. No: ask directly. Beneath: also count replies cut by the token limit.

Use it when the answer needs several steps and the direct answer is often wrong. On the word problems, both models went from a handful right to 37 of 40. That was the largest change in this lab.

Use it when a weak prompt leaves the model with one favourite answer. Llama's "returns" habit on prompt C mostly went away when it had to describe the message first. But fixing the prompt, as lesson 1's context lines did, is cheaper: one word per answer instead of about 87 tokens.

Do not use it when the direct answer is already right. On qwen with prompt D, 40 of 40 both ways, step by step only added a median of 78.5 tokens to every reply.

Do not assume it transfers between models. On the same 40 sums, it clearly helped llama and did not clearly help qwen.

Do not trust the working as an explanation. I found two replies where the working reached one answer and the final line gave another. The answer line is what your program uses; judge the prompt by that.

What This Lab Can and Cannot Tell You

A two-column page headed read before you quote a number from this lesson, titled what was measured, and what was not. Measured: two small models, two runs; sums, word problems, sorting; one wording of the request; tokens written, not time. Not measured: larger or thinking models; hard, many-step problems; other wordings, or examples; speed on a machine to itself.

This lab ran two small models, twice, at temperature 0, on four tasks of 40 questions. I wrote all 40 word problems for this test, and after the runs I found that two of them are ambiguous. One is: "A delivery van makes 14 stops a day, 5 days a week. It misses 9 stops for a holiday. How many stops does it make in 4 weeks?" Both models read the 9 missed stops as happening every week and answered 244; I meant once, 271. A clearer wording might have scored one more right on each model. The other says "his mother's age is 3 times his age minus 5 years", which can also be read as 3 × (14 − 5), giving 33 in six years instead of 43; both models read it as I meant (3 × 14 − 5 = 37 now, 43 in six years), so no score changes.

It tried one wording of the step-by-step request. Other wordings, or a worked example of the steps (the few-shot idea from lesson 4 applied to working), may do better or worse. It did not test larger models, or "thinking" models such as qwen3, which reason at length whether you ask or not. It measured tokens, not time, because other work shared the machine. And the word problems are short, two or three steps: on a problem with ten steps, the chance of a slip somewhere in the working grows, and this lab says nothing about that.

What to Do Next

A hand-drawn list headed for your own prompt, titled four things to do. 1, both ways: score the prompt direct and step by step on the same cases. 2, answer line: end with Answer:, and parse only that line. 3, count: tokens written, and replies cut by the limit. 4, read: the working of a few wrong answers. Beneath: keep the steps only where they buy right answers.

Take a prompt of yours whose answers need some working out: a calculation, a date, a rule applied to a case. Collect 40 cases with the answers you expect. Run them direct and step by step, with an answer line and a generous token limit. Count right answers both ways, question by question, and count the tokens written. Work out the price of each extra right answer, the way the math slide does. Then read the working of five wrong answers: you will see whether the model slipped on a calculation, followed a wrong plan, or reasoned correctly and then wrote a different final line.

A closing card headed to keep, titled steps where the answer needs steps. In large type: 2 → 37. Beneath: word problems right of 40 on qwen2.5:3b, at a median 204 tokens instead of 4. Then: measure both ways.

Knowledge Check

Knowledge Check

4 questions - Score 80% to pass

Q1

On 40 word problems, qwen2.5:3b got 2 right when asked for the number only and 37 when asked to think step by step. What did the extra right answers cost?

Q2

On the 40 sums, qwen2.5:3b went from 30 to 32 right: 6 sums fixed and 4 broken. What can you conclude?

Q3

Llama's working for the pizza problem reached 3 pizzas, but its last line said 'Answer: 2 whole pizzas.' What does your program store?

Q4

With lesson 1's prompt D, qwen2.5:3b sorted 40 of 40 messages right both ways. Should that prompt ask for step-by-step working?

This is a real run in VS Code's terminal.

A real screenshot of VS Code's terminal after running python stepwise_demo.py. Pens: direct 31, 3 tokens written; step by step, Maya gets $29 in change., 93 tokens written. 35 x 61: direct 2165, 5 tokens written; step by step 2135, 200 tokens written. Reset: direct billing, 2 tokens written; step by step account, 127 tokens written.

All three questions change answer, and all three changes are fixes. The pens problem is 31 directly, which is wrong, and 29 step by step. Notice that the step-by-step answer line was a sentence, "Maya gets $29 in change.", not a bare number: that is why the lab takes the first number from the answer line instead of comparing the whole line. 35 x 61 is one of the three sums that none of the three direct models got right in the previous chapter's lesson 11: 2165 directly, 2135 (right) step by step, at 200 tokens against 5. The password message is "billing" directly and "account" step by step, one of the four prompt C messages that step by step fixed for qwen.

On my laptop, all six answers and all six token counts matched both of the lab's runs. The first version of this script used the card message instead of the password message, and its step-by-step answer disagreed with the lab's first run. That disagreement is why the lab has a second run. I changed the script to a message whose reply was the same in both runs, so what you see here is repeatable.

repeat

A sequence diagram with three columns: the lab, the model and the check. Step one, the lab sends the model the question plus one rule. Step two, the model returns the reply and a token count. Step three, the lab sends the check the final answer and the right one. Step four, the check returns right? cut? slips?. Beneath: 640 calls in each run, 4 tasks times 2 ways times 40 times 2 models; then all of it again.

Two brand cards headed the tools, with their logos, titled what the test ran on. Ollama: qwen2.5:3b and llama3.2:3b, Apple M4, 24 GB. Python: 1,280 chat calls, 2 runs of 640.