Prompting

Examples in a Prompt: How Many, Which Ones, in What Order

0 of 20 complete

0%

Contents

Back|PromptingExamples in a Prompt: How Many, Which Ones, in What Order
1/20
45 min left
Prerequisites
System Message or User Message: Where the Rules Gorequired
Related Topics
Chat Templates: The Text a Conversation BecomesHow Models GenerateWhat a Language Model Actually Outputs: Odds for Every Next TokenHow Models GenerateThe Context Window: What Happens When a Prompt Does Not FitHow Models GenerateQuantization: The Same Model in Fewer BitsHow Models GenerateOne Call, End to End: The Chapter in a Single RequestHow Models Generate
1 of 20

Show, Then Ask

When a new colleague joins a team, you can explain the job in words, or you can work next to them through a few real cases together. Most people learn faster from the cases. The words say what to do; the cases show what "done" looks like, including all the small details nobody thought to write down.

A language model can be taught the same way inside a prompt. Before the real question, you include a few finished cases: a message, then its right answer, written in exactly the form you want back. These are called examples, and a prompt that contains them is often called a few-shot prompt.

An illustration of two people at a whiteboard covered with a diagram, one pointing at it with a pen while the other holds a laptop and listens. Under the heading examples change the answers, and not only the score. Beneath: llama3.2:3b, 40 messages, 20 right with no examples, 28 to 36 with four, depending on which four and their order.

Lesson 1 added two examples to a prompt that was already perfect, so it could not tell whether they helped. This lesson tests examples properly, on a prompt that still makes mistakes, and on two models. It asks three questions: does adding examples help, does it matter which ones you choose, and does it matter what order they are in?

Examples in a Prompt

A hand-drawn list headed five words for this lesson, titled examples in a prompt. Example: a finished case, a message and its right answer. Zero-shot: a prompt with no examples. Few-shot: a prompt with a handful of examples. Order: which example comes first, and which comes last. Choice: which messages you picked to be the examples. Beneath: in this field, a shot means one example.

Example. One finished case inside the prompt: a message, then the answer you would want for it, written in the same format you want the model to use.

Zero-shot. A prompt with no examples at all, only the instruction. In this field a "shot" means one example.

Few-shot. A prompt with a handful of examples, often one to five.

Order. Which example comes first in the prompt and which comes last, just before the real question.

Choice. Which messages you picked to be the examples. Two sets can cover the same categories and still be quite different messages.

One more idea from earlier lessons matters here: the model does not learn anything permanent from the examples. It only reads them, as part of the text of this one request, and they change which next token is most likely. Remove them from the next request and they are gone.

The Lab: A Prompt With Room to Improve

To see whether examples help, the prompt must still make mistakes without them. So this lab uses lesson 1's prompt C: the instruction naming the four categories (billing, delivery, returns, account) and the format rule ("Answer with the category word only, in lower case, and nothing else"), but not the context lines that made prompt D perfect. The test set is the same 40 customer messages, 10 per category.

An editorial frame labelled one user message, top to bottom, headed set A, as the model reads it, titled the rules, four examples, then the message. Rules: the four categories and the format rule, lesson 1's prompt C. Billing example: My receipt shows a different amount from my card statement. Delivery and returns examples: The tracking number you sent does not work, and Can I send back a kettle that I opened? Account example: I cannot find the settings page to change my name. The message: Message, the customer's text, then Category:. Beneath: each example is written as Message: ..., then Category: and the answer.

I wrote 10 new messages to use as examples. None of them is one of the 40 test messages, and the lab checks that before it runs; an example that is also a test message would make the test meaningless. The check only catches exact copies, though. A few examples are close in meaning to a test message (for instance "Please stop sending me newsletters" and the test message "How do I turn off your marketing emails?"), so some scores may be a little flattering. Each example is written exactly like the question: "Message:", the text, then "Category:" and the answer.

The lab tries six versions of the prompt:

  • no examples: prompt C alone.
  • set A: one example per category, in the order billing, delivery, returns, account.
  • set A, reversed: the same four, in the order account, returns, delivery, billing.
  • set B: four different examples, one per category, in the same order as set A.
  • three billing and three account: three examples that all share the same answer.

With No Examples, It Said Returns

A sketched bar chart headed llama3.2:3b, no examples, how often each category was answered, titled with no examples, it said returns. Billing: 3. Delivery: 5. Returns: 30, the longest bar. Account: 2. Beneath: bar length is answers; each category really had 10 messages.

Start with llama3.2:3b and no examples. It got 20 of 40 right, and its mistakes were not spread out. It answered "returns" 30 times, though only 10 messages were about returns. Seven billing messages, five delivery messages and eight account messages all came back as "returns".

This is worth noticing, because it is typical of what goes wrong. A model with an unclear task often keeps giving one answer. The fallback here happened to be "returns"; with another model or another prompt it could be any of the four. The score, 20 of 40, hides this. Counting which answers came back shows it immediately.

qwen2.5:3b got 36 of 40 on the same prompt, with a different pattern of mistakes: three account messages answered "billing", and one billing message answered "returns". That is the same prompt C that lesson 1 measured, with the same result.

Four Examples

Two panels headed llama3.2:3b, lesson 1's prompt C, 40 messages, titled four examples, thirteen more right. Left, no examples: 20 of 40, answered returns 30 times. Right, four examples, set A: 33 of 40, one per category. Beneath: on qwen2.5:3b, 36 and 39 of 40.

Adding set A, one example per category, took llama from 20 to 33 right. The "returns" habit disappeared: it answered returns exactly 10 times. One possible reason: with one example of each category in front of it, the model saw all four answers being used.

But the mistakes did not vanish. They moved. With set A, seven account messages came back as "billing". None of the account messages had been answered "billing" before. The examples fixed one pattern of mistakes and created another.

On qwen, set A took the score from 36 to 39. A model that already does the task well has less room to improve, and examples added only a little.

Order Moved the Mistakes

A two-column page headed llama3.2:3b, the same four examples, two orders, titled order moved the mistakes. Set A: 33 of 40 right; account to billing, 7. Set A, reversed: 28 of 40 right; delivery to returns, 5; billing to returns, 4; billing to account, 3. Beneath: billing example first, account example last; account example first, billing example last.

Now the same four examples, in reverse order: account first, billing last. Nothing else changed, not one word of the examples or the rules. Llama's score fell from 33 to 28.

And again, the mistakes moved. In set A's order, the seven mistakes were all account messages answered "billing". In the reversed order, the account messages were all right, but five delivery messages and four billing messages went back to "returns", and three billing messages were answered "account".

It is tempting to explain this with a simple rule, such as "the model copies the last example". The data does not fit that rule. In the reversed order the last example was billing, yet billing messages were answered wrongly more often, not less. Nor does "copies the first example" hold up: set A over-answered billing (17 times) and set A reversed over-answered account (13 times), but set B, also with billing first, answered billing exactly 10 times. What the data does show is simpler and more useful: on this model and this task, the order of the examples changed both the score and which messages were wrong. On qwen, both orders gave 39.

What the Examples Change: the Odds

Why would the order of four examples change an answer? The previous chapter's first lesson gives the tool to look. At every step a model gives odds for every possible next token, and at temperature 0 it takes the most likely one. The lab asked Ollama for those odds on the first token of the answer, for three messages, under three of the prompts.

For "Please delete my account and all my data", with no examples, llama gave "returns" a probability of 0.95 and "account" only 0.03. With set A, "billing" came first at 0.53, with "account" second at 0.25. With set A reversed, "account" jumped to 0.89. The same message and the same rules, and the right answer went from 3 in 100 to 89 in 100, depending only on which examples came before it and in what order.

For "Please explain the two separate amounts on my statement", a billing question, the three prompts gave three different winners: "returns" at 0.59 with no examples, "billing" at 0.91 with set A, and "account" at 0.47 with set A reversed, with "billing" second at 0.23. For "The courier left the box at the wrong address", "delivery" won only with set A, at 0.77; with no examples and with the reversed set, "returns" won at 0.92 and 0.73.

Two things follow from these numbers. First, examples do not flip a switch inside the model; they move the odds, sometimes a little and sometimes a lot. Second, a big lead under one prompt is no protection against a change to the prompt. Under set A, "billing" led at 0.91 against 0.04 for the statement message, and reversing the order alone made "account" the winner. The margin (how far the winning answer is ahead of the second one) tells you how sure the model is under this exact prompt, not whether the answer will survive a different one.

So there is no shortcut: you cannot tell from one prompt's odds how fragile (easily broken by a small change) it is. Ollama can return these odds with logprobs (the odds, stored on a log scale), and they are useful for seeing how sure the model is, but the only way to find out whether an answer survives a different order or a different set is to run that order or that set.

Which Four Examples

Four isometric blocks headed llama3.2:3b, right answers of 40, titled which four, and in what order. No examples: 20 of 40. Set A, reversed: 28 of 40. Four, set A: 33 of 40. Four, set B: 36 of 40, the tallest. Beneath: height is right answers, full height 40; three prompts with four examples (two sets, one in two orders), 28 to 36.

Set B used four different messages, one per category, in the same order as set A. I wrote set B after seeing set A's results, so it is not a blind test; keep that in mind when you compare the two. Its messages were: "Why was I charged a cancellation fee?", "Can I choose a delivery time slot?", "Is there a fee for sending back a mattress?" and "Please stop sending me newsletters." Llama got 36 of 40 with set B, against 33 with set A. Its mistakes shrank to four, each a different kind.

So with four examples, one per category, llama scored anywhere from 28 to 36 depending on which four messages were used and in what order. That spread of 8 is more than half of the 13-point gain that adding examples gave in the first place. The examples helped, but how much they helped depended on which examples were picked and on their order, which neither the instruction nor the categories say anything about.

On qwen, the three prompts with four examples gave 38 to 39. The same pattern, but small.

A bar chart headed right answers of 40, six ways of adding examples, two models, titled the weaker model moved most. For none, set A, A rev, set B, 3 bill and 3 acct, two bars each: qwen2.5:3b 36, 39, 39, 38, 37, 36; llama3.2:3b 20, 33, 28, 36, 22, 34.

Three Examples With One Answer

A table headed llama3.2:3b, three examples that all share one answer, titled neither label was overused. Three billing: 22 of 40 right, answered billing 4 times. Three account: 34 of 40 right, answered account 7 times. No examples: 20 of 40 right, answered returns 30 times. Beneath: three examples with one label did not make the model overuse that label; it never passed the 10 real messages.

A common worry about examples is that the model will copy their answers: show it three billing examples, and it will say "billing" to everything. The last two prompts test that.

With three billing examples, llama answered "billing" only 4 times, fewer than the 10 billing messages in the set. Its old habit came back: it answered "returns" 26 times, and it got 22 of 40 right, barely better than no examples at all.

With three account examples, llama answered "account" only 7 times, and got 34 of 40 right, nearly as good as the best balanced set. I cannot tell from this lab why three account examples worked so much better than three billing examples. One possibility is that the account examples happened to break llama's "returns" habit across all categories (returns answers fell from 30 to 12), while the billing examples did not (26). With one run and three examples each, that stays a guess.

What the numbers do show is that the fear of copying did not come true here, on either model. Qwen scored 37 and 36 with these prompts, close to its 36 with no examples.

The Stronger Model Barely Moved

A table headed qwen2.5:3b, the same six prompts, titled the stronger model barely moved. No examples: 36 of 40 right. Four, any set: 38 to 39 of 40 right. Three, one label: 37 and 36 of 40 right. Beneath: here, the examples helped the weaker model far more.

Put the two models side by side and the main lesson is clear. On qwen2.5:3b, which already got 36 of 40 without examples, no set of examples changed the score by more than 3. On llama3.2:3b, which got only 20, examples added up to 16 right answers, and choosing and ordering them differently moved the result by up to 8.

In this lab, examples helped the weaker model most, and that is also where their choice and order mattered most. A prompt that works only because of one set of examples in one order that happened to work is fragile: change the examples a little, or change the model, and the result can move a long way.

What Four Examples Cost

A two-column page headed median prompt tokens, llama3.2:3b, titled what four examples cost. Prompt, tokens and right answers: no examples, 73 tokens, 20 right; four, set A, 138 tokens, 33 right; difference, 65 tokens, 13 more right. Beneath: read on every call; fixed text, reused if nothing that changes comes before it (previous chapter, lesson 6).

Four short examples added 65 tokens to llama's prompt: 73 tokens with none, 138 with set A (the median, the middle value of the 40, counted the way Ollama reports them, including the chat template, the hidden text Ollama wraps around your message). On this task, those 65 tokens bought 13 more right answers, a good trade.

Examples are fixed text: they are the same on every call. As lesson 6 of the previous chapter showed, fixed text at the start of a prompt can be reused from one call to the next, as long as nothing that changes comes before it. So put the examples after the rules and before the customer's message, and keep them word for word the same.

Try It Yourself

This script sends three messages to llama3.2:3b with no examples, with set A, and with set A reversed.

A real screenshot of VS Code with fewshot_demo.py open, lines 1 to 33 visible. It defines RULES, the four categories and the format rule; EXAMPLES, set A's four messages with their categories; three messages about two amounts on a statement, a courier who left a box at the wrong address, and deleting an account; a function ask that sends a text to llama3.2:3b on Ollama's chat endpoint at temperature 0 with up to 20 tokens; and a function prompt that joins the rules, the examples and the message. Beneath: copy it from the box on the slide.

Before you run this lab. It uses llama3.2:3b, running in Ollama on your own computer. If you have not set that up yet, the lab setup guide shows how to install Ollama, download a model and check that everything works, on macOS, Windows or Linux; download this one with ollama pull llama3.2:3b. You can use a different model instead, and your numbers will differ from the ones in this lesson.

"""The same prompt with no examples, four examples, and the same four examples in reverse order.

This one uses llama3.2:3b, where the effect is easy to see. Pull it first (see the lab setup guide):
    ollama pull llama3.2:3b
    python fewshot_demo.py
"""
import json
import urllib.request

RULES = ("Put this customer message into one of these categories: billing, delivery, returns, account.\n\n"
         "Answer with the category word only, in lower case, and nothing else.")
EXAMPLES = [("My receipt shows a different amount from my card statement.", "billing"),
            ("The tracking number you sent does not work.", "delivery"),
            ("Can I send back a kettle that I opened?", "returns"),
            ("I cannot find the settings page to change my name.", "account")]
MESSAGES = ["Please explain the two separate amounts on my statement.",     # billing
            "The courier left the box at the wrong address.",               # delivery
            "Please delete my account and all my data."]                    # account


def ask(text):
    body = {"model": "llama3.2:3b", "stream": False, "messages": [{"role": "user", "content": text}],
            "options": {"temperature": 0, "num_predict": 20}}
    req = urllib.request.Request("http://localhost:11434/api/chat", data=json.dumps(body).encode(),
                                 headers={"Content-Type": "application/json"})
    return json.loads(urllib.request.urlopen(req).read())["message"]["content"]


def prompt(examples, msg):
    shots = [f"Message: {m}\nCategory: {c}" for m, c in examples]
    return "\n\n".join([RULES] + shots + [f"Message: {msg}\nCategory:"])


for msg in MESSAGES:
    none = ask(prompt([], msg))
    four = ask(prompt(EXAMPLES, msg))
    reverse = ask(prompt(EXAMPLES[::-1], msg))
    print(f"\n{msg}")
    print(f"  none: {none.strip()!r:<12} four: {four.strip()!r:<12} reversed: {reverse.strip()!r}")

The Lab Report

A real terminal recording of python fewshot.py report. For llama3.2:3b on Apple M4, 24 GB, lesson 1's prompt C plus examples, 40 messages, 10 per category: examples, right, usable, and how often each of billing, delivery, returns, account and other was answered. Zero 20 of 40, answered 3, 5, 30, 2, 0; balanced 33, 17, 10, 10, 3, 0; balanced-rev 28, 3, 5, 19, 13, 0; balanced-b 36, 10, 11, 11, 8, 0; all-billing 22, 4, 9, 26, 1, 0; all-account 34, 9, 12, 12, 7, 0. For qwen2.5:3b: zero 36, balanced 39, balanced-rev 39, balanced-b 38, all-billing 37, all-account 36. Beneath: the lab's own report, you do not need to run it.

The lab is scripts/labs/prompting/fewshot.py. It imports lesson 1's rules, test messages and scoring, so every lesson in this chapter measures the same task. It stores one results file per model, and the report mode prints both. A separate odds mode asks for the first token's probabilities on the three demo messages under three of the prompts, which is where the numbers on the odds slide come from, and a choice mode adds set B to an existing results file. In the report, "balanced" is set A, "balanced-rev" is set A reversed, and "balanced-b" is set B. The "usable" column equals "right" in every row, because every reply was a single category word. The columns on the right count how often each category was answered; the "other" column would count replies that named no category, and it stayed at 0: with the format rule in place, both models always answered with one of the four words.

A sequence diagram with three columns: the lab, the model and the check. Step one, the lab sends the model rules, examples and a message. Step two, the model returns a category. Step three, the lab sends the check the answer and the right category. Step four, the check returns right, and which one. Beneath: 480 calls, 6 prompts times 40 messages times 2 models.

Two brand cards headed the tools, with their logos, titled what the test ran on. Ollama: qwen2.5:3b and llama3.2:3b, Apple M4, 24 GB. Python: 480 chat calls, 6 prompts, 40 messages, 2 models.

Count the Mistakes in Your Browser

This box has no model. It holds llama's real answers from the lab, one letter per message, and counts the right answers and the most common mistakes for each version of the prompt.

It prints the same numbers as the lab: 20 right with no examples, where the most common mistakes are account answered as returns (8) and billing answered as returns (7); 33 with set A, whose mistakes are all account answered as billing (7); and so on. Try changing most_common(2) to most_common(4) to see every kind of mistake, or count how many times each letter appears in a line to see which answer the model preferred.

The Code, Part by Part

The rules. RULES is lesson 1's prompt C: the instruction and the format rule, with no context lines.

The examples. EXAMPLES is set A, four pairs of a message and its category. They are messages that are not in the test set.

Building the prompt. prompt writes each example as "Message: ...", a new line, then "Category: ..." and the answer. It puts a blank line between the rules, each example and the real message. It ends with "Category:", so the model's next word is the answer. EXAMPLES[::-1] is Python for the same list in reverse order.

The call. ask sends the whole prompt as one user message to llama3.2:3b at temperature 0, so it always takes the most likely token and, on one machine, the same prompt gives the same answer. It allows up to 20 tokens, plenty for one word.

How to Choose Examples

A flowchart headed choosing examples, titled which examples, and how to know. Test with no examples, then read the mistakes, then add one example per category, then try a second set and a second order, then: same result? Yes: keep it, then test on new messages. No: the prompt is fragile, add context or test more. Beneath: one example set, tested once, is one data point.

Start with no examples, and read the mistakes. Llama's mistakes with no examples pointed straight at the problem: almost everything became "returns". That tells you what the examples need to show.

Add one example per category. A balanced set shows every answer in use. In this lab, a balanced set helped both models; three examples with the same answer did not help reliably.

Try a second set and a second order. This is the step most people skip, and the one this lab says matters most. If a second set or a second order gives a similar score, the prompt is stable. If it moves a lot, as llama's did (28 to 36), the prompt is fragile, and the examples are covering up a weak prompt instead of fixing it. Clearer instructions or context lines, as in lesson 1's prompt D, may be the real fix.

Test on messages the examples never saw. Examples written by looking at your own test messages can make the score look better than it will be on new ones. Keep a set of test messages that you never use for examples.

A sketch of four stacked boxes joined by arrows, headed sketched, building a few-shot prompt, titled rules, examples, message. The rules. Message: example 1, Category: billing. ... one example per category .... Message: the customer's text, Category:. Beneath: the examples use exactly the format you want back.

A table headed what the examples did in this lab, titled examples, measured. Count: none to four, 20 to 33 right on llama3.2. Choice: set A against set B, 33 against 36. Order: set A forwards and reversed, 33 against 28. Strong model: qwen2.5, 36 to 39. Beneath: measured once per prompt, on two small models.

When to Use Examples, and When Not To

Use examples when the model gets the task wrong in a way you cannot easily describe. If you can write the fix as a sentence, such as "refunds of money are billing", a context line says it directly, and in lesson 1 context lines took qwen to 40 of 40 (llama was not tested with them). Examples are for patterns that are easier to show than to say.

Use examples when the format is hard to describe. A complicated output shape is often easier to show once than to explain in words.

Do not rely on one set in one order. On the weaker model here, the same four examples in a different order cost 5 right answers. Test at least two sets or two orders before you trust a score.

Weigh the gain against the cost when a prompt already works well. On qwen, set A fixed 3 of the 4 remaining mistakes but added about 65 tokens to every call. Whether that is worth it depends on how costly each mistake is, and every example is more text to keep correct.

Keep examples out of your test set. An example that is also a test message tells you nothing about how the prompt does on new messages.

What This Lab Can and Cannot Tell You

A two-column page headed read before you quote a number from this lesson, titled what was measured, and what was not. Measured: two small models, once; two sets of four examples; two orders of one set; two sets of three with one shared answer. Not measured: more examples than four; many random orders; a fresh set of test messages; examples written blind.

This lab ran each prompt once, at temperature 0, on two small models. It tried two sets of four examples, two orders of one of them, and two sets of three that share an answer. Set B was written after I had seen set A's results. It did not try more than four examples, many random orders, or a second set of test messages. With only a few versions of each, the numbers show that choice and order can matter a great deal on a weaker model, but not how much on average, or which order is best in general. A proper answer to that would run many random orders and report the spread, which is the kind of test a later lesson in this chapter builds.

What to Do Next

A hand-drawn list headed for your own prompt, titled four things to do. 1, start at zero: measure the prompt with no examples. 2, aim: pick examples like the mistakes, then score on cases you did not read. 3, shake: try another set and another order. 4, hold out: test on messages the examples never saw. Beneath: if the order changes the score, you have learned something about the prompt.

Take a prompt of yours that uses examples, or one that could. Score it with no examples on 20 cases and count where the answers go, not only how many are right. Add one example per category, chosen from the kind of case it got wrong. Then score it on a second set of 20 cases that you did not look at while choosing. Swap the order of the examples and score again, and swap in a different set and score again. If the scores are close, keep the examples. If they are far apart, the examples are holding up a prompt that does not really explain the task, and the better fix is usually clearer instructions.

A closing card headed to keep, titled which examples, which order. In large type: 20 → 28 to 36. Beneath: right of 40 on llama3.2:3b, no examples, then three prompts with four examples. Then: test more than one set.

Knowledge Check

Knowledge Check

4 questions - Score 80% to pass

Q1

With no examples, llama3.2:3b got 20 of 40 right and answered 'returns' 30 times. What does counting the answers show that the score alone does not?

Q2

The same four examples, reversed, took llama from 33 to 28 right. What is the lesson?

Q3

On qwen2.5:3b, examples moved the score from 36 to at most 39. On llama3.2:3b, from 20 to as much as 36. Why the difference?

Q4

Why must example messages never also be test messages?

Every prompt went to both llama3.2:3b and qwen2.5:3b at temperature 0: 6 prompts times 40 messages times 2 models is 480 calls. Each reply is scored right or wrong, and the lab also counts which category each reply named, to see where the mistakes went.

This is a real run in VS Code's terminal.

A real screenshot of VS Code's terminal after running python fewshot_demo.py. Please explain the two separate amounts on my statement: none returns, four billing, reversed account. The courier left the box at the wrong address: none returns, four delivery, reversed returns. Please delete my account and all my data: none returns, four billing, reversed account.

The first message is a billing question, and it got three different answers: "returns" with no examples, "billing" with set A, and "account" with set A reversed. The same message, the same rules, the same four examples: only their order changed the last answer. The courier message was right only with set A. The account message was right only with set A reversed. On my laptop these matched the lab's stored answers for the same messages; on your computer, small differences are possible.