Prompting

Instructions Around a Long Document: Where to Put the Rules

0 of 20 complete

0%

Contents

Back|PromptingInstructions Around a Long Document: Where to Put the Rules
1/20
52 min left
Prerequisites
Small Rewordings: The Same Instruction, Seven Waysrequired
Related Topics
The Context Window: What Happens When a Prompt Does Not FitHow Models GenerateThe Silent Cut: How Much of Your Text an Embedding Model Really ReadsTokens and EmbeddingsChat Templates: The Text a Conversation BecomesHow Models GenerateContext Rot: Measuring How Long Context Degrades Your AgentAgents in ProductionWhat a Long Prompt Really Costs: Wait, Writing Speed and MemoryHow Models Generate
1 of 20

A Long Printout and a Short Question

Imagine a colleague walks over with a very long printout. Somewhere on it is a short note: "Which department should this customer letter go to?" If the note is clipped to the very top, you read it first and then skim the rest. If it is buried halfway down, you might read two thousand words before you find out what you are meant to do. If it is at the very bottom, you read everything, and the question is the last thing on your mind when you answer.

A person copes with all three, because a person can go back and look again. A language model reads the whole printout once, in order, and then writes its answer straight after the last word it read. So it is fair to ask whether the place you put your instructions, relative to a long document, changes the answer.

An illustration of a man in a room with a bookcase and a window, pulling a very long strip of paper out of a printer on a desk; the paper runs across the floor in loops. Under the heading the same rules, with a long document in four places, titled where the document sits. Beneath: 2,000 words of unrelated text: nowhere, before, between, after. Usable answers of 40: qwen2.5 40, 40, 40, 36; llama3.2 36, 38, 35, 34.

This lesson measures it. I took the sorting prompt from lesson 1, pasted a long piece of text into it in four different places, and sent it with the same 40 customer messages to two small models. Then, because the answer turned out to be "not much, in this test", I show the case where placement matters a great deal: a document too long for the window.

Five Words for This Lesson

A hand-drawn list headed five words for this lesson, titled a long document in a prompt. Document: a long text pasted into the prompt, here one that has nothing to do with the task. Placement: where the document sits: before, between or after the rules and message. Control: the same prompt with no document, to compare against. Window: num_ctx: the most tokens the model reads in one request. Silent cut: a prompt too long for the window, shortened with no error. Beneath: usable, the reply is exactly the right word, as in lesson 1.

Document. A long text pasted into a prompt. In real apps it is often a help article, a contract or a web page that the model should use. In this lab it is on purpose a text that has nothing to do with the task, so that any change in the answers comes from its length and position, not from what it says.

Placement. Where the document sits in the prompt. There are three parts that stay the same (the rules, the message, and the word "Category:" that asks for the answer) and the document moves around them.

Control. The same prompt with no document. Every placement is compared with it, message by message.

Window. The most tokens a model reads in one request, prompt and reply together. Ollama calls it num_ctx. The previous chapter's lesson 7 measured it.

Silent cut. When a prompt is longer than the window, Ollama shortens it and still returns a normal reply, with no error. The same lesson 7 showed which part it keeps.

As in lesson 1, a reply is usable when the whole reply, in lower case and without spaces at either end, is exactly the correct category word.

The Lab: One Document, Four Places

An editorial frame labelled top of the prompt to bottom, headed one prompt, three parts, four orders, titled what moved, and what stayed. None, the control: rules, then Message: and Category:. Before: document, then rules, then Message: and Category:. Between: rules, then document, then Message: and Category:. After: rules, then Message: and Category:, then document. Beneath: Rules: lesson 1's prompt D. Document: 2,000 words of this course's system design text.

The rules are lesson 1's prompt D: the instruction that names the four categories, the format rule that asks for the category word only, and the four context lines that say what belongs in each category. I chose prompt D because it is the strongest prompt in the chapter so far. On qwen it scored 40 of 40 in lesson 1, so any loss would show.

The document is 2,000 words of this course's own text, taken in order from 17 short sections on system design: CQRS, , consensus, health checks, fault tolerance, several kinds of backup, and so on (you do not need to know what these are; that is the point). It is introduced with one line, "Here is a document from our knowledge base:". It says nothing about bills, parcels, returns or passwords. The lab simply stops at the 2,000th word, so the document ends in the middle of a sentence, "roughly 60,000".

There are four prompts:

  • none, the control: the rules, then "Message:" with the customer's text and "Category:" on the next line.
  • before: the document first, then the rules, then the message.
  • between: the rules, then the document, then the message.
  • after: the rules, then the message and "Category:", then the document.

Each prompt went with each of lesson 1's 40 messages to qwen2.5:3b and llama3.2:3b through Ollama's chat endpoint (the /api/chat address used in every lesson here), at temperature 0 (the model always takes its most likely token), with room for 30 written tokens and a window of 8,192 tokens. Four prompts, 40 messages and two models make 320 calls. The design, including the window size, is written in the lab file's own description, dated before the run. Scoring is lesson 1's, unchanged.

The Scores: Small Moves, Both Ways

A bar chart headed usable answers of 40, four placements, two models, titled small moves, both ways. For none, before, between and after, two bars each: qwen2.5:3b 40, 40, 40, 36; llama3.2:3b 36, 38, 35, 34. Beneath: the same numbers, in the order shown.

On qwen2.5:3b, the control scored 40 of 40, as it did in lesson 1. With the document before the rules, 40. With the document between the rules and the message, 40. With the document after the message, 36. So on qwen, 2,000 words of unrelated text changed nothing with the document before or between; with it last, 4 answers were lost (the next slides ask whether that is more than luck).

On llama3.2:3b, the control scored 36. With the document before, 38; between, 35; after, 34. Read that list carefully. The highest score, 38, came with a document, and the control sits in the middle of the range. If the document were simply a distraction that makes the model worse, the control would be the best score. It is not.

So the first honest summary is that, in this test, the placement of a long unrelated document moved the score by a few answers, and not always in the same direction. The next slides look at which answers moved, and ask whether any of these moves is bigger than luck.

One thing makes the comparison fair: the control repeats earlier runs exactly. Qwen's 40 control replies are exactly the same as prompt D's replies in lesson 1, and llama's 40 are identical to its direct answers with prompt D in lesson 7, although both earlier runs used a smaller window. So the changes below come from the document, not from a different machine state.

Qwen: The Four Answers Lost With the Document Last

Two panels headed qwen2.5:3b, the same 40 messages, titled four answers lost with the document last. Left, no document: 40 of 40 usable answers. Right, document after: 36 of 40 usable answers. Beneath: 0 fixed, 4 broken, sign test p = 0.125: all one way, but too few to rule out luck.

Qwen's four lost answers are worth reading one by one, because a score of 36 hides what kind of mistake was made.

A table headed qwen2.5:3b, the document after the message, titled the 4 answers it lost. Where do I print the label to send an item back?: returns wanted, delivery given. Do I have to pay postage to return an item?: returns wanted, billing given. I returned two items but the return shows only one.: returns wanted, billing given. I cannot sign in with Google any more.: account wanted, signing in given. Beneath: right under the other three placements. 3 are wrong categories; 1 is not a category word.

Three of the four are returns messages that came back as another category. The postage question came back as billing, which a person can understand: postage is money. The label to print came back as delivery, and a label is also something you see on parcels. The two items that were returned but show as one came back as billing. None of these three is a strange answer. Each is a message that could belong to either of two categories, and here it went to the other one.

All four messages were right in the control and right with the document before or between. With the document after, qwen fixed nothing and broke 4. That looks one-sided, and it is. But the sign test from lesson 7 asks how often luck alone would give a split this lopsided. Four changes all going one way happens by chance with a probability of 2 × 1/16 = 0.125, about one time in eight (the p value is that chance; the 2 is there because all four could also have gone the other way). That is not rare enough to call it a real effect. With 40 messages and a model that was already perfect, four is simply too few changes for this test to decide.

An Answer Copied From the Rules

A table headed qwen2.5:3b, I cannot sign in with Google any more., titled it copied a phrase from the rules. The rule: account: signing in, passwords, profile details, privacy and emails from us. None: account. Before: account. Between: account. After: signing in. Beneath: right meaning, wrong form: the code wants the word account.

The fourth lost answer is a different kind of mistake. For "I cannot sign in with Google any more.", qwen replied "signing in". That is not one of the four categories. It is a phrase copied from the context line for account: "account: signing in, passwords, profile details, privacy and emails from us."

In a way the model understood the message: signing in is exactly what the customer is talking about, and the rule that mentions signing in is the account rule. But it gave the rule's words instead of the rule's name. For a person reading the answer, "signing in" is clear. For the code that checks whether the reply is one of four exact words, it is a failure, just like a wrong category.

The other three placements gave "account" for the same message. So the document placed after the message did not make qwen misread this customer. It changed the form of the answer. Lesson 1 showed how much the format rule matters for exactly this reason: adding it took qwen from 19 to 36 usable answers, because a reply that means the right thing but has the wrong shape is still unusable. With the document last, the format rule is about 2,000 words before the point where the model writes its answer. That distance might matter. That is a guess, and the next slides say why I cannot test it with this data.

Llama: The Mistakes Moved Around

A two-column page headed llama3.2:3b, every message that was wrong in at least one placement, titled llama: the mistakes moved around. Left, message; right, wrong under. Why is there a fee of 4.99 on my invoice that I never agreed to?: none (returns). My card was declined but the money still left my bank.: between (account), after (account). The discount code did not apply and I paid full price.: after (account). Please explain the two separate amounts on my statement.: before (account), after (account). Do you accept payment in instalments?: between (account). I need a refund of the shipping charge you added by mistake.: none (returns), between (returns), after (returns). Can you send my order to my office instead of my home?: between (account). Only one of my three items has arrived so far.: after (returns). The driver never rang the bell and took the parcel back.: none (returns), between (returns), after (returns). Do I have to pay postage to return an item?: before (delivery). Please delete my account and all my data.: none (returns). Beneath: 11 of 40 messages; wrong in all four: 0.

On llama the picture is noisier. Across the four prompts, 11 of the 40 messages were wrong at least once, and 29 were usable every time. Not one message was wrong in all four. The mistakes moved around.

Without a document, llama's 4 mistakes were all the answer "returns": the fee on the invoice, the refund of the shipping charge, the driver who took the parcel back, and deleting an account. Lessons 4, 7 and 8 met a much stronger form of this with prompt C (about 30 "returns" answers); with the context lines here, only 4 remain, leaning the same way.

With the document in place, a different mistake appeared. Seven wrong answers across the three placements were "account" given to a billing or delivery message: the declined card, the discount code, the two amounts on a statement, payment in instalments, and sending an order to an office. The control never gave "account" wrongly. Meanwhile, with the document before the rules, llama fixed all four of its "returns" mistakes, which is why that placement scored 38.

I noticed this swap between "returns" and "account" only after reading the replies, and one run cannot tell why it happened. What the data does show is that the document changed which messages llama got wrong, more than how many.

Is Any of This More Than Luck?

A table headed each placement against no document, message by message, titled none of the moves beats luck. Before: qwen2.5 40 to 40: 0 fixed, 0 broken, p = 1. llama3.2 36 to 38: 4 fixed, 2 broken, p = 0.688. Between: qwen2.5 40 to 40: 0 fixed, 0 broken, p = 1. llama3.2 36 to 35: 2 fixed, 3 broken, p = 1. After: qwen2.5 40 to 36: 0 fixed, 4 broken, p = 0.125. llama3.2 36 to 34: 2 fixed, 4 broken, p = 0.688. Beneath: no p value is below 0.05. llama3.2's control, 36, sits inside its own spread of 34 to 38.

Lesson 8 showed that a single score is one sample from a spread, and that the fair way to compare two prompts is message by message: count the messages each one fixed and broke, and run the sign test. Here every placement is compared with the control.

For llama, the document before the rules fixed 4 messages and broke 2 (p = 0.688). Between, it fixed 2 and broke 3 (p = 1). After, it fixed 2 and broke 4 (p = 0.688). A p value near 0.7 or 1 means that luck alone would give a split like this more often than not. Llama's control score, 36, sits inside the range of the three document scores, 34 to 38. That is exactly what you would expect if the document made no real difference and the rest is noise.

For qwen, before and between changed nothing at all. After broke 4 and fixed none, p = 0.125, the closest any result came, and still not below the usual line of 0.05.

So the careful conclusion is this. On these 40 messages and these two models, I cannot tell any placement of this document apart from no document at all. The one hint worth remembering is that the only placement that cost qwen anything was the one that put the document last. It is a hint, not a finding. To turn it into a finding you would need many more messages, or more runs of different documents, and this lab has neither.

Why Placement Could Matter

A sketch of four stacked boxes on paper, headed sketched: a 4,000-word document, a 4,096 window, titled what a cut would keep. About 5,258 tokens, cut to 2,050: first 4, last 2,046. Before: the rules and message survive. Between: the message survives, the rules do not. After: only the document's end survives. Beneath: the rule from How Models Generate, lesson 7. The demo on this page shows it for real.

It is still worth thinking about why the order could matter, as long as I label these as possible reasons and not as what the lab showed.

The last text before the answer. A model writes its answer one token at a time, straight after the last token of the prompt. In the control, and with the document before or between, the last thing in the prompt is "Category:", which says: the next word is a category. With the document after, the last thing is the end of the document, in the middle of a sentence ("...roughly 60,000") in a section about hot backups. "Category:" is still there, but two thousand words further back. One possible reason for qwen's four losses is that the model had to reach much further back to know what kind of word to write.

The reminder in lesson 6. Lesson 6 found something similar with a customer's text that tried to take over the prompt. Putting the task and the format rule again after the customer's text, a reminder, took qwen from 0 to 40 usable answers under the attack. That lesson could not separate "being last" from "being said twice", and neither can this one. But both results point the same way: the words closest to the answer seem to carry the most weight.

Distance from the rules. With the document between, the rules are 2,000 words above the message. Qwen did not mind at all here, and llama's changes were within luck. So this lab gives no support to the idea that distance alone hurts, at this length.

None of these reasons was tested directly. The lab changed where the document sat, and the only placement with even a hint of harm is the one that moved "Category:" away from the end.

Every Prompt Fitted, on Purpose

A sketched bar chart headed qwen2.5:3b, median prompt tokens, against the windows, titled the document was 95% of the prompt. No document: 138. With document: 2,707. Default window: 4,096. The lab's window: 8,192. Beneath: the document added 2,569 tokens to 138. The longest prompt was 2,716, 33% of the window: nothing was cut.

The lab set the window to 8,192 tokens for one reason: to make sure the whole prompt was read, so that the test measured placement and not a cut. The lab stored prompt_eval_count, the number of prompt tokens the model actually read, for every call, so this can be checked instead of assumed.

On qwen, the control prompt had a median of 138 tokens (the median is the middle value of the 40 calls). With the document, it had 2,707. The document added 2,569 tokens with the document before or between, and 2,568 after (one token fewer, where the two parts join), for every message, which is what you expect if nothing was cut: a cut would have removed a different amount from prompts of different lengths. The longest prompt of all was 2,716 tokens, 33% of the window. Llama's counts were a little smaller, 134 and 2,657, because its tokenizer (the part that splits text into tokens) splits text differently.

Four isometric blocks headed qwen2.5:3b, tokens, drawn to scale, titled the prompt, and two windows. No document: 138 tokens, a thin slab. With document: 2,707 tokens. Default window: 4,096 tokens. Lab window: 8,192 tokens, the tallest. Beneath: height is tokens. A 3,000-word document would come to about 3,978, just under the default window; 4,000 words, about 5,258, would not fit.

Notice how the document dominates. 2,569 of 2,707 tokens, about 95% of the prompt, is text that has nothing to do with the task. The rules and the message, the parts that decide the answer, are the other 5%. In a real app the document is usually the part you do not control, and it is usually the part that grows.

The Real Danger: A Document That Does Not Fit

A two-column page headed qwen2.5:3b, this lab's own token counts, titled does the document fit? Left, what; right, tokens. Rules and message: 138 tokens. 2,000 words of document: 2,569 tokens. Tokens per word: 2,569 ÷ 2,000 = 1.28. 3,000 words, window 4,096: about 3,978: fits, about 100 to spare. 4,000 words, window 4,096: about 5,258: cut to 2,050. Beneath: the lab used 8,192; estimates, not measured.

The previous chapter's lesson 7 measured what Ollama does when a prompt does not fit: it returns a normal reply, with no error, after cutting the prompt to half the window plus 2 tokens, keeping the first 4 tokens and the end. On my laptop, a request that does not set num_ctx gets a window of 4,096. That lesson also found that the chat endpoint, with one long message as here, cuts the same way.

Here is the arithmetic with this lab's own counts. The document came to 2,569 tokens for 2,000 words, which is 2,569 ÷ 2,000 = 1.28 tokens per word, counting the one-line introduction. The rules and a message took 138.

  • A 3,000-word document at the default 4,096 window. About 138 + 1.28 × 3,000 ≈ 3,978 tokens. That fits, with about 100 tokens to spare (fewer for the longest messages), and the reply needs up to 30 of them. It is right at the edge: roughly 70 more words of document and it would not fit, and lesson 7 did not measure exactly where the line between "fits" and "cut" sits.
  • A 4,000-word document at the default window. About 138 + 1.28 × 4,000 ≈ 5,258 tokens. That does not fit. By lesson 7's rule, Ollama would keep 4,096 ÷ 2 + 2 = 2,050 tokens: the first 4, then the last 2,046.

These two lines are estimates from this lab's counts, not measurements; your text will have its own tokens per word.

Now the order matters enormously. With the document before the rules, the last 2,046 tokens are the end of the document, then the rules and the message, so the rules and the question survive. With the document between, the message survives but the rules are gone. With the document after, the last 2,046 tokens are all document: the rules and the customer's message are both gone. The model is left answering a question it was never shown. The demo on the next slide shows exactly this.

Try It Yourself

This script sends two of lesson 1's messages to qwen2.5:3b with lesson 1's rules. Each message goes three ways: with no document, with a long document before the rules, and with the document after the message. It does this twice, once with a window of 8,192 tokens, where everything fits, and once with 4,096, where the document does not.

A real screenshot of VS Code with longdoc_demo.py open, lines 1 to 33 visible. It defines RULES, lesson 1's prompt D: the instruction, the format rule and the four context lines; PARAGRAPH, a short paragraph about backups; DOCUMENT, a one-line introduction and that paragraph repeated 60 times; MESSAGES, two messages with the category I labelled them with: I cannot sign in with Google any more, account, and Do I have to pay postage to return an item?, returns; a function ask that sends a prompt to qwen2.5:3b at temperature 0 with up to 30 tokens and a given num_ctx, and returns the prompt tokens read and the reply. The function build, which places the document nowhere, before the rules or after the message, comes after line 33, in the box on the slide. Beneath: copy it from the box on the slide.

Before you run this lab. It uses qwen2.5:3b, running in Ollama on your own computer. If you have not set that up yet, the lab setup guide shows how to install Ollama, download the model and check that everything works, on macOS, Windows or Linux. You can use a different model instead: the guide shows the one line to change, and your numbers will differ from the ones in this lesson.

"""Lesson 1's rules and two customer messages, with a long unrelated document in three places and two windows.

This one uses qwen2.5:3b. Pull it first (see the lab setup guide):
    ollama pull qwen2.5:3b
    python longdoc_demo.py
"""
import json
import urllib.request

RULES = """Put this customer message into one of these categories: billing, delivery, returns, account.

Answer with the category word only, in lower case, and nothing else.

billing: charges, payments, invoices, prices and refunds of money.
delivery: shipping, couriers, tracking and parcels that have not arrived or arrived damaged.
returns: sending an item back, exchanges and the return process.
account: signing in, passwords, profile details, privacy and emails from us."""

PARAGRAPH = ("A backup is a copy of data kept somewhere else, so that it can be restored after a failure. "
             "A full backup copies everything. An incremental backup copies only what changed since the last "
             "backup. A snapshot records the state of a disk at one moment, and a hot backup is taken while "
             "the system keeps serving traffic. Each kind trades storage space against the time a restore takes. ")
DOCUMENT = "Here is a document from our knowledge base:\n\n" + PARAGRAPH * 60     # about 4,100 words, unrelated

MESSAGES = [("I cannot sign in with Google any more.", "account"),
            ("Do I have to pay postage to return an item?", "returns")]


def ask(text, num_ctx):
    body = {"model": "qwen2.5:3b", "stream": False, "messages": [{"role": "user", "content": text}],
            "options": {"temperature": 0, "num_ctx": num_ctx, "num_predict": 30}}
    req = urllib.request.Request("http://localhost:11434/api/chat", data=json.dumps(body).encode(),
                                 headers={"Content-Type": "application/json"})
    d = json.loads(urllib.request.urlopen(req).read())
    return d["prompt_eval_count"], d["message"]["content"].strip()


def build(place, msg):
    ask_part = f"Message: {msg}\nCategory:"
    parts = {"none": [RULES, ask_part], "before": [DOCUMENT, RULES, ask_part], "after": [RULES, ask_part, DOCUMENT]}
    return "\n\n".join(parts[place])


print(f"document: {len(DOCUMENT.split())} words")
for num_ctx in (8192, 4096):
    print(f"\nnum_ctx {num_ctx}")
    for msg, gold in MESSAGES:
        for place in ("none", "before", "after"):
            kept, reply = ask(build(place, msg), num_ctx)
            cut = "  CUT" if kept == num_ctx // 2 + 2 else ""
            print(f"  {msg[:22]:<23} {place:<7} {kept:>5} tokens{cut:<5} {gold:<8} got {reply.splitlines()[0][:30]!r}")

The Lab Report

A real terminal recording of python longdoc_report.py. For qwen2.5:3b, lesson 1's prompt D plus 2,000 words of unrelated text, num_ctx 8,192, a table of document, right, usable, tokens, all stop, and fixed, broke and sign test against none: none 40, 40, 138, 40/40, the control; before 40, 40, 2707, 40/40, 0, 0, p = 1.000; between 40, 40, 2707, 40/40, 0, 0, p = 1.000; after 36, 36, 2706, 40/40, 0, 4, p = 0.125. Usable in all 4 placements: 36; wrong in at least one: 4. The control repeats lesson 1, prompt D: 40 of 40 replies identical. Then the four wrong replies under after: the label to print, delivery; postage, billing; two items returned, billing; sign in with Google, signing in. For llama3.2:3b: none 36, 36, 134, the control; before 38, 38, 2657, 4, 2, p = 0.688; between 35, 35, 2657, 2, 3, p = 1.000; after 34, 34, 2656, 2, 4, p = 0.688. Usable in all 4 placements: 29; wrong in at least one: 11. The control repeats lesson 7, prompt D, direct: 40 of 40 identical. Then all 17 wrong replies, by placement. Last, the window for qwen2.5:3b: longest prompt 2,716 of 8,192 tokens (33%), nothing cut; the document added 2,569 tokens, 1.28 per word; rules and message 138; 3,000 words at the default 4,096: about 3,978 tokens, fits, about 100 to spare; 4,000 words: about 5,258 tokens, cut to 2,050 (first 4 + last 2,046). Beneath: the lab's own report, you do not need to run it.

The measurement is scripts/labs/prompting/chapter_batch.py, mode longdoc. It is one of three labs for this chapter that I designed together and wrote down, in the file's own description, before any of them ran. It imports lesson 1's rules, messages and scoring, takes the first 2,000 words of the course text, builds the four prompts, and stores every reply and its prompt token count in one results file per model. It never reports a time, because other labs were sharing Ollama while it ran.

The report is a separate file, scripts/labs/prompting/longdoc_report.py. It reads the two results files and never calls a model. For each model it prints each placement's right and usable scores, the median prompt tokens, how many replies stopped on their own, and the fixed and broken counts and sign test against the control. Then it prints how many messages were usable everywhere, the check that the control repeats the earlier lessons, every wrong reply, and the window arithmetic.

I want you to know which parts came after I saw the data. The four placements, the document, the window, the messages and the scoring were all fixed before the run. The report was written afterwards, and so were its choices: the sign tests against the control, the list of wrong replies, the repeat check, and the window arithmetic. The swap between "returns" and "account" on llama was found by reading the replies. The demo script, with its repeated paragraph and its 4,096 window, was also written after the run, to show the cut that the lab avoided on purpose.

Compare Placements in Your Browser

This box has no model. It holds the lab's real answers for all four placements on both models, one letter per message, with "?" for a reply that is not one of the four words. It prints each placement's usable score and its fixed, broken and sign test against the control. Then it applies lesson 7's cut rule to a document length and window that you choose, using this lab's own token counts.

It prints the same numbers as the lab: qwen 40, 40, 40 and 36 with p = 0.125 for the document after, and llama 36, 38, 35 and 34 with no p value below 0.05. At the bottom, with a 4,000-word document and a 4,096 window, it estimates about 5,258 tokens, 2,050 kept, and says which parts survive for each placement. Try DOC_WORDS = 3000, which just fits, or set NUM_CTX = 8192 and watch the cut disappear. Or change one "?" in qwen's "after" row to the right letter and see how far the p value moves.

The Code, Part by Part

The rules. RULES is lesson 1's prompt D, word for word: the instruction, the format rule and the four context lines, separated by blank lines as in the lab. Keeping it identical is what lets the demo's control replies be compared with the lab's stored rows.

The document. PARAGRAPH is a short paragraph about backups, unrelated to customer support. DOCUMENT puts the lab's one-line introduction in front of it and repeats it 60 times, about 4,100 words. Repeating one paragraph keeps the script short enough to read. It is a different text from the lab's document, and the lesson says so wherever the two are compared.

The call. ask sends one prompt to qwen2.5:3b through Ollama's chat endpoint at temperature 0, with room for 30 tokens and the window you pass in. It returns two things: prompt_eval_count, the number of prompt tokens the model actually read, and the reply.

The placement. build puts the parts in order: the document nowhere, before the rules, or after the message. The message part always ends with "Category:".

The cut check. If the tokens read equal num_ctx // 2 + 2, the script prints "CUT". That is the same check lesson 7 recommended for your own code, and it is the only sign in the reply that anything was lost.

How to Put a Long Document in a Prompt

A table headed what a long document did in this lab, titled a long document, measured. No document: usable qwen2.5 40, llama3.2 36. With document: qwen2.5 36 to 40, llama3.2 34 to 38. Clear moves: none: every sign test p above 0.05. Prompt size: 2,716 tokens at most, of 8,192: nothing cut. Beneath: one run each, lesson 1's 40 messages, temperature 0.

Make it fit first. Count the tokens of your longest real prompt, document included, with the model's own tokenizer (the Tokens and chapter shows how), and set num_ctx above it with room for the reply. This matters far more than anything else on this page: in this lab, placement moved at most a few answers, while a cut in the demo removed the question itself.

Check what was read. After every call, look at prompt_eval_count. If it equals num_ctx // 2 + 2, or is far below what you expected, the prompt was cut and the reply cannot be trusted. (It will never exactly equal a count of your own text, because it includes the chat template's markers.)

Put the document first, and the rules and the question after it. If the prompt ever does get cut, the end survives, so put what matters most at the end. When it fits, this order also leaves "Category:", or whatever asks for the answer, as the last thing the model reads. In this lab that made no difference I could tell from luck; the reason to prefer it is the cut.

Or repeat the instructions after the document. If your rules must come first, for example in a system message (the separate instructions message from lesson 3), say the task and the format rule again after the document, as lesson 6's reminder did. It costs a few dozen tokens.

Measure on your own task. Score the prompt with and without the document on the same cases, count fixed and broken, and use the sign test before you decide anything.

A flowchart headed before you send a long document, titled fit first, then order. Count the prompt's tokens, then: prompt plus reply under num_ctx? No: raise num_ctx, or send less text, and count again. Yes: document first; rules and question last, then score it on your own test cases. Beneath: then check prompt_eval_count on every reply.

When This Matters, and When It Does Not

It matters most when the prompt might not fit. A chat that grows, a document that users upload, a search step (a step that looks up pages and pastes them into the prompt) that sometimes returns ten long pages instead of three: each can push a prompt past the window without warning. Then the order decides whether the model loses some of the document, the rules, or the question.

It matters when the reply must have an exact form. Qwen's "signing in" was the right idea in the wrong shape. If your code needs one exact word or a strict format, consider repeating the format rule after the document, as lesson 6's reminder did (this lab did not test it), and check every reply in code.

It matters less for a short document that fits, on a clear task. With lesson 1's full rules, 2,000 unrelated words before or between made no difference to qwen at all, and llama's changes were within luck. If your prompt is like this one, reordering will not help much; write test cases instead.

It may matter more with a document that is related to the task. This lab used unrelated text on purpose. A real help article about refunds and passwords could pull answers towards its own words much more strongly, and this lab says nothing about that. Measure it.

Do not trust an order because it worked for someone else. Two small models here already disagreed on which messages went wrong. A different model, a longer document or a different task can change the picture.

What This Lab Can and Cannot Tell You

A two-column page headed read before you quote a number from this lesson, titled what was measured, and what was not. Measured: two small models, one run; one task, 40 messages; one unrelated text, 2,000 words; prompts that fit the window. Not measured: larger models; a document the task needs; longer texts, other lengths; a cut: only in the demo.

This lab ran two small models, once each, at temperature 0, on one task of 40 short messages that I wrote. It used one document of one length, unrelated to the task, and one wording of the rules. With 40 messages and a control that was already near the top, the sign test can only see large differences, so "cannot tell from luck" here means "too small for 40 messages to show", not "no difference". A real effect of two or three answers in forty could be there and this lab would miss it.

The lab did not test a document that the task needs, which is the usual case in real apps, or documents much longer than 2,000 words that still fit, or larger models. It never cut a prompt; the only cut in this lesson is in the demo, with two messages. The arithmetic for 3,000 and 4,000 words is an estimate from this lab's own tokens per word and lesson 7's rule, not a measurement. What carries over is the method: make the prompt fit and check that it did, compare each change with a control on the same cases, and read the wrong replies before you explain them.

What to Do Next

A hand-drawn list headed for your own prompt, titled four things to do. 1, count: count the tokens of your longest real prompt. 2, window: set num_ctx above it, with room for the reply. 3, order: document first; the rules and question after it, or repeat them. 4, measure: score it with and without the document, on your own cases. Beneath: check prompt_eval_count on every reply.

Find the prompt in your own work that carries the longest text: a pasted document, search results, or a long chat history. Count its tokens on the longest real input you have seen, not a typical one, and check that num_ctx is set above that count with room for the reply. Add one line of code that logs a warning when prompt_eval_count equals num_ctx // 2 + 2 or is far below what you expected. Then look at the order. If the document sits after your question, move it before, or repeat the question and the format rule after it. Finally, score the prompt on 20 to 40 of your own cases with the document and without it, count fixed and broken, and see whether the difference is bigger than luck before you change anything else.

A closing card headed to keep, titled fit first, then put the question last. In large type: 2,716 of 8,192. Beneath: tokens in this lab's longest prompt: it fitted, and the placement changed a score by at most 4 of 40. Then: a prompt that does not fit loses far more.

Knowledge Check

Knowledge Check

4 questions - Score 80% to pass

Q1

On llama3.2:3b, the four placements scored 36 (no document), 38, 35 and 34. What is the fairest reading?

Q2

With the document after the message, qwen2.5:3b answered 'signing in' for 'I cannot sign in with Google any more.' Why does this count as a failure?

Q3

A 4,000-word document at Ollama's default 4,096 window, by this lab's counts, is about 5,258 tokens. If the document is placed after the message, what does the model see?

Q4

Which advice does this lesson put first for a prompt with a long document?

This is a real run in VS Code's terminal.

A real screenshot of VS Code's terminal after running python longdoc_demo.py. Document: 4148 words. With num_ctx 8192: I cannot sign in with Google, none 136 tokens, got account; before 4706 tokens, got account; after 4706 tokens, got signing in. Do I have to pay postage, none 138 tokens, got returns; before 4708 tokens, got returns; after 4708 tokens, got billing. With num_ctx 4096: none 136 and 138 tokens, got account and returns; before 2050 tokens, CUT, got account and returns; after 2050 tokens, CUT, got Given the repetitive nature of, for both messages.

With the 8,192 window, the demo's document is about 4,100 words and fits: 4,706 and 4,708 tokens were read. Both messages were right with no document and with the document before. With the document after, the sign-in message came back as "signing in" and the postage question as "billing". Those are the same two wrong replies the lab stored for these two messages with the document after (I chose these two messages from the four qwen got wrong, so matching is a repeat of the lab, not a new finding), even though the demo's document is a different text: a paragraph about backups repeated 60 times, not the lab's 2,000 words of course text. The control replies and their token counts, 136 and 138, are identical to the lab's stored rows. I checked this by comparing the output with the results file.

With the 4,096 window, every prompt with the document was cut to exactly 2,050 tokens, the "CUT" in the output, as lesson 7's rule says. With the document before, both answers were still right, because the rules and the message were at the end and survived. With the document after, both replies began "Given the repetitive nature of", a comment about the document. The model never saw the rules or the question, so it wrote about the only text it had. There was no error in either reply. On your computer, a different Ollama version or chip can change a close choice between two tokens, so a small difference is possible.

A sequence diagram with three columns: the lab, the model and the check. Step one, the lab sends the model rules plus document plus message. Step two, the model returns the reply plus prompt tokens. Step three, the lab sends the check the reply and the right label. Step four, the check returns right? usable?. Beneath: 320 calls: 4 placements × 40 messages × 2 models, temperature 0, window 8,192.

Two brand cards headed the tools, with their logos, titled what the test ran on. Ollama: qwen2.5:3b and llama3.2:3b, Apple M4, 24 GB. Python: 320 chat calls, window 8,192.