Prompting

Prompts as Code: Templates That a Customer Cannot Break

0 of 20 complete

0%

Contents

Back|PromptingPrompts as Code: Templates That a Customer Cannot Break
1/20
56 min left
Prerequisites
Instructions Around a Long Document: Where to Put the Rulesrequired
Related Topics
Chat Templates: The Text a Conversation BecomesHow Models GenerateMarking Untrusted Text: Spotlighting, Measured on Small ModelsAI Security and Agent SafetyDeterministic Scaffolding: The LLM Explains, It Does Not ArbitrateAgents in ProductionWhat a Language Model Actually Outputs: Odds for Every Next TokenHow Models GenerateThe Context Window: What Happens When a Prompt Does Not FitHow Models Generate
1 of 20

A Printed Form With One Empty Box

Think of a printed form at a bank counter. Most of the page is printed once, for everybody: the bank's name, the questions, the small print at the bottom. There is one empty box that says "Write your name here". A thousand customers fill in the same form, and each one writes something different in the box. The bank never reprints the page for each customer. It only reads what is in the box.

This works because the printed part and the box stay separate. Whatever a customer writes in the box, it cannot change the questions printed above it, or the small print below it. Now imagine a strange form where the words you write in the box could reach out and change the printed text: write the right word, and a line from the bank's private instructions appears on your copy. Nobody would design a paper form like that. But it is surprisingly easy to build a prompt like that in code.

An illustration of a classroom. A young woman at the left desk writes on a sheet covered in small sketches; beside her, a young man rests his head on his hand, a pencil at his lip, looking at a nearly empty sheet with one small box on it. Under the heading one template, filled for every customer, titled what goes in the box. Beneath: 47 inputs, four ways to fill a template. format crashed on 47; a two-step format crashed on 2 and leaked a note once; replace and an f-string: 0 and 0.

In every lesson of this chapter, my code built each prompt the same way: fixed words written once, and a gap where the customer's message goes. This lesson looks at that code. I tested four common ways of filling the gap against the kinds of text real customers type, and then measured what happens to the model's answers when the fixed words change by only a space or a line break.

Five Words for This Lesson

A hand-drawn list headed five words for this lesson, titled prompts as code. Template: a prompt written once, with a gap filled in for every request. Placeholder: the marked gap, here the word message in curly brackets. Fill: the code that puts the real text into the gap. Untrusted text: text you did not write, a customer's message, as in lesson 6. Diff: the exact characters that differ between two versions. Beneath: usable, the reply is exactly the right word, as in lesson 1.

Template. A prompt that you write once and use for every request, with a gap where the changing part goes. The fixed part is like the printed form; the gap is the box.

Placeholder. The marker in the template that shows where the gap is. In Python it is often a name in curly brackets, such as {message}.

Fill. The code that puts the real text into the gap. In Python there are several ways to do it, and this lesson tests four of them. The difference between them turned out to matter a great deal.

Untrusted text. Text that you did not write and cannot control, such as a customer's message. Lesson 6 used the same idea: a customer can type anything, including characters that mean something to your code.

Diff. A list of the exact characters that differ between two versions of a text. A person reading two prompts side by side can easily miss a space at the end of a line. A diff made by a program does not miss it.

As in lesson 1, a reply is usable when the whole reply, in lower case and with no spaces at either end, is exactly the correct category word.

The Lab: Two Parts

An editorial frame labelled the lab's PART 1, no model, headed one template, four ways to fill it, titled the same gap, four pieces of code. Four boxes: format, TEMPLATE.format(message=msg); format2, fill the message, then .format() again for the policy note; replace, TEMPLATE.replace with the placeholder and msg; fstring, a function that builds the text with an f-string. Beneath: inputs: lesson 1's 40 messages and 7 a customer could type, with braces, a dollar sign, a backslash, nothing, or 400 words.

The lab has two parts, and I wrote both down in the lab file's own description before running either.

Part 1 needs no model. It uses two templates. format, replace and the f-string fill a short sorting prompt with a JSON example in it, in the style of lesson 5: "Reply as JSON, for example" followed by the example with the category billing, in curly brackets. The four ways are:

  • format: Python's str.format, the most common way to fill a named gap.
  • format2: a second template, with no JSON example and two gaps, the message and an internal policy note. The code fills the message first and leaves the policy gap for later; a second call to format, in another part of the code, fills the policy note. Pipelines built from several steps (a pipeline is a chain of functions, each passing its result to the next) often end up like this.
  • replace: the text method replace, which swaps the exact characters {message} for the customer's text and does nothing else.
  • : a function that builds the whole text with an f-string, Python's way of writing values straight into a string.

format Crashed on Every Input

A table headed 47 inputs per method, counted in code, titled two methods broke, two did not. Format: 47 crashed, 0 changed, 0 leaked. Format2: 2 crashed, 1 changed, 1 leaked. Replace: 0 crashed, 0 changed, leak n/a (one gap). Fstring: 0 crashed, 0 changed, leak n/a (one gap). Beneath: format: 47 of 47, because of the JSON example's braces. format2: only on what a customer typed.

str.format crashed on all 47 inputs, including all 40 ordinary messages. The error was the same every time: KeyError: '"category"'.

The reason is in the template, not in the messages. To str.format, every pair of curly brackets is a gap to fill. The template's JSON example has curly brackets around "category": "billing", so format read that as a gap whose name is "category", looked for a value with that name, found none, and stopped. The customer's text never mattered: the template could not be filled at all.

The usual fix inside format is to double every bracket that is not a gap, so that the example reads {{"category": "billing"}} in the template. That works, and it is also the kind of fix that breaks again the next time someone pastes a new JSON example into the prompt without doubling its brackets. Lesson 5 put JSON examples in prompts on purpose, so this is not a rare situation.

Two panels headed 47 inputs, two templates, titled the loud bug and the quiet one (two different templates). Left, format: 47 of 47, crashed: seen at once. Right, format2: 44 of 47, looked fine. Beneath: format2 broke on 3 inputs, all typed by a customer: 2 crashed, 1 pulled in the internal note.

A bug that breaks every input is the good kind, because you find it the first time you run the code. The next method is the dangerous kind. It worked on 44 of the 47 inputs, including all 40 ordinary messages, and broke only on text that a customer chose to type.

format2: The Customer's Brackets Were Filled Too

The two-step template broke on three of the seven tricky inputs, and on none of lesson 1's messages.

Two of them crashed. "My order {12345} was charged twice." raised IndexError: Replacement index 12345 out of range..., and "Please ignore {0} and {1}, just refund me." raised the same error for index 0. After step 1, the customer's text sits inside the prompt, and step 2 calls format on the whole prompt. To format, {12345} and {0} are gaps that ask for numbered values, and there are none. So the request fails, and the customer gets an error page instead of an answer, because they typed curly brackets. An order number in curly brackets is not an attack; people type things like that.

A two-column page headed format2, one customer message, titled the customer's braces were filled too. Left, step; right, text. The customer typed: What does, the word policy in curly brackets, say about refunds? After step 1, Message: holds the same text. After step 2, Message: holds What does INTERNAL NOTE: refunds above 50 dollars need a manager. say about refunds? The internal note appears 2 times, once inside the message. Beneath: step 2 fills every policy gap, not only yours.

The third input is worse, because nothing fails. The customer wrote "What does {policy} say about refunds?". Step 1 put that text into the prompt as it was. Step 2 then filled every {policy} it could find, and it could not tell the one I wrote from the one the customer wrote. So the internal note, "INTERNAL NOTE: refunds above 50 dollars need a manager.", was pasted into the middle of the customer's message. The model received a message the customer never wrote, containing a line the customer was never meant to see.

What Did Not Break, and Why

replace and the f-string had no crashes and no changes on all 47 inputs, including the brackets, the dollar signs, the backslashes, the empty message and the 400 repeated words. Their template has no second gap, so they could not leak; the fair leak test is replace in two steps on the two-gap template, below, and it leaked just like format2.

They work for the same reason. TEMPLATE.replace("{message}", msg) looks for those exact nine characters in the template and swaps in the customer's text wherever they appear (once, here). It does not read the customer's text at all, so nothing in it can be taken as a gap. An f-string is even stricter: the gaps are written in the Python code itself, so the only values that can be filled are the ones the programmer named. The customer's brackets are just characters inside a value.

I want to be careful not to turn this into "replace is safe and format is not". After the run, I added one check that was not in the lab's design: the same two-step template, filled with replace twice instead of format twice, first the message and then the policy. It leaked in exactly the same way, on the same input, because the second replace also found the customer's {policy}. The problem is not the method. The problem is filling anything after the customer's text is already in the prompt.

So the rule I take from this is short. Fill every gap in one pass that never re-reads what it inserted, and never run a second fill over text that already holds an inserted value. One format(...) call with every value at once works (I checked it on the two-gap template: 0 problems in 47), as long as the template's own brackets are doubled; so does one replace when there is only one gap, or one f-string. What breaks is filling in stages. The same goes for any value you did not write, such as a pasted document from lesson 9: filling in stages can pull the customer's text into it too. The f-string has a trap of its own: its JSON example also needs doubled brackets, and pasting in an undoubled example crashes it just like , and it cannot be kept in a separate template file. Of the four, only needs no escaping in the template itself. If the text will sit between tags, escape it first, as lesson 6 did: to , then to and to , so the customer cannot type your tags either.

Part 2: Changes You Would Not Notice

A table headed what each version sends after the rules, invisible characters in words, titled six versions, a few characters apart. Clean: Message: msg, new line, Category:. Trailing-space: the same, then a space after Category:. No-newline: Message: msg, space, Category:. Raw-input: Message:, 3 spaces, new line, msg, 2 spaces, 3 new lines, Category:. No-label: msg, new line, Category:. Double-rules: Message: msg, new line, Category:, with the rules sent twice above it. Beneath: msg in curly brackets is the customer's message. Median prompt tokens on qwen2.5, in this order: 138, 139, 138, 140, 136, 232.

Part 1 was about code that breaks. Part 2 is about code that works, and still sends slightly different text from what you think. Lesson 8 showed that rewording an instruction moves the answers. Here nothing is reworded. Each version is lesson 1's prompt D, the instruction, the format rule and the four context lines, with a change to the characters around the customer's message that a person reviewing the prompt could easily miss:

  • clean: the reference, "Message:" with the text, then "Category:" on the next line, exactly as in lessons 1 and 9.
  • trailing-space: one space after "Category:". Text editors often add or keep one.
  • no-newline: the line break before "Category:" replaced by a space, as happens when a template is joined onto one line.
  • raw-input: the message with the spaces and blank lines a web form can leave around it: three spaces after "Message:", then a line break, the text, two spaces and three line breaks.
  • no-label: the message with no "Message:" in front of it, as if someone tidied the template.
  • double-rules: the rules pasted twice, the kind of thing combining two people's edits by hand can do.

Each version went with lesson 1's 40 messages to both models through Ollama's chat endpoint at temperature 0, with a limit of 30 tokens for each reply: six versions, 40 messages and two models make 480 calls. Scoring is lesson 1's, unchanged.

The Scores: Small Moves, Mostly on Llama

A bar chart headed usable answers of 40, six versions, two models, titled small moves, mostly on llama. For clean, trailing, no-newline, raw-input, no-label and double, two bars each: qwen2.5:3b 40, 40, 39, 40, 40, 39; llama3.2:3b 36, 37, 36, 39, 34, 33. Beneath: the same numbers, in the order shown.

On qwen2.5:3b, clean scored 40 of 40, and so did trailing-space, raw-input and no-label. No-newline and double-rules each scored 39. Across the five changed versions, only 2 of qwen's 200 replies differed at all from clean, letter for letter. Both were returns messages that came back as delivery: "I sent the lamp back ten days ago, has it reached you?" with no-newline, and "Do I have to pay postage to return an item?" with the rules pasted twice.

On llama3.2:3b, clean scored 36. The other versions scored 37, 36, 39, 34 and 33. Here 18 of the 200 replies differed from clean. The best score, 39, came from raw-input, the version with extra spaces and blank lines. The worst, 33, came from pasting the rules twice.

A sketched bar chart headed replies that were not letter for letter the same as clean, of 40, titled llama's replies moved more. Trailing-space: llama3.2 1, qwen2.5 0. No-newline: llama3.2 4, qwen2.5 1. Raw-input: llama3.2 3, qwen2.5 0. No-label: llama3.2 4, qwen2.5 0. Double-rules: llama3.2 6, qwen2.5 1. Beneath: bar length is llama3.2's count. In all, 18 of 200 replies moved on llama3.2, 2 of 200 on qwen2.5.

Before reading anything into these numbers, one check. Clean is the same prompt that lesson 9 used as its control, and its 40 replies are identical, letter for letter, to lesson 9's on both models. So the machine was giving the same answers to the same text, and any change here came from the characters that changed.

Which Answers Moved on Llama

A two-column page headed llama3.2:3b, every message whose reply moved, titled which answers moved, and to what. Left, message; right, version, clean then new. Why is there a fee of 4.99 on my invoice that I never agreed to?: was returns; trailing-space billing, no-newline billing, raw-input billing, no-label billing. The package arrived crushed and the corner is torn open.: was delivery; no-newline returns, no-label returns. Only one of my three items has arrived so far.: was delivery; no-newline returns. Please delete my account and all my data.: was returns; no-newline account, raw-input account. The driver never rang the bell and took the parcel back.: was returns; raw-input delivery, double-rules delivery. The discount code did not apply and I paid full price.: was billing; no-label returns, double-rules returns. Please explain the two separate amounts on my statement.: was billing; no-label billing and more, double-rules account. My card was declined but the money still left my bank.: was billing; double-rules delivery. I cancelled my subscription but you took another payment.: was billing; double-rules returns. I need a refund of the shipping charge you added by mistake.: was returns; double-rules delivery. Beneath: 10 of 40 messages; the rest never moved.

Only 10 of llama's 40 messages ever moved; the other 30 got the same reply under all six versions. Reading the 10, two patterns stand out, and they are the same two that lessons 4, 7, 8 and 9 kept meeting.

Llama's four mistakes under clean were all the answer "returns": the fee on the invoice, the shipping charge refund, the driver who took the parcel back, and deleting an account. Every one of the 8 fixes across the five versions was one of these four messages turning right. The fee on the invoice was fixed by four of the five versions. So llama is almost undecided on these four messages, and a very small change can switch its answer.

The breaks mostly went the other way: of the 9 breaks, spread over 6 messages, 6 became "returns". The crushed parcel, the partial delivery, the discount code and the cancelled subscription all fell into llama's favourite wrong answer at least once. One break was a different kind: with no label, "Please explain the two separate amounts on my statement." got the right word, "billing", followed by a blank line and a paragraph of explanation, so it was right but not usable.

In short, the small characters did not teach llama anything new. They moved a few messages that were already near an edge, in both directions.

Is Any of This More Than Luck?

A table headed each version against clean, message by message, titled none of the moves beats luck. Trailing-space: qwen2.5 40 to 40: 0 fixed, 0 broken, p = 1. llama3.2 36 to 37: 1 fixed, 0 broken, p = 1. No-newline: qwen2.5 40 to 39: 0 fixed, 1 broken, p = 1. llama3.2 36 to 36: 2 fixed, 2 broken, p = 1. Raw-input: qwen2.5 40 to 40: 0 fixed, 0 broken, p = 1. llama3.2 36 to 39: 3 fixed, 0 broken, p = 0.25. No-label: qwen2.5 40 to 40: 0 fixed, 0 broken, p = 1. llama3.2 36 to 34: 1 fixed, 3 broken, p = 0.625. Double-rules: qwen2.5 40 to 39: 0 fixed, 1 broken, p = 1. llama3.2 36 to 33: 1 fixed, 4 broken, p = 0.375. Beneath: no p value is below 0.05, on either model.

The fair way to compare two prompts on the same messages is the one lesson 7 used: count the messages each version fixed and broke compared with clean, and run the sign test. The sign test asks how often luck alone would split the changed messages this unevenly, if the change made no real difference. The answer is the p value; below 0.05, one time in twenty, is the usual line for calling a difference real.

On qwen, every version's p value is 1: one message broken at most, and one message is never evidence of anything. On llama, the p values are 1, 1, 0.25, 0.625 and 0.375. None is below 0.05. So the honest summary is that on these 40 messages, I cannot tell any of the five versions apart from clean, on either model. Most of these differences are the size that luck produces all the time.

A two-column page headed llama3.2:3b, clean to raw-input, titled is 36 to 39 more than luck? Counted, and the sign test: usable, 36 to 39; messages that changed, 3: 3 fixed, 0 broken; went one way, 3 of 3; chance of that by luck, 2 × 1/8 = 0.25; verdict, cannot tell from luck. Beneath: on the same 40 messages; lesson 7's sign test, p = 0.25.

The case most likely to mislead is raw-input on llama, which went from 36 to 39. It is tempting to read that as advice: leave the spaces and blank lines in, and llama does better. Please do not. Three changes that all go one way happen by luck with a chance of 1/2 × 1/2 × 1/2 = 1/8 in one direction, and either direction counts, so 2 × 1/8 = 0.25: one time in four. On qwen the same version changed nothing. It was also one of five versions, and when you look at five, the best of them will usually look good by luck. My reading is that llama's 39 is most likely noise that happened to fix three messages sitting near llama's edge, not a real effect of extra spaces. The same holds for the losses: double-rules going from 36 to 33 is not enough, on 40 messages, to prove it hurts.

What the data does show is simpler. The characters around a message are part of the prompt, they change what the model reads, and on a model that is unsure about some messages they change some answers. That is the same warning as lesson 8, at an even smaller scale.

Keep the Template in One Place, and Diff What You Send

If a space can change a reply, then two copies of "the same" prompt are not the same prompt unless they match character for character. That leads to three habits, none of them about wording.

One template, one file, under version control. Keep each template in a single place, for example one file or one constant, and keep that file in Git, the version control tool most teams use. Then every change to it has a date, an author and a record of exactly which characters changed. When a score moves, you can find out what changed in the text instead of guessing. A prompt copied into three files will drift apart, one space at a time.

Diff the exact text you send. Do not review the template; review what it produces. Build the prompt for one real message from the old version and the new version, and compare the two strings with a program. Python's difflib does it in a few lines, and printing the end of each string with repr shows spaces and line breaks as visible characters, such as \n for a new line. In the demo on the next slides, no-newline and clean are exactly the same length, so a length check misses the change. The line diff reports that something changed, and repr shows what: the new line became a space.

A sequence diagram with three columns: template file, your code and the model. Step one, the template file sends your code one version, from git. Step two, your code sends the model the text, filled once, and logs it. Step three, the model returns the reply, which is checked. Step four, your code sends the template file a change: diff, test, commit. Beneath: the customer's text goes in at step 2, last, escaped, and nothing is filled after it.

Log the text you actually sent, with the template's version. When a reply is wrong in production, the first question is what the model was really given. If you store the final prompt, or at least the template version and the inputs, you can answer that exactly.

And whenever the template changes, score the new version on your fixed test cases and compare it with the old one message by message, as on the previous slide. A change of a space that moves nothing is fine. A change that moves many messages, even if the total stays the same, deserves a closer look.

Try It Yourself

This script does both parts on a small scale. First, with no model, it fills the lab's template four ways for three tricky messages and prints what happened. Then it builds two of the lab's small changes for two messages, prints the number of changed lines and the last characters of each prompt, and sends both versions to qwen2.5:3b.

A real screenshot of VS Code with templates_demo.py open, lines 1 to 37 visible. It defines TEMPLATE, the sorting prompt with a JSON example; TWO_STEP, the template with a message gap and a policy gap; SECRET, the internal note; FILL, the four ways to fill a template, format, format2, replace and fstring; and TRICKY, three customer messages, then the start of the loop that tries each fill and catches a crash. Beneath: copy it from the box on the slide.

Before you run this lab. It uses qwen2.5:3b, running in Ollama on your own computer. If you have not set that up yet, the lab setup guide shows how to install Ollama, download the model and check that everything works, on macOS, Windows or Linux. You can use a different model instead: the guide shows the one line to change, and your numbers will differ from the ones in this lesson. Part 1 of the script needs no model at all.

"""Four ways to fill a prompt template, three tricky customer messages, and two tiny edits sent to a model.

Part 1 needs no model. Part 2 uses qwen2.5:3b. Pull it first (see the lab setup guide):
    ollama pull qwen2.5:3b
    python templates_demo.py
"""
import difflib
import json
import urllib.request

TEMPLATE = ('Classify the customer message. Reply as JSON, for example {"category": "billing"}.\n\n'
            "Message: {message}")
TWO_STEP = "Classify the customer message.\n\nMessage: {message}\n\nPolicy: {policy}"
SECRET = "INTERNAL NOTE: refunds above 50 dollars need a manager."

FILL = {
    "format": lambda m: TEMPLATE.format(message=m),
    "format2": lambda m: TWO_STEP.format(message=m, policy="{policy}").format(policy=SECRET),
    "replace": lambda m: TEMPLATE.replace("{message}", m),
    "fstring": lambda m: f'Classify the customer message. Reply as JSON, for example {{"category": "billing"}}.'
                         f"\n\nMessage: {m}",
}
TRICKY = ["My order {12345} was charged twice.", "What does {policy} say about refunds?",
          "You took $49.99 instead of $39.99."]

print("Part 1: fill the template (no model)")
for msg in TRICKY:
    print(f"\n{msg}")
    for name, fill in FILL.items():
        try:
            text = fill(msg)
        except Exception as err:          # a crash: the customer gets an error page, not an answer
            print(f"  {name:<8} CRASH  {type(err).__name__}: {str(err)[:40]}")
            continue
        verdict = "ok" if msg in text else "CHANGED"
        if text.count(SECRET) > 1:
            verdict += ", internal note now inside the message"
        print(f"  {name:<8} {verdict}")

RULES = """Put this customer message into one of these categories: billing, delivery, returns, account.

Answer with the category word only, in lower case, and nothing else.

billing: charges, payments, invoices, prices and refunds of money.
delivery: shipping, couriers, tracking and parcels that have not arrived or arrived damaged.
returns: sending an item back, exchanges and the return process.
account: signing in, passwords, profile details, privacy and emails from us."""
VERSIONS = {
    "clean": lambda m: f"{RULES}\n\nMessage: {m}\nCategory:",
    "no-newline": lambda m: f"{RULES}\n\nMessage: {m} Category:",
    "double-rules": lambda m: f"{RULES}\n\n{RULES}\n\nMessage: {m}\nCategory:",
}
TESTS = [("I sent the lamp back ten days ago, has it reached you?", "returns", "no-newline"),
         ("Do I have to pay postage to return an item?", "returns", "double-rules")]


def ask(text):
    body = {"model": "qwen2.5:3b", "stream": False, "messages": [{"role": "user", "content": text}],
            "options": {"temperature": 0, "num_predict": 30}}
    req = urllib.request.Request("http://localhost:11434/api/chat", data=json.dumps(body).encode(),
                                 headers={"Content-Type": "application/json"})
    return json.loads(urllib.request.urlopen(req).read())["message"]["content"].strip()


print("\nPart 2: the exact text, then the model (qwen2.5:3b)")
for msg, gold, other in TESTS:
    a, b = VERSIONS["clean"](msg), VERSIONS[other](msg)
    changed = [d for d in difflib.ndiff(a.splitlines(), b.splitlines()) if d[0] in "+-"]
    print(f"\n{msg}  (labelled {gold})")
    print(f"  clean vs {other}: {len(changed)} changed lines, {len(b) - len(a):+} characters")
    print(f"  the last 24 characters sent: {a[-24:]!r}")
    print(f"  {'':<28} {b[-24:]!r}")
    print(f"  {'clean':<12} got {ask(a)!r}")
    print(f"  {other:<12} got {ask(b)!r}")

The Lab Report

A real terminal recording of python templates_report.py. PART 1: fill one template, 47 inputs, 40 of lesson 1's messages and 7 tricky ones, no model. Method, crashed, changed, note leaked, rerun now: format 47, 0, 0; format2 2, 1, 1; replace 0, 0, 0; fstring 0, 0, 0; each same as stored. Format: every input crashed, 40 of them normal messages, KeyError category. Format2: the order number, IndexError replacement index 12345 out of range; ignore 0 and 1, IndexError replacement index 0 out of range; the policy question was sent as Message: What does INTERNAL NOTE: refunds above 50 dollars need a mana. Added after the run, replace() in two steps: 0 crashed, 1 changed, 1 leaked. PART 2, qwen2.5:3b, against clean: version, right, usable, tokens, replies differ, fixed, broke, sign test: clean 40, 40, 138, 0; trailing-space 40, 40, 139, 0, 0, 0, p = 1.000; no-newline 39, 39, 138, 1, 0, 1, p = 1.000; raw-input 40, 40, 140, 0, 0, 0, p = 1.000; no-label 40, 40, 136, 0, 0, 0, p = 1.000; double-rules 39, 39, 232, 1, 0, 1, p = 1.000. Clean repeats lesson 9's control: 40 of 40 replies identical. The two moved replies: the lamp message and the postage message, returns to delivery. PART 2, llama3.2:3b: clean 36, 36, 134, 0; trailing-space 37, 37, 135, 1, 1, 0, p = 1.000; no-newline 36, 36, 134, 4, 2, 2, p = 1.000; raw-input 39, 39, 136, 3, 3, 0, p = 0.250; no-label 35, 34, 132, 4, 1, 3, p = 0.625; double-rules 33, 33, 228, 6, 1, 4, p = 0.375. Clean repeats lesson 9's control: 40 of 40 identical. Then all 18 moved replies, with the message, the right category, the clean reply and the new one. Beneath: the lab's own report, you do not need to run it.

The lab is scripts/labs/prompting/templates.py. Its code mode is Part 1: it fills the template four ways for the 47 inputs and stores the counts in templates-code.json. It never calls a model and gives the same result every time. Its drift mode is Part 2: run once with MODEL=qwen2.5:3b and once with MODEL=llama3.2:3b, it imports lesson 1's rules, messages and scoring from parts.py, builds the six versions and stores every reply and its prompt token count. Both modes, the inputs and the six versions are written in the file's own description, which I wrote before the run; I edited the file afterwards to fix its report heading and a typo in the description ("Three ways" for the four fills).

Break a Template in Your Browser

This box has no model. The top half holds the lab's template and its four fill methods, exactly as in the lab, and runs them on one message that you can change. The bottom half holds the lab's real answers for all six versions on both models, one letter per message, with "?" for a reply that is not one of the four words alone, and prints each version's usable score with its fixed, broken and sign test against clean.

It prints what the lab found: for the policy question, format crashes, format2 changes the message and pulls in the internal note, and replace and the f-string are fine. Below that, qwen's scores of 40, 40, 39, 40, 40 and 39, and llama's 36, 37, 36, 39, 34 and 33, with no p value below 0.05. Try changing MESSAGE to your own text: an order number in curly brackets, {0}, a price with a dollar sign, or a line of JSON. Then change one "?" in llama's rows to the right letter and watch how little it takes to move a p value.

The Code, Part by Part

The template. TEMPLATE is a short sorting prompt with a JSON example in it and one gap, {message}. TWO_STEP has two gaps, the message and the policy note, and SECRET is the note. They are the lab's exact strings.

The four fills. FILL maps each method's name to a small function. format calls TEMPLATE.format. format2 fills the message and puts the text {policy} back into the policy gap, then calls format again on the result to fill the policy, as a second step elsewhere in a pipeline would. replace swaps the exact placeholder. fstring writes the whole prompt as an f-string; notice that its JSON example needs doubled brackets, {{ and }}, because inside an f-string a single bracket means a value.

The checks. For each tricky message, the loop tries each fill. An exception is printed as a crash with its error. Otherwise, if the customer's text is not in the result exactly as typed, it prints "CHANGED", and if the internal note appears more than once, it says so. These are the lab's crash, change and leak checks.

How to Build a Prompt in Code

A flowchart headed text from a customer, on its way into a template, titled escape, fill once, then test. Customer's text, then escape it in code (lesson 6), then fill every gap in one pass, then: tricky inputs pass, diff read? No: fix the code, and back to the fill. Yes: send it; log the exact text sent. Beneath: never call .format() on text that holds a customer's words.

Never run a fill over text that already contains an inserted value. That includes a template that already had the message filled in. Fill every gap in one pass: one format call with all values at once (brackets in the template doubled), one replace for a single gap, or one f-string. After that, the string is finished.

Escape before you insert. If your template puts the customer's text inside tags, escape &, < and > first, as lesson 6 did, so the customer cannot write your tags. If it uses another marker, escape or remove that marker instead.

Test the fill with tricky inputs. Keep a short list of messages like the lab's seven: curly brackets, numbered brackets, your own placeholder names, dollar signs, backslashes, an empty message, a very long one, and your own tag. Run every template through them in an automatic test, and check that nothing crashes and the text arrives unchanged.

Keep one copy, under version control. One template in one file, changed only through a reviewed commit.

Diff the exact text, then score it. Before shipping a template change, compare the old and new prompts for a real message character by character, and score both on your test cases, message by message, with the sign test.

A table headed what this lab measured, titled templates, measured. Format: crashed on 47 of 47. Format2: 2 crashed, 1 leaked. Replace, f-string: 0 crashes or changes; one gap, so no leak test. Small edits: usable qwen2.5 39 to 40, llama3.2 33 to 39; no p below 0.05. Beneath: PART 1 has no model; PART 2 is one run each at temperature 0.

When This Matters, and When It Does Not

It matters whenever a customer's text goes into a template. That is almost every real prompt. The fill is a few lines of code that people rarely check, and the lab found a crash that anyone typing curly brackets could trigger and a leak that a curious customer could trigger on purpose. These are ordinary bugs with ordinary fixes, and the fixes cost nothing.

It matters most in pipelines with several steps. The two-step bug only appears when text is filled, passed on, and filled again. If your prompt is built by several functions, or by a library that fills a template for you, find out whether any step fills after the customer's text is in. This lab did not test any template library; it tested plain Python.

Small characters moved the less accurate model more, here. In this lab, the five small changes moved 2 of qwen's 200 replies and 18 of llama's. The model that already made mistakes moved more. With two models, that is a pattern to watch for, not a rule.

It matters less for the exact choice between two harmless versions. In this lab, no small change was clearly better or worse than clean. Do not spend a day choosing between a trailing space and none. Spend it on the habits that catch the change you did not mean to make: one copy, a diff, and a test run.

Pasting the rules twice always costs tokens. Whatever it did to accuracy, double-rules added 94 tokens to every call. A mistake that makes every request bigger is worth catching, even when the score cannot see it.

What This Lab Can and Cannot Tell You

A two-column page headed read before you quote a number from this lesson, titled what was measured, and what was not. Measured: four fills, 47 inputs; five small edits, one run; two small models; one task: 40 messages. Not measured: template libraries; other models, more runs; attacks that name a category; chat templates, system messages.

Part 1 is exact: it is plain Python with no model, and it gives the same counts every time. But it tested four ways of filling one template in plain Python, on 47 inputs I chose. It did not test template libraries, which have their own rules for brackets and escaping, or other languages. The leak it found moved a note that was already in the prompt; it shows the mechanism, not harm to a real app.

Part 2 ran each version once, at temperature 0, on two small models and one task of 40 short messages. With 40 messages and a clean prompt that was already near the top, the sign test can only see large differences, so "cannot tell from luck" here means "too small for 40 messages to show", not "no difference". The five changes are five of the many small differences a template can pick up. The lab put the whole prompt in one user message; it did not test changes to a system message or to the chat template, the marker text the server adds around each message. What carries over is the method: fill every gap in one pass, test the fill with tricky text, keep one copy under version control, and compare the exact text and the scores before and after every change.

What to Do Next

A hand-drawn list headed for your own prompts, titled four things to do. 1, find: every place a prompt is filled more than once. 2, one pass: fill every gap at once; never fill text that already holds a value. 3, one file: keep each template in one file, under Git. 4, test, diff: run tricky inputs; print the exact text before and after a change. Beneath: then score it on the same cases, with the sign test.

Search your own code for every place a prompt is built. For each one, find out how the customer's text gets in, and whether anything is filled after it. If you see format called on a string that already contains a customer's words, change it so that the customer's text is inserted last, with replace or an f-string. Then write a small test with the lab's seven tricky messages and your own placeholder names, and run every template through it. Move each template into one file under version control, if it is not already. Finally, the next time someone changes a template, print the old and new text for one real message with repr, compare them, and score both versions on your test cases before the change goes live.

A closing card headed to keep, titled fill every gap in one pass. In large type: 3 of 47. Beneath: inputs where a customer's own braces broke the two-step fill. Then: test with the text a customer could type.

Knowledge Check

Knowledge Check

4 questions - Score 80% to pass

Q1

A template contains a JSON example with curly brackets. What happened when the lab filled it with str.format?

Q2

In the two-step template, a customer asked: What does {policy} say about refunds? What went wrong?

Q3

On llama3.2:3b, the raw-input version scored 39 against clean's 36: 3 fixed, 0 broken, p = 0.25. What should you do?

Q4

Which habit would have caught the no-newline change, where the prompt had exactly the same length as clean?

fstring

The inputs are lesson 1's 40 customer messages and 7 more that a real customer could type: an order number in curly brackets, the text {0} and {1}, two prices with dollar signs, a Windows folder path full of backslashes, the question "What does {policy} say about refunds?", an empty message, and "Hello." repeated 400 times. For each method the lab counts three things: a crash (Python raised an error, so no prompt was built), a change (the customer's text is not in the prompt exactly as typed), and a leak (the internal note appears more than the one time it should).

Part 2 uses the models. It sends lesson 1's prompt D and five versions of it that differ only in small characters to qwen2.5:3b and llama3.2:3b. That part starts on a later slide.

In this small lab the note was already in the prompt once, in its own "Policy:" line, so the model saw nothing it would not have seen anyway. What changed is who decides where internal text goes. A customer who guesses the name of a gap can move your text into their own words. If your app ever repeats the customer's question back, logs it for a human, or has a gap for something more private than this note, the same bug shows that text in places you did not plan.

A sketch of four stacked boxes joined by arrows, headed sketched: format2 on the policy message, titled filling twice lets the text in. Customer: What does, policy in curly brackets, say. Step 1: Message: What does, policy in curly brackets, say. Step 2 fills both policy gaps. The note now sits inside the message. Beneath: in the lab: 1 input of 47. replace() in two steps, checked after the run: 1 too.

This is the same mistake as the fake closing tag in lesson 6, seen from the other side. There, the customer typed my tag and ended my quote early. Here, the customer typed my placeholder and my code filled it. In both cases, text I did not write was treated as part of my structure.

format
replace
&
&amp;
<
&lt;
>
&gt;

The lab's other tricky inputs did not break any method. That does not mean they are fine to send. An empty message and 400 copies of "Hello." both went into the prompt exactly as typed, because that is what a fill should do. Whether the model should ever see them is a different check, before the template: reject empty text, and cut or refuse text longer than you planned for, as lesson 9's window arithmetic showed.

Five isometric columns headed qwen2.5:3b, median prompt tokens, drawn to scale, titled only one version cost many tokens. Clean: 138 tokens. Trailing-space: 139 tokens. Raw-input: 140 tokens. No-label: 136 tokens. Double-rules: 232 tokens, the tallest. Beneath: height is tokens; no-newline, 138, is not drawn. The rules pasted twice added 94 tokens to every call; the other versions moved the count by -2 to +2.

Every version except no-newline changed the number of tokens the model read, even the single space: 139 instead of 138 (the median, the middle value when all 40 counts are put in order). Pasting the rules twice added 94 tokens to every call on both models. That is a cost you pay on every request, and nothing in a normal code review would show it as a number.

This is a real run in VS Code's terminal.

A real screenshot of VS Code's terminal after running python templates_demo.py. Part 1, fill the template, no model. My order 12345 in curly brackets was charged twice: format CRASH KeyError category; format2 CRASH IndexError, replacement index 12345 out of range; replace ok; fstring ok. What does policy in curly brackets say about refunds?: format CRASH; format2 CHANGED, internal note now inside the message; replace ok; fstring ok. You took $49.99 instead of $39.99: format CRASH; format2, replace and fstring ok. Part 2, the exact text, then the model. The lamp message: clean vs no-newline, 3 changed lines, +0 characters; the last 24 characters end in new line Category: and in space Category:; clean got returns, no-newline got delivery. The postage message: clean vs double-rules, 9 changed lines, +467 characters; the same last 24 characters; clean got returns, double-rules got delivery.

Part 1 shows all three failures from the lab in a few lines: format crashes on every message, format2 crashes on the order number and pulls the internal note into the policy question, and replace and the f-string are fine. The dollar signs break nothing.

Part 2 picks the only two messages that moved on qwen in the lab, so it is chosen after the run to show a change, not to measure one. For the lamp message, clean and no-newline have exactly the same number of characters, and a line-by-line diff reports 3 changed lines; only the last characters, printed with repr, show the new line turned into a space. Clean got "returns" and no-newline got "delivery". For the postage message, pasting the rules twice added 467 characters and changed "returns" to "delivery", while the end of the prompt is identical. All four replies match the lab's stored replies for these messages and versions. I checked this by comparing them with the results file, and the output was the same on two separate runs. The script leaves out the lab's seed and 4,096-token window and uses Ollama's defaults; for these one-word replies it made no difference. On your computer, a different Ollama version or chip can change which of two nearly equal tokens is chosen, so a small difference is possible.

The report is a separate file, scripts/labs/prompting/templates_report.py. It reads the stored results and never calls a model. For Part 1 it also runs the four fills again in memory and checks that they give the stored counts, which they did, and prints which inputs crashed and what the leaked prompt said. For Part 2 it prints each version's scores, median tokens, how many replies differ from clean, the fixed and broken counts and the sign test, and every reply that moved.

I want you to know which parts came after I saw the data. The four fills, the 47 inputs, the six versions, the messages and the scoring were fixed before the run. The report was written afterwards, and so were its choices: the sign tests against clean, the list of moved replies, the check that clean repeats lesson 9's control, and the rerun of Part 1. The check of replace in two steps was added after the run and was not part of the design; the report says so on its own line. The lab's own report mode first had a small mistake in its heading, saying 6 tricky inputs where there are 7; I fixed the heading after the run, and the counts always used all 7. The demo script, with its choice of messages, was also written after the run.

Three brand cards headed the tools, with their logos, titled what the test ran on. Python: 47 inputs, four fills, no model. Ollama: qwen2.5:3b and llama3.2:3b, 480 calls. Git: not in the test: where a template should live.

The small changes. RULES is lesson 1's prompt D. VERSIONS builds clean, no-newline and double-rules exactly as the lab does. For each test message the script counts changed lines with difflib.ndiff, prints the difference in length, and prints the last 24 characters of both prompts with !r, which is repr: it shows a new line as \n and makes a space visible by its position.

The call. ask sends one prompt to qwen2.5:3b through Ollama's chat endpoint at temperature 0 with room for 30 tokens and returns the reply with spaces at either end removed.