Serving And Inference

Load Testing Honestly: Is Your Load Test Lying?

0 of 32 complete

0%

Contents

Back|Serving And InferenceLoad Testing Honestly: Is Your Load Test Lying?
1/32
83 min left
  1. Home
  2. AI Engineering: Data, RAG and Agents
  3. Serving and Inference Basics
  4. Load Testing Honestly: Is Your Load Test Lying?
Prerequisites
Cold Start: The First Request Waited for Python to Import Its LibrariesrequiredWorkers and Threads: Processes, Threads, or Both on One Core?requiredBatching Requests: Does It Help a Small CPU Model?requiredTail Latency: What Sits in the Slowest 1% of Requestsrequired
Related Topics
Model Signatures: The Right Numbers in the Wrong Shape, and What a Schema Check CatchesPackaging, Registry and VersioningHistogram MetricsObservability & MonitoringPercentilesObservability & Monitoring
1 of 32
Previous lessonCold Start: The First Request Waited for Python to Import Its Libraries

System Design

  • Foundation
  • Intermediate
  • Advanced
  • Capstone

AI Engineering

  • Foundation
  • Data, RAG and Agents
  • Evaluation, LLM Ops and Security

systemdesign.academy

  • Home
  • Glossary
  • Interview prep
  • Reviews
  • About
  • Privacy
  • Terms

The Inspector Who Stands in the Line

Let me start at a post office.

The manager wants to know how long customers wait at the one counter. She sends an inspector with a stopwatch. The inspector has a simple plan. He joins the line, waits, gets served, writes down his wait, and joins the line again. He does this all day.

A flat illustration of an engineer at a desk with two monitors full of charts: rings, bars and lines, a plant, a lamp and a coffee mug beside the keyboard. Below the picture: an engineer reads two screens of charts about a test; every chart is built from the requests the tester sent and timed; a request the tester never sent is in none of them.

At eleven o'clock the clerk goes to the back room for five minutes. The inspector is at the counter at that moment, so he waits five minutes and writes it down. That is one long wait in his notebook. But in those five minutes, thirty other people came in and stood in the line behind him. Each of them waited too, some for almost five minutes. None of them is in the notebook, because the inspector was busy waiting himself. At the end of the day his notebook says: almost every wait was short, and one was long.

A second inspector sits by the door instead. She writes down when each person walks in and when each person is served. Her notebook holds all thirty long waits.

A load tester for a web service can work like either inspector. In this lesson I test one small service both ways, and with a real tool, and I make the service stop on purpose, like the clerk. Then I check how far each report is from the truth.

Where This Lesson Starts

This is lesson 7 of the chapter on serving a model. It uses the same model and the same service as the lessons before it, so I do not explain them again. The model is the features chapter's gradient boosted tree model. It scores one customer of a real online shop, with a test AP of 0.5450. The service is the small FastAPI web service from lesson 2, latency anatomy. It runs with one worker process and one thread. Lesson 5, workers and threads found that to be the best setting on one core.

Every earlier lesson needed a load test. Lesson 4, batching requests, built the open-loop client I use again here, and lesson 3, tail latency, explained percentiles. Lesson 6, cold start, measured what the first request pays. This lesson turns round and tests the testers.

A percentile says what share of requests were faster than a given time. The median, also called the p50, is the time that half of the requests beat, and the p99 is the time that 99 of every 100 beat.

The machine has one core. My AWS account may run only one virtual CPU at a time, so I rent the same one-core box as before. There is no second box for the tester. The tester and the service share the core, and that is one of the things I measure.

The question: when a load tester gives you a p99, how close is it to what users would really see?

The Words You Need First

Please read this slide slowly if any word is new. Every slide after it uses these words.

A hand-drawn grid of twelve cards, each a word with its meaning: load test, tester, closed loop, open loop, paced loop, stall, scheduled time, coordinated omission, p99, correction, warm-up and latency bound. Below: request, latency, p50 and microsecond (us) mean what they meant in lessons 2 to 6; 1 s = 1,000 ms = 1,000,000 us.

A load test sends many requests to a service on purpose, to see how it behaves under traffic. The program that sends them and times each answer is the tester. A connection is one open line between the tester and the service, and in this lab each connection carries one request at a time.

A closed loop tester is the first inspector. Each connection sends its next request only after the last answer has come back.

An open loop tester is the second inspector. It decides in advance when each request should go out, as if users arrived on their own.

My open loop draws its arrival times as Poisson arrivals: random arrivals with no pattern, like people walking into a shop, at a steady average rate. Then it sends each one at that time, whether or not earlier answers are back. The moment a request was meant to go out is its scheduled time, and an open loop times each request from there. A paced loop is a closed loop that also aims at a rate. Each connection tries to send one request every so many milliseconds, but it still waits for its answer first.

A stall is a moment when the service answers nothing. Real services stall during a long garbage collection pause (the language tidying up memory nobody uses any more) or while they wait for a lock. Here I make the service stall on purpose, for 100 ms at a time.

Coordinated omission is the name Gil Tene gave to the inspector's mistake: the tester waits together with the service, so the requests it never sent are never timed. p99 is the time that 99 of every 100 requests beat. A correction tries to add back the requests a closed loop skipped. A warm-up is requests sent first and not counted. A is the slowest p99 you will accept.

The Headline: The Same Service, Very Different p99s

Here is the main result first. Each row is one kind of tester, and each dot is one of five rounds. All of them measured the same service, which stalled for 100 ms at the same random moments in every run of a round.

A ruler titled five rows, one service, five different answers, on a log scale (each mark is ten times the one before) from 1 ms to 1,000 ms with five rows, each with five dots, one per round. Closed, 1 connection: dots near 1 ms, labelled 0.97 ms. Paced, 8 connections: near 7.55 ms. Locust, 8 users: near 9.06 ms. Open, 546 a second: near 84.95 ms. Open, 1,074 a second: near 2,016 ms. Below: without any stall the one-connection closed loop also reported 0.97 ms; it timed each stall only once, as its slowest request, 101.5 to 101.6 ms in every round, far too rarely to move its p99; the open loop at 546 a second reported 84.95 ms with them and 5.29 ms without; at 1,074 a second, the rate the one-connection tester itself reached, arrivals on their own clock kept up in 1 of 5 rounds. Caption: by design a closed loop cannot send while it waits, 0 requests in 152 stalls; the open loop at 546 a second sent 2,057 during its 38.

The closed loop with one connection reported a p99 of 0.97 ms. With no stall at all, the same tester also reported 0.97 ms. Compared round by round, the two cannot be told apart. The stalls were there, 100 ms each, and this tester's p99 could not show them.

An open loop at 546 requests a second reported 84.95 ms. Without the stalls, at the same arrival times, it reported 5.29 ms. So the open loop's p99 showed the stalls very clearly.

Two paced testers aiming at 546 a second (they reached 514 and 525) reported 7.55 ms (mine) and 9.06 ms (locust). Both were far below the open loop's 84.95 ms, told apart in every round.

Why, in one count. By design, a closed loop cannot send while its requests wait. So the four closed loops sent 0 requests during the 152 stalls they met; that is a check, not a finding. The finding is what this does to the p99. Each stall reached the closed loop's record once (its slowest request, 101.5 ms, in every round), far too rarely to move its p99. The open loop at 546 a second sent 2,057 requests during its 38 stalls.

How the Lab Was Built

I wrote the lab's design into the docstring of scripts/labs/serving/load_testing.py on 2026-10-08, before any run on the box. It names every tester, every rate, the rule for "told apart from noise", what counts as the truth, and eighteen guesses.

A narrow title column reading load_testing.py, designed before the first run: twenty arms, 5 rounds; and six labelled zones. The service: lesson 2's handler, 1 worker, 1 thread; stalls of 100 ms at random times, on average every 4 s. Closed loops: 1, 4, 16 and 64 connections, each waiting for its answer. Paced loops: 8 connections aiming at 546 a second, my own client and locust. Open loops: 328, 546, 764, 983 and 1,074 a second, latency from the scheduled time. The order: 20 arms a round, 5 rounds, each round shuffled; 30 s each, two of 120 s. The rule: told apart only if all 5 paired differences agree in sign and the two sets do not overlap. Caption: 18 guesses written first; 13 came out right.

The stall. The tester sends a small "arm" request to the service at the start of each counted run. The service then draws stall times from a seed, a number that fixes a random draw so it can be repeated exactly. The stalls come on average every 4 seconds, never two within half a second. At each one, the service's event loop, the one thread that reads and answers every request in turn, runs time.sleep(0.1). For 100 ms nothing is read or answered. Sleep gives the core away, so the tester on the same core keeps its own clock during a stall.

One seed per round. Every run of a round with the stall on used the same seed, so every tester met the stalls at the same moments after its own start. And the open runs at one rate used the same arrival times with the stall on and off. So I could compare two testers, or stall on and off, on matched runs.

What counts as the truth. For a given rate, I take the open loop's p99 at that rate, in the same round, with the same stalls. That is what users arriving on their own clock at that rate saw.

The rule. A paired difference is the difference between two arms within the same round. A difference is "told apart from run-to-run noise" only if all five paired differences have the same sign, AND the two sets of five values do not overlap. When I say one tester reported more than another, I compared those two directly.

The Machine, and the Tester That Shares It

Like lessons 2 to 6, every timing behind a finding comes from a small machine I rented from Amazon Web Services. The one run on my laptop, the demo near the end, is labelled as such.

Four rows, each with a logo. EC2 c7g.medium, us-east-1: 1 vCPU (Neoverse-V1), 2 GiB, not burstable; the account allows 1 vCPU, so the tester shares the one core; on for 1.411 hours, $0.0512. Ubuntu 24.04, arm64: core at least 99.5% idle before every one of 110 runs; host steal 0.00%; timers stopped. Python 3.13.15, scikit-learn 1.9.1: numpy 2.5.3; locust 2.46.7 in its own venv; my testers in the service's venv. FastAPI 0.142.2, uvicorn 0.54.0: lesson 2's handler and model file, plus a stall the tester can switch on. Caption: every timing behind a finding comes from this box; the one laptop run is labelled as such.

The box is an EC2 c7g.medium. AWS's own API lists it with 1 virtual CPU, 1 core and 2,048 MiB of memory, and not burstable. I checked the account's limit with the AWS service-quotas tool on the day: 1 vCPU for standard machines. So the tester could not run on a second box.

Before every run, a fresh service was started. Then the core had to be at least 95% idle over two seconds before the tester began. In all 110 runs it was at least 99.5% idle. Steal, time the machine under my box gives to other customers' boxes, was 0.00%. No SSH session was open at the start or end of any run.

The system log had 64 lines during runs, in 11 of the 110 runs. Some came from a statistics job that cron starts on a schedule. Some came from the AWS management agent failing to reach AWS, because my box had no permission to talk to it. One came from systemd skipping a cache clean-up, and one from the kernel.

The Whole Lab in One Picture

Before the recordings, here is the whole lab as one numbered sequence. The numbers on the arrows match the numbers on the recordings that follow.

A sequence diagram with five lifelines: my laptop, AWS, the box, tester and service, and thirteen numbered arrows. 1, my laptop asks AWS whether any box is running. 2, key, group, SSH rule. 3, image, run-instances. 4, my laptop to the box: wait, then SSH. 5, Python, locust, files. 6, stop timers, look. 7, start, log out. 8, my laptop to AWS: read the console. 9, tester to service: arm the stall, requests. 10, service to tester: answers, stall times. 11, my laptop to the box: the demo. 12, the box to my laptop: raw files. 13, my laptop to AWS: terminate, delete. Below: arrows 9 and 10 are the timed work, 100 counted runs with 2,729,222 answers after 10 short calibration runs, with nobody connected; the tester and the service run on the box's one core; step 8 reads the serial console through AWS and never touches the box. Caption: steps 1 to 8 and 11 to 13 are the recordings that follow.

Steps 1 to 7 happen once: check, rent the box, set it up, and start the schedule. Step 8 is how I watched the schedule without touching the box. Arrows 9 and 10 are the timed work. Each counted run begins with the tester arming the stall. It ends with the service handing back the exact times of its stalls, read on the same clock as the tester's. Steps 11 to 13 come after the schedule: the student demo, the copy of the results to my laptop, and the deletion of everything.

Build the Lab Box Yourself

The rented box is part of the lab, so I show every step of it. You can rent the same kind of box and run the same schedule.

Which box you see. Every recording comes from the box that measured this lesson's numbers. No recording and no SSH happened during a timed run. Two steps were recorded twice. In step 2, the first recording lost its top lines. So, before renting anything, I deleted that key and group and made them again for the recording you see. In step 5, the same thing happened, and I ran the step again on the same box; it is safe to repeat.

What you need first. An AWS account, the AWS command line tool (aws) set up with your keys, and ssh, scp, rsync and curl. All the steps live in one file, scripts/labs/serving/load_files/load_aws_steps.sh. You run it from the scripts/labs/serving folder, one step at a time: bash load_files/load_aws_steps.sh check, then access, and so on. Every code block below is copied word for word from that file.

The blocks use a few names the file sets at its top. Run these lines first, from the scripts/labs/serving folder, if you paste the blocks by hand. All of them are copied from the file, except HERE=$PWD: the file works out that folder from its own location, which a pasted line cannot do.

export AWS_DEFAULT_REGION=us-east-1 AWS_PAGER=""
NAME=ai-research-course-load
HERE=$PWD
STATE=${LOAD_STATE:-$HOME/lab-data/serving/load}
KEY=$HOME/lab-data/serving/keys/$NAME.pem
K=${K:-5}; SEC=${SEC:-30}; LONG=${LONG:-120}
mkdir -p "$STATE" "$HOME/lab-data/serving/keys"
SSHO=(-o LogLevel=ERROR -o StrictHostKeyChecking=accept-new -o "UserKnownHostsFile=$STATE/known_hosts")
[ -f "$STATE/id" ] && ID=$(cat "$STATE/id")
[ -f "$STATE/ip" ] && IP=$(cat "$STATE/ip")
[ -f "$STATE/sg" ] && SG=$(cat "$STATE/sg")
B=/home/ubuntu/load

Steps 1 to 3: Check, Lock the Door, Rent the Box

Step 1: is any other course box running? My account allows one box of this kind at a time, and another lesson may be using it. So the first command counts the course's boxes that are not yet deleted. It must print 0.

A real terminal recording of step 1. The shell prints the command aws ec2 describe-instances with filters for the tag Project=ai-research-course and the states pending, running, stopping, stopped and shutting-down, asking for the number of instances; the answer is 0.

The recording shows the command and its answer, 0, so no other course box was alive. This is the command:

  aws ec2 describe-instances --filters "Name=tag:Project,Values=ai-research-course" \
    "Name=instance-state-name,Values=pending,running,stopping,stopped,shutting-down" \
    --query "length(Reservations[].Instances[])"

Step 2: a key and a locked door. A key pair is how you prove to the box that you are allowed in. AWS keeps one half, and you keep the other half in a file only you can read. A security group is a door rule for the box. Mine opens only port 22, the SSH port, and only to my own address. Both get tags, labels that say which project and lesson they belong to.

A real terminal recording of step 2, with the address, network and group ids replaced by placeholders. create-key-pair for ai-research-course-load with the tags Project and Lesson writes the key to the key file without printing it; chmod 600 on the key file; curl reads my address; describe-vpcs finds the default network; create-security-group makes the group ai-research-course-load with its tags; authorize-security-group-ingress prints a small table: from my address /32, port 22. Last line: the key file, readable only by its owner, 388 bytes.

The key material went straight into the file and was never printed, and the door rule names only my own address. These are the commands:

  aws ec2 create-key-pair --key-name ai-research-course-load --key-type ed25519 \
    --tag-specifications "ResourceType=key-pair,Tags=[{Key=Project,Value=ai-research-course},{Key=Lesson,Value=load}]" \
    --query KeyMaterial --output text > ~/lab-data/serving/keys/ai-research-course-load.pem
  chmod 600 ~/lab-data/serving/keys/ai-research-course-load.pem
  MYIP=$(curl -sf https://checkip.amazonaws.com)
  VPC=$(aws ec2 describe-vpcs --filters Name=isDefault,Values=true --query "Vpcs[0].VpcId" --output text)
  SG=$(aws ec2 create-security-group --group-name ai-research-course-load --vpc-id "$VPC" \
    --description "lesson 7 load testing, ssh from one address" \
    --tag-specifications "ResourceType=security-group,Tags=[{Key=Project,Value=ai-research-course},{Key=Lesson,Value=load}]" \
    --query GroupId --output text)
  aws ec2 authorize-security-group-ingress --group-id "$SG" --protocol tcp --port 22 --cidr "$MYIP/32" \
    --query "SecurityGroupRules[].{port:FromPort,from:CidrIpv4}" --output table

Steps 4 to 7: Wait, Set Up, Look, Start

Step 4: wait until it runs. The AWS tool waits until the box is running. Then the script reads the box's public address into IP, and tries SSH every five seconds until the box answers.

A real terminal recording of step 4, with the instance id and address replaced by placeholders. aws ec2 wait instance-running; describe-instances reads the public address into IP; describe-instances prints running, c7g.medium, us-east-1a. Then ssh with ConnectTimeout=5 runs true and fails once with a sleep 5, and the next ssh answers: ssh works, aarch64, Ubuntu 24.04.5 LTS.

The first SSH try failed because the box was still starting, and the second answered. These are the main commands:

  aws ec2 wait instance-running --instance-ids "$ID"
  IP=$(aws ec2 describe-instances --instance-ids "$ID" \
    --query "Reservations[0].Instances[0].PublicIpAddress" --output text)
  until ssh -i ~/lab-data/serving/keys/ai-research-course-load.pem "${SSHO[@]}" -o ConnectTimeout=5 \
    ubuntu@"$IP" true 2>/dev/null; do sleep 5; done
  echo "$IP" > "$STATE/ip"

Step 5: Python, locust and the lab files. The box gets the same Python, 3.13.15, and the same pinned library versions as lessons 2 to 6, through uv, a fast installer for Python. locust gets a second, separate environment, so it cannot change the service's libraries. Then scp, a copy command that works through SSH, sends the files over. The last command builds the SQLite feature table, computes the reference score of every request row, and prints the files' sha256 fingerprints.

A real terminal recording of step 5, with the address replaced by a placeholder and my home folder shortened to a tilde. Over ssh: make the folders; install uv, Python 3.13.15 and a venv; install the pinned packages; make venv-locust and install locust 2.46.7 with numpy, pandas and pyarrow, which prints locust 2.46.7. Then scp copies lesson 2's model, data and service files, this lesson's load_files, the features chapter's files and data, and the demo. Last: store: 26851 rows; raw/ref.npy: 26851 single-row scores, equal to lesson 2's expected column 26818 of 26851; the sha256 of model.pkl starts bec4d11e, store.parquet 62aadc30, requests.parquet 407d62e6.

Steps 8 and 11 to 13: Watch, Demo, Bring It Home, Delete

Step 8: watch without touching. While it runs, the schedule writes one progress line to the box's serial console after every run. The serial console is a log that AWS keeps for the box, and you can read it through AWS without logging in.

A real terminal recording of step 8, with the instance id replaced by a placeholder. aws ec2 get-console-output with --latest, piped through grep LOAD-PROGRESS and tail -12, prints progress lines with times: o3-off-r2, o5-on-r2, o9-off-r2, o5-off-r2, o7-off-r2, c64-on-r2, o3-on-r2, p5-on-r2 and c4-on-r2 done between 04:05:53 and 04:10:31, round 2 finished (20 runs), then p5-off-r3 and o3-off-r3 done.

These lines came from AWS's copy of the console while the runs went on, so the box was never touched. This is the command:

  aws ec2 get-console-output --instance-id "$ID" --latest --output text | grep LOAD-PROGRESS | tail -12

Step 11: the student demo, on the same box. After the schedule, I ran the demo you will meet later, so you can see what it prints on one core.

A real terminal recording of step 11, with the address replaced by a placeholder: python load_demo.py on the box, its output also written to load-demo-run.txt. It prints the machine, 1 cores, load 0.37, and test AP 0.5450; the service stalls 100 ms every 2 s; A sent 903 requests a second and B's random arrivals came at 451 a second. A closed loop: 9,029 sent, p50 1.05 ms, p99 1.14 ms, max 101.7 ms, mean 1.11 ms, 0.04% over 50 ms. B open loop: 4,439 sent, p50 1.64, p99 91.71, max 104.2, mean 6.57, 4.78% over 50 ms. A after the HdrHistogram correction: 384 values added, p99 77.59 ms.

This run happened after the last timed run had finished, so it could not disturb any timing. Its first line says "1 cores", which is how my demo prints the count; it is the box's one core. This is the command:

  ssh -i ~/lab-data/serving/keys/ai-research-course-load.pem "${SSHO[@]}" ubuntu@"$IP" \
    'cd load/serving/examples && ../../venv/bin/python load_demo.py --save ~/load/raw/load-demo.json \
     | tee ~/load/raw/load-demo-run.txt'

Step 12: bring the results home first. Lesson 3 lost a whole box of results because it was deleted before anyone copied them. So the moment the runs finish, the raw files come back, before anything else.

A real terminal recording of step 12, with the address replaced by a placeholder. Over ssh, the box records its info after the schedule and prints the size of its raw folder, 14M; then rsync copies the box's load/raw folder to my laptop; ls counts 260 files; du prints 13M for the copy.

The Lab's Report, Running

This is a real recording of the report script, load_report.py. It ran on my laptop, but it times nothing. It reads the raw files the box measured and does all the arithmetic again with its own code.

A terminal recording of load_report.py in nine numbered sections. 1: the box, c7g.medium, 1 vCPU, the tester shares the one core; 1.411 h at $0.0363 an hour = $0.0512 plus disk $0.0025, everything deleted; 110 runs, core at least 99.5% idle before each, 0 ssh sessions, 141 stored texts scanned, 0 ids, addresses or keys. 2: S = 1092.1, R1 = 1074.4 = 0.984 x S, open rates 328, 546, 764, 983 and 1074. 3: 100 counted runs, 2,729,222 answers, 0 scores not equal to the bit, 523 stalls of 100.0 to 100.1 ms, and a table of p50, p99, max and mean for nine arms. 4: seven p99 comparisons, six told apart and c1-on against c1-off cannot be told apart. 5: the correction for p5-on, l5-on and c1-on; locust's own 99% 9 ms, from its raw times 9.04 ms. 6: the other lies, with a line added after review: o7-off p99 within 10 ms in 4 of 5 rounds, locust counted 28.49 to 29.61 s. 7: 13 guesses right, 5 wrong. 8: the demo and the fact-check. 9: the playground. Last line: all 974 checks agree with the stored lab.

The report does not import the lab. It reads the small copy of the raw files kept in the repo. It computes its own percentiles, its own correction, its own paired differences and its own verdicts. Then it compares every one with the lab's results file. If any number disagreed, it would stop with an error. I checked that too. I changed one p99 in a copy of the results file by 1 microsecond, and the report printed "1 MISMATCHES" and stopped with an error.

Section 1 also scans every stored file, including the text inside the compressed .npz files and the terminal recordings. It looks for my account number, IP addresses, host names, AWS ids, key fingerprints and my home folder. The first scan found 3 hits. All three were the version string of a maths library, "0.3.34.106.0": five numbers in a row, like an address with one extra part. I made the scan match exactly four numbers, and it then found 0.

During a Stall, the Closed Loop Stops Sending

Why did the closed loop's p99 not show the stalls? The tester wrote down when it sent every request, and the service wrote down when every stall began. So I could line up every stall and count the requests sent around it.

A line chart of requests sent per 10 ms against milliseconds from the start of a stall, from -150 to 350, with dashed lines at 0, labelled stall, and at 100, labelled ends. The closed loop with 1 connection runs near 10.8 per 10 ms, drops to 0 at the stall and stays at 0 until 100 ms, then jumps back. The paced loop with 8 connections runs between about 3 and 7, sends a little just after the stall starts, then 0 until 100 ms, then rises and swings. The open loop at 546 a second stays near 5 to 6 the whole time, through the stall. Below: the closed loop sent about 10.8 requests per 10 ms before a stall and none inside it, 0 in 38 stalls, the same for 4, 16 and 64 connections; the paced loop sent at most 8 into one stall; the open loop's arrivals kept coming, 2,057 in 38 stalls. Caption: a request that is never sent is never slow.

The closed loop's line falls to zero for exactly the length of the stall. Its one connection had a request in flight when the stall began, and it could not send another until the answer came. This is not a finding: a closed loop works this way by design. It is a check that the stall happened and the tester behaved as designed.

The paced loop sent a few requests just after the stall began, then stopped. Its 8 connections were each waiting for their next time slot. The ones whose slot came during the stall sent their request, and then waited like the closed loop. It sent at most 8 requests into any one stall.

The open loop's line did not move. Its arrivals came on their own clock, about 5 to 6 per 10 ms, and it sent them all, 2,057 requests into 38 stalls. Each was timed from its scheduled moment, so each one's wait behind the stall was counted. That is the whole difference between the testers, in one picture.

Who Waits Behind the Stall

Here is the same idea as a drawing, for one stall.

A hand-drawn timeline. At the top, a long box labelled the service stops for 100 ms. Below it, a row labelled closed, 1 conn. with one small box and the words 1 request waits; no new ones. Below that, a row labelled open loop with nine small boxes side by side and the words arrivals keep coming and queue. An arrow along the bottom is labelled time. Below: at 546 a second the open loop scheduled 54.1 requests inside each 100 ms stall on average; it sent them all and timed each from its arrival, so their waits run from almost 100 ms down to nothing; the closed loop had one request in flight, timed that one, and sent nothing else until the stall ended; in the open runs at 546 a second, 2.7% of requests overlapped a stall; in the one-connection closed loop, 0.025%. Caption: this is coordinated omission, the tester waits together with the service.

Now count. A p99 is the time that 99 of every 100 requests beat, so it looks at the slowest 1%. In the open loop at 546 a second, 2.7% of all requests overlapped a stall. That is more than 1%, so the p99 landed inside the stalled requests: 84.95 ms.

In the closed loop with one connection, only 0.025% of requests overlapped a stall: one per stall, out of about a thousand answers a second. That is far less than 1%, so the slowest 1% was made only of ordinary requests. The p99 was 0.97 ms, the same as with no stall at all.

Gil Tene named the problem in a post to the mechanical-sympathy mailing list on 4 August 2013. He described it as a "common measurement technique problem" and wrote that it "can often render percentile data useless". My lab shows that sentence in numbers: the closed loop's p99 was a correct percentile of the requests it sent, and those requests were the wrong sample.

What Each Tester Reported

Here is every tester's p99 with the stall on, every round as a dot.

A dot chart on a log scale (each mark is ten times the one before) from 1 to 1,000 ms with nine columns: C1, C4, C16, C64, P8, L8, O328, O546 and O764, five dots each. C1 near 1 ms, C4 near 4.6, C16 near 15.6, C64 near 160, P8 near 7.5, L8 near 9, the three open loops near 79, 85 and 98. Below: C means closed loop with that many connections, P8 my paced client, L8 locust, O the open loop at that many a second; medians C1 0.97, C4 4.62, C16 15.62, C64 159.96, P8 7.55, L8 9.06 ms, the open loops 78.67, 84.95 and 97.55 ms; P8 and L8 against O546, both lower, told apart; P8 and L8 aimed at 546 a second and reached 514 and 525. Caption: the open loops' p99 showed the stall at every rate; of the closed loops' p99s, only 64 connections' did, on top of its own queue.

The open loops' p99 showed the stall at every rate. At 328, 546 and 764 a second their p99 was 78.67, 84.95 and 97.55 ms. Without stalls, the same arrivals gave 3.39, 5.29 and 9.77 ms. At every rate, stall on against stall off was told apart.

The paced testers' p99 missed it at about the same rate. They aimed at 546 a second and reached 514 and 525. My paced client reported 7.55 ms and locust 9.06 ms, against the open loop's 84.95 ms at 546 a second. Both were told apart from the open loop, round by round. They were 0.086 and 0.107 of it, as medians of the five rounds.

The closed loops' p99 grew with their connections, and that is the next slide. Only the one with 64 connections reached the stall's size.

One thing about the paced client surprised me. Even with no stall, it reported a p99 of 7.21 ms against the open loop's 5.29 ms at the same rate, told apart. And it sent 524.6 requests a second, not the 546 it aimed at. One possible reason is that it waits with a timer that rounds up to whole milliseconds, so its sends drift late and bunch up. I did not test that.

More Connections: The Closed Loop's Own Queue

What if the closed loop opens more connections? Then more requests are in flight when a stall begins, and more of them wait through it.

Four hand-drawn bars for closed loops with 1, 4, 16 and 64 connections, stall on, each split into a lower part for p50 and the whole bar for p99, median of 5 rounds. The 1 and 4 connection bars are very short, 0.97 and 4.62 ms; 16 connections 15.62 ms; 64 connections 159.96 ms, about three times the height of its lower part. Below: p50 0.93, 3.67, 14.57 and 58.96 ms; with N connections each request waits behind about N - 1 others, so the p50 grew with N and the answers a second stayed near S, 1,048, 1,062, 1,070 and 1,057; only with 64 did more than 1% of requests overlap a stall, 1.6%, so only its p99 showed one, 159.96 ms, about its own p50 of 58.96 ms plus one 100 ms stall. Caption: a closed loop's latency is partly made by the tester itself.

With 64 connections the p99 was 159.96 ms. That is about 59 ms of waiting behind the other 63 requests, plus one 100 ms stall. So here the closed loop did see a stall, but mixed with a queue that it made itself. Its p50, 58.96 ms, is all its own queue. Users arriving on their own at that rate would not wait behind 63 others, unless they really came at the same moment.

The answers a second hardly moved: 1,048, 1,062, 1,070 and 1,057 for 1, 4, 16 and 64 connections. One connection already kept the one core almost fully busy, because the tester and the service take turns on it.

So a closed loop gives you two dials, and neither is "what users see". Few connections hide the stall. Many connections add a queue of their own. A closed loop tells you most about the most a service can answer, which is how lessons 4 and 5 used it. It tells you least about the waits of users.

The HdrHistogram Correction

There is a known repair for a closed or paced tester. HdrHistogram is a library for recording latencies, written by Gil Tene. Its method recordValueWithExpectedInterval takes each recorded v and an expected interval I, the time between two requests when nothing goes wrong. If v is larger than I, it also records v - I, v - 2I, and so on, while they are at least I. Its documentation says it will "auto-generate an additional series of decreasingly-smaller" values, standing in for the requests the tester should have sent while it waited.

Three panels, each with a large number. Paced, 8: 69.79 ms; recorded 7.55 ms; true 84.95 ms; corrected over true 0.76 to 0.83. Locust, 8: 71.10 ms; recorded 9.06 ms; true 84.95 ms; corrected over true 0.72 to 0.86. Closed, 1: 64.34 ms; recorded 0.97 ms; the open loop at 1,074 a second kept up in 1 of 5 rounds. Below: for each recorded latency v the correction also counts v - I, v - 2I and so on while they are at least I; I was the pacing interval, 14.65 ms, for the paced testers and the run's own median for the closed loop; it stayed below the open loop's p99; one possible reason is that it assumes each skipped request would have been answered at once, while the open loop also saw the queue that builds behind a stall. Caption: a correction is an estimate; an open loop is a measurement.

I applied it to the stored latencies, with I chosen before the runs. For the paced testers I was the pacing interval: 8 connections at 546 a second is 14.65 ms. For the closed loop it was the run's own median latency, which was about 0.3% below the median gap between its sends.

For the paced testers it got most of the way. My client went from 7.55 ms to 69.79 ms, and locust from 9.06 ms to 71.10 ms. The open loop at the same rate said 84.95 ms. Corrected over true was 0.76 to 0.83 for mine and 0.72 to 0.86 for locust, across the five rounds, and still below the truth in every round. Even with no stall, my paced client's p99 was 1.36 times the open loop's, so part of its 7.55 ms is its own pacing.

For the one-connection closed loop, the correction went from 0.97 ms to 64.34 ms. But the truth for its rate is hard to state. It sent 1,048 requests a second with the stalls. Arrivals on their own clock at 1,074 a second were more than the service could keep up with. The open loop at that rate kept up in only 1 of 5 rounds. Its p99, 2,016 ms, only says that the queue never emptied. So the corrected 64.34 ms is far below that truth.

I guessed the correction would land within 0.8 to 1.2 of the truth for the paced client. It landed at 0.76 to 0.83, so that guess was wrong, but not by much.

A Real Tool: locust

I wanted one real, widely used tool on the same service. locust is a load-testing tool written in Python and installed with pip.

Its users run tasks, and a task waits for its answer before the user goes on. In its source at version 2.46.7, a user runs a task and then calls self.wait() before the next one. Its constant_pacing wait aims at a fixed time between task starts. Its own docstring says: "If a task execution exceeds the specified wait_time, the wait will be 0 before starting" the next one. Its FastHttpUser times a request from start_perf_counter = time.perf_counter(), just before the request goes out. So locust is a paced closed loop, the same model as my paced client.

A table under the title a real tool, measured the same way, locust 2.46.7, 8 users, constant_pacing, aiming at 546 a second, stall on, medians of 5 rounds. Rows: what locust printed as its 99%, 9, 9, 9, 8, 9 ms, one per round; the same from locust's own raw times, 9.04 ms; timed by me around each call, 9.06 ms, my paced client 7.55 ms, told apart; what users arriving on their own saw, 84.95 ms (open loop, same rate, same stall seed); locust's CPU per request 369 us, my paced client 81 us. Below: a locust user waits for each answer before its next task, so locust sent at most 8 requests into any stall, then waited with the service; its own table rounds to whole milliseconds and agreed with its raw times; its p99 was higher than my paced client's, one possible reason being that it used 4.50 times the CPU per request on the shared core, which I did not test; locust's round 1 counted 28.49 s and met 5 stalls, the others about 29.6 s. Caption: the tool reported faithfully what it measured; it just did not measure the stall.

locust's own table said 9 ms in four rounds and 8 ms in one. Its raw response times gave 9.04 ms, and my own clock around each call gave 9.06 ms. So locust reported faithfully what it measured. It simply measured the same sample my paced client did, and that sample left out the stall.

One caution about locust's numbers. Its round 1 counted only 28.49 s, because its 30 s includes its own start. That round met 5 stalls where the open loop's round 1 met 6; the other rounds counted about 29.6 s. I divided its answers by a fixed 30 s, so its 525 a second reads a little low: by its own counted span it was 531 to 537. Round 1 also gave its lowest corrected ratio, 0.72.

Other Lie 1: The Run Was Too Short

Coordinated omission is the big lie here, but not the only one. The next four slides each take one more, measured on the same data.

A dot chart of the p99 of each 10 s window of five 120 s runs of the open loop at 546 a second, stall on, against seconds into the run, with a dashed line labelled whole 120 s, median 85.3 ms. Large dots at the first window: four between about 79 and 93 ms and one near 5 ms. Smaller dots for the other windows spread from about 5 to 97 ms, most between 60 and 97 ms. Below: a 10 s test is the first window; its p99 ranged from 5.4 to 93.1 ms across the 5 rounds, 17.1 times apart, and in round 1 no stall fell in the first 10 s; the whole 120 s gave 71.6 to 88.6 ms, 1.24 times apart; without stalls the first 10 s was within 0.98 to 1.06 of the 120 s p99. Caption: a rare event needs a run long enough to contain it many times.

Two of the arms ran for 120 seconds instead of 30, with the stall on and off. I cut each run into twelve windows of 10 seconds. The first window is exactly what a 10-second load test would have reported.

With stalls, a 10-second test could say anything. One of five 10-second tests missed every stall and reported 5.4 ms; the other four reported 79.5 to 93.1 ms, so the five were 17.1 times apart. In that round no stall happened in the first 10 seconds, so a 10-second test would have said the service never stalls. Over the whole 120 seconds, the p99 ranged only from 71.6 to 88.6 ms, 1.24 times apart.

Without stalls, 10 seconds was enough. The first window was within 0.98 to 1.06 of the 120-second p99 in every round. So run length matters most when the slow event is rare. Here a stall came on average every 4 seconds. A real pause that comes once a minute would need a run of many minutes to show up even a few times.

Other Lie 2: The Average

Many dashboards still show the average , called the mean: the total of all latencies divided by how many there were.

Four isometric towers for the open loop at 546 a second, stall on, median of 5 rounds, heights in milliseconds: p50 1.34 ms, a flat tile; mean 4.47 ms, a low tile; p99 84.95 ms, tall; max 104.95 ms, tallest. Below: the mean was 4.47 ms, above the p50 because the stalled requests pull it up, but 19.0 times below the p99; a report with only the mean would say about 4 ms for a service where 1 request in 100 waited 85 ms or more; without stalls the same arm's mean was 1.65 ms. Caption: report percentiles; here the mean sat far from both the typical and the slow request.

With the stall on and the open loop at 546 a second, the mean was 4.47 ms. The p50 was 1.34 ms, the p99 84.95 ms and the slowest request 104.95 ms.

The mean matched almost no request. Most requests took about a millisecond. The ones caught by a stall took tens of milliseconds. The mean, 4.47 ms, sits between the two groups, where few requests actually were. A report with only the mean would say "about 4 ms" for a service where 1 request in 100 waited 85 ms or more.

The mean did move with the stalls, from 1.65 ms without them to 4.47 ms with them, so it is not useless. But it hides the shape. Report the p50, the p99 and the max, and look at all three.

Other Lie 3: Throughput With No Latency Bound

The most common number in a load test report is : how many requests a second the service handled. A closed loop with many connections gives a big one.

A two-column ledger titled a throughput number needs a latency bound, stall off. What the closed loop said: closed loop, 16 connections, 1,098 answers a second; its p50 was 14.55 ms and its p99 15.56 ms, made by its own queue of 16; so 1,098 a second is 2.01 times the highest rate I tested that met the bound; the true limit lies between 546 and 764. What the bound allowed: open loop, p99 within 10 ms in all 5 rounds at 328 and 546 a second; at 764 a second its p99 was within 10 ms in 4 of 5 rounds, one round reached 10.97 ms; at 983, 45.31 ms; say the rate together with the p99 it held, and how the arrivals were sent.

The closed loop with 16 connections, with no stall, reported 1,098 answers a second. That is a true number: the service did answer that many. But every one of those answers waited behind 15 others, so the p50 was 14.55 ms.

With a bound, the answer changes. Before the runs I chose a bound: p99 at most 10 ms. Then I asked the open loop.

At 328 and 546 arrivals a second, the p99 stayed within 10 ms in all five rounds. At 764 a second its p99 was within 10 ms in 4 of 5 rounds; one round reached 10.97 ms. At 983 a second it was 45.31 ms. So the honest statement is "546 a second at a p99 of 5.29 ms", not "1,098 a second". The closed loop's number was 2.01 times the highest rate I tested that met the bound. I tested only 328, 546, 764 and 983 a second, so the true limit lies between 546 and 764 a second.

A throughput number without a latency bound answers the question "how busy can I make it?". Users ask a different question: "how long will I wait?"

Other Lie 4: The Tester Shares the Core

On this box there was no second machine for the tester, so every tester ran on the same one core as the service. Every microsecond the tester used was one the service did not get.

Five hand-drawn bars of the tester process's CPU time as a share of the run, median of 5 rounds, stall on: closed 1 0.08, paced 8 0.04, locust 8 0.20, the tallest, open 546 0.05, open 1,074 0.09. Below: per request, closed 1 72 us, paced 8 81 us, locust 8 369 us, open 546 98 us, open 1,074 92 us; the account allows one vCPU, so there was no second box for the tester; every microsecond the tester used was one the service did not get, and the open loop's own sender lag, sent minus scheduled, had a p99 of 0.61 ms without stalls. Caption: put the tester on its own machine when you can, and say so when you cannot.

My testers used 0.04 to 0.09 of the core. locust used 0.20, with 369 microseconds of CPU per request against 81 for my paced client, 4.50 times as much, told apart.

The open loop has its own small lie too. It has to send each request at its scheduled moment, but on a shared core it is sometimes a little late. I measured the gap between scheduled and actually sent: with no stall at 546 a second its p99 was 0.61 ms. That lateness is counted in the , because the clock starts at the schedule, which is honest. But it is the tester's time, not the service's.

I cannot show here what a separate tester machine would change, because I could not rent one. What I can say is the size: in these runs the tester took up to a fifth of the core.

One Lie That Turned Out Small: The Warm-Up

The last lie on my list was counting the warm-up. A fresh service is slow on its first requests, so a test that counts them reports a slower service than the one users meet later. I ran the open loop at 546 a second with no stall, counted from the very first request a fresh service ever answered. Then I compared it with the normal runs, which had 400 warm-up requests first.

Two panels. Warm-up not counted: 5.29 ms, p99; max 9.9 ms; mean 1.65 ms. First requests counted: 5.37 ms, p99; max 10.2 ms; mean 1.66 ms. Below: p50, p99, max and mean could not be told apart, round by round; the max was higher without the warm-up in 3 of 5 rounds; lesson 6 found the first request of a fresh service slower than the second; among about 16,160 requests that is one value, and it cannot move a p99. Caption: a lie that turned out small here is still worth checking on your service.

Here it made no measurable difference. The p99 was 5.37 ms against 5.29 ms, the max 10.2 against 9.9 ms, and none of p50, p99, max or mean could be told apart. The max was higher without the warm-up in only 3 of 5 rounds. I had guessed 4 of 5, so that guess was wrong.

The reason is size. A 30-second run at 546 a second holds about 16,000 requests. Even if the first few are slow, they are a handful of values among sixteen thousand. On a service with a slow first minute, such as a large model loading on demand, or with short tests, the warm-up could matter a lot. Check it on your own service rather than trust my small number.

My Eighteen Guesses Before the Run, Checked

I wrote eighteen guesses into the lab before it ran. Each is quoted here word for word from the docstring. Thirteen were right and five were wrong.

  1. "G1. S between 1,000 and 1,200 answers a second (lesson 5: 1,077.1)." Right: 1,092.1.

  2. "G2. R1 between 0.70 and 0.95 x S." Wrong: 0.984. One connection kept the one core almost as busy as sixteen.

  3. "G3. every armed run records all its planned stalls, each between 100 and 110 ms long." Right, 100.0 to 100.1 ms, but only after I fixed my check (see the limits slide).

  4. "G4. o5-on p99 between 60 and 100 ms in every round (reasoned: 100 - 0.01 x (1 - 0.5) x 4,000 = 80 ms)." Right: 75.8 to 88.7 ms.

  5. "G5. c1-on p99 below 3 ms in every round, below oc-on's p99, told apart; oc-on's at least 20 x c1-on's." Right, but only because the open loop at that rate was overloaded, so this says nothing about the stall.

  6. "G6. c1-on p99 at most 1.25 x c1-off p99 in every round (one connection does not see the stall in its p99)." Right: at most 1.006 times.

  7. "G7. p5-on p99 below 10 ms in every round; o5-on's at least 5 x p5-on's, told apart." Right.

  8. "G8. l5-on p99 (locust's raw response times) below 10 ms in every round; locust's own "99%" within 1 ms of it." Right.

  9. "G9. corrected p5-on p99 between 0.8 and 1.2 x o5-on p99 in every round." Wrong: 0.76 to 0.83.

Try It Yourself

The full lab needs a rented box and about an hour and a half. The demo, load_demo.py, runs on your own computer. It starts a small stalling service and measures it with a closed loop and an open loop.

A narrow title column reading load_demo.py, designed before it ran: two testers on your own machine; and five labelled zones. Train: the features chapter's model; it must score test AP 0.5450. A stalling service: a small FastAPI service in a second process; 100 ms stall every 2 s. A, closed loop: one connection for 10 s; latency from the send. B, open loop: random arrivals at half of A's rate for 10 s; latency from the scheduled time. Print: p50, p99, max, mean, share over 50 ms, and A after the correction. Caption: on the box (1 core), A's p99 1.14 ms, B's 91.71 ms.

I wrote the demo's design into its docstring after the lab's design and before the demo first ran.

A real screenshot of VS Code with load_demo.py open at the top of the file, showing its docstring: what it needs, how to run it, and the design written before it first ran.

Before you run this lab. The demo uses lesson 2's Python environment, with the same library versions as the lab's box. If you made it for lessons 2 to 6, use it again. If not, make it once, inside the scripts/labs/serving/examples folder:

python3 -m venv venv-sv
source venv-sv/bin/activate        # on Windows: venv-sv\Scripts\activate
pip install -r lat_requirements.txt

I used Python 3.13. Then run python ../../features/fetch_data.py once, which downloads the shop data. Now run python load_demo.py. It needs no GPU and no cloud account. It starts its own small service on your machine only (127.0.0.1) and stops it at the end. If a package is missing, it prints one line saying why and stops.

The first line it prints says the times come from your machine, with its number of cores and its load at that moment. Read the two testers against each other, on your machine. Does the closed loop's p99 stay near its p50? Does the open loop's p99 land near the stall's 100 ms? How close does the correction get?

Simulate a Stalling Server

This box runs in your browser and needs nothing but Python. It is a simulation: a pretend server that answers one request at a time, in the time the lab's box took per request. It stalls for STALL_MS at random moments. A pretend closed loop and a pretend open loop measure it with a pretend clock. It has no model and no network, so it shows the mechanism, not the box's numbers. The box's real p99s are printed under it.

Press Run. Then try CONNS = 64, and watch the closed loop's p99 grow from its own queue. Try RATE = 1000, close to the most this pretend server can answer, and watch the open loop's p99 grow. Try STALL_MS = 0, and the two testers almost agree.

The report script writes this box from the lab's results and runs it. It checks that the simulated open loop sees the stall and the simulated closed loop does not.

The Lab's Code, Piece by Piece

The lab has one file that runs on my laptop, a few small files that run on the box, and the step script you saw above.

load_testing.py, on my laptop, holds the design in its docstring, written before any run. Its commands are check (lesson 2's files match by sha256), status (reads the box's serial console), collect, cost, factcheck and restat. collect calls load_stats.py, which copies the box's raw files into results/load-raw/, passes every log line through load_redact.py, and does all the arithmetic.

load_files/load_service.py is the service. It imports lesson 2's service file unchanged, so the model, the SQLite table and the handler are lesson 2's. It adds the stall. /arm draws the stall times from a seed and schedules each on the event loop. /stalls returns when each one really started and ended.

load_files/load_client.py holds my three testers. The open loop is lesson 4's client, copied with credit. It draws arrival times first, has a sender thread and a receiver thread and 512 kept-alive connections, and times from the scheduled arrival. The closed loop and the paced loop are new. All three check every score against the box's own single-row score.

How to Set Up a Load Test You Can Trust

Here is the order I would follow, using only what this lab measured.

A flowchart titled five checks before you trust a number. Does the tester send on its own clock? No leads to use an open loop, or correct and say so. Yes, and that box, lead to latency from the scheduled time? Then long enough to see the rare event many times? Then p50, p99 and max, never only the mean. Then a rate together with the p99 it held. Below: here, at 546 a second, my paced client reported 7.55 ms and the open loop 84.95 ms; 10 s windows ranged 17.1 times apart; the mean was 4.47 ms. Caption: ask how the number was measured before you ask what it is.

The chart starts with the question this whole lesson is about: does the tester send on its own clock? Here are the same checks in words, each with the number from this lab behind it.

  1. Find out your tool's loop. Read its documentation or its source. Does a connection or a user wait for its answer before sending again? Here, locust did, and at 546 a second it reported 9.06 ms against the open loop's 84.95 ms.

  2. Time from the schedule. Use a tester that sends on its own clock and times each request from when it should have gone out. If you must use a closed or paced tester, apply the correction and say that you did. Here it reached 0.72 to 0.86 of the truth.

  3. Run long enough. Long enough to hold the rare slow event many times. Here, 10-second runs ranged 17.1 times apart, and 120-second runs 1.24 times.

  4. Report percentiles. The p50, the p99 and the max, never the mean alone. Here the mean was 4.47 ms while the p99 was 84.95 ms.

  5. Give a rate with its bound. Say "546 a second at a p99 of 5.29 ms", and say how the arrivals were sent. Here a closed loop said 1,098 a second, 2.01 times the rate that met a 10 ms bound.

  6. Say where the tester ran. Here it shared the one core, and took up to a fifth of it.

When Each Kind of Tester Is the Right One

A closed loop is right for one question: the most a service can answer. Keep enough connections busy and you learn its ceiling, as lessons 4 and 5 did. Its latencies are then mostly its own queue, so do not report them as what users see.

An open loop is right when you want to know what users will wait. Users do not wait for each other before they click. They arrive on their own, and an open loop copies that. Here it was the only tester whose p99 showed the stalls at every rate.

A paced loop, such as locust with constant_pacing, is right for scripted user journeys. In those, each step really does wait for the last: log in, then search, then buy. There it models one user honestly. But a fleet of such users still goes quiet together when the service stalls, so its percentiles can miss the stall. Here it reported 7.55 ms and 9.06 ms where the open loop said 84.95 ms.

The correction is right when you cannot change the tester. It moved the paced results most of the way, to 0.72 to 0.86 of the truth. It is an estimate, and it needs an expected interval you choose.

None of them is right at a rate the service cannot hold. At 1,074 arrivals a second, near this service's limit with stalls, the open loop kept up in 1 of 5 rounds. Its p99 only described a growing queue.

What This Lab Cannot Tell You

Two columns titled shows and cannot show. Shows: one small model on one 1-vCPU box; the tester shares the core. One kind of stall: the event loop sleeps 100 ms, giving the core away. Steady rates for 30 s (two runs of 120 s), 5 rounds. Cannot show: a tester on its own machine, or a real network between them. A stall that keeps the core busy, as some garbage collection pauses do, which would slow the tester too. Traffic that rises and falls, or stalls rarer than once a minute.

One core, and the tester on it. The account's limit made this a one-core lesson. With the tester on another machine, a real network would add its own delays and the service would get the whole core.

One kind of stall. My stall sleeps, so it gives the core to the tester. A stall that burns the core would also slow the tester on the same core, which is a different lie that I did not test.

Rates near the limit. The open loop at 1,074 a second, the rate the one-connection tester reached, kept up in only 1 of 5 rounds. So for the closed loop with one connection I have no clean "true" p99 to compare with, only a queue that grew for the whole run.

Labelled changes after the results. My first check of the stalls compared every PLANNED stall with the ones that happened. But the service plans stalls up to 2 seconds past the end of a run. So in 12 of 55 armed runs, one planned stall fell after the run had ended and was counted as missing. I changed the check to "every stall planned inside the counted part happened", and then it held. The leak scan's address pattern was narrowed after it matched a version string, as the report slide says. The step 2 and step 5 recordings were each the second try.

What to Do on Monday

A hand-drawn grid of six cards, titled five habits for an honest load test. 1, know the loop: does your tool wait for answers? read its docs or source. 2, time from the plan: latency from when a request should have gone out. 3, run long: long enough to hold the rare event many times. 4, percentiles: p50, p99 and max; never the mean alone. 5, rate with a bound: 546 a second at p99 5.3 ms, not 1,098 a second. The reason: at 546 a second locust said 9.06 ms; arrivals on their own clock said 84.95 ms. Caption: a load test measures the requests it sends; make it send them when users would.

If you take one thing to work on Monday, open your last load test report and find out how its tool sends requests. If each connection or user waits for its answer before the next one, its p99 may be hiding every pause the service had.

Then run the same test with an open-loop tester at the same rate, for long enough, and put the two p99s side by side. If they agree, you have learned that your service has no long pauses at that rate. If they do not, you have found the pauses your users already feel.

The one idea to keep: a tester can only time the requests it sends. If it stops sending when the service stops answering, the slow moments never reach the report.

Knowledge Check

Knowledge Check

4 questions - Score 80% to pass

Q1

With the stall on, the closed loop with one connection reported a p99 of 0.97 ms. Why did it miss the 100 ms stalls?

Q2

At 546 requests a second with the stalls on, which tester's p99 showed the stalls?

Q3

What did the HdrHistogram correction do to the paced client's p99?

Q4

A closed loop with 16 connections reported 1,098 answers a second. What is the honest way to state this service's capacity, by this lesson?

bound

What it costs. A c7g.medium costs $0.0363 an hour on demand in us-east-1, from AWS's price list, and AWS bills it by the second. My box was on for 1.411 hours, from launch to terminated, which is $0.0512. Its 16 GiB disk, deleted with it, added $0.0025, so $0.0537 in all. These numbers are in results/load-cost.json.

What the recordings hide. Every line passed through a small filter, load_redact.py, a copy of lesson 5's wrk_redact.py. In the AWS recordings it hides my account number and every IP address. It also hides every name AWS makes up for a resource, host names, my user name and anything from a key file. The two VS Code shots, taken on my laptop, are not filtered and show my own prompt. Where you see <ip> or i-<id>, your terminal shows the real value. Lines that start with + are the shell printing each command just before it runs it.

Keep the key file outside any git folder. Mine lives in ~/lab-data/serving/keys/, and the last step deletes it.

Step 3: rent the box. First the script asks AWS's parameter store (SSM) for the newest Ubuntu 24.04 image for arm processors. So you never copy an id that has gone out of date. Then run-instances asks for one c7g.medium with a 16 GiB disk that is deleted with the box, with the tags on both.

A real terminal recording of step 3, with the image, group and instance ids replaced by placeholders. aws ssm get-parameter reads the Ubuntu 24.04 arm64 image id into AMI; aws ec2 run-instances with that image, type c7g.medium, the key ai-research-course-load, a 16 GiB gp3 disk deleted on termination and the tags Project, Lesson and Name, asking only for the instance id, which goes into ID; describe-instances prints the id, c7g.medium, pending and the image id.

The box started in the "pending" state, which is normal for the first few seconds. These are the two commands:

  AMI=$(aws ssm get-parameter --query Parameter.Value --output text \
    --name /aws/service/canonical/ubuntu/server/24.04/stable/current/arm64/hvm/ebs-gp3/ami-id)
  ID=$(aws ec2 run-instances --image-id "$AMI" --instance-type c7g.medium --count 1 \
    --key-name ai-research-course-load --security-group-ids "$SG" \
    --block-device-mappings '[{"DeviceName":"/dev/sda1","Ebs":{"VolumeSize":16,"VolumeType":"gp3","DeleteOnTermination":true}}]' \
    --tag-specifications \
    "ResourceType=instance,Tags=[{Key=Project,Value=ai-research-course},{Key=Lesson,Value=load},{Key=Name,Value=ai-research-course-load}]" \
    "ResourceType=volume,Tags=[{Key=Project,Value=ai-research-course},{Key=Lesson,Value=load}]" \
    --query "Instances[0].InstanceId" --output text)
  echo "$ID" > "$STATE/id"

The last line writes ID into a small file in ~/lab-data/serving/load/, as step 2 did for SG, so the later steps can read them back.

The three sha256 fingerprints match the ones lesson 2 recorded, so the box serves the same model. One line is worth a note: 26,818 of the 26,851 reference scores equal the scores lesson 2 stored, not all of them. The other 33 differ in the last bits, as the packaging chapter saw between machines with different library builds. The box's own reference is what every answer here is checked against. These are the main commands of the step, run after step 4 has set IP:

  S=(ssh -i "$KEY" "${SSHO[@]}" ubuntu@"$IP")
  C=(scp -q -i "$KEY" "${SSHO[@]}")
  "${S[@]}" "mkdir -p $B/raw $B/features/results $B/serving/examples ~/lab-data/features"
  "${S[@]}" "curl -LsSf https://astral.sh/uv/install.sh | sh > /dev/null 2>&1 && \
    ~/.local/bin/uv python install 3.13.15 > /dev/null 2>&1 && \
    ~/.local/bin/uv venv --allow-existing --python 3.13.15 $B/venv > /dev/null 2>&1"
  "${C[@]}" "$HERE/examples/lat_requirements.txt" ubuntu@"$IP":$B/
  "${S[@]}" "~/.local/bin/uv pip install --python $B/venv/bin/python -q -r $B/lat_requirements.txt"
  "${S[@]}" "~/.local/bin/uv venv --allow-existing --python 3.13.15 $B/venv-locust > /dev/null 2>&1 && \
    ~/.local/bin/uv pip install --python $B/venv-locust/bin/python -q locust==2.46.7 numpy==2.5.3 pandas==3.0.6 \
    pyarrow==25.0.1 && $B/venv-locust/bin/locust --version"

The full list of files to copy is in the copy step of load_aws_steps.sh.

Step 6: stop the timers and look at the box. Ubuntu runs small jobs on a timer, such as checking for updates. One of those starting in the middle of a run would land in the results, so the script stops every timer first. Then it prints lscpu, nproc, uptime and every installed package with its version.

A real terminal recording of step 6, with the address replaced by a placeholder. Over ssh: stop every active timer and the unattended-upgrades service, after which systemctl prints 0 timers listed; lscpu prints architecture aarch64, 1 CPU, vendor ARM, model name Neoverse-V1, 1 thread per core, 1 core per socket; nproc prints 1; uptime prints up 4 min with load average 0.18, 0.20, 0.09. Then the box info is saved, and Python 3.13.15 and the full pip freeze are printed, including fastapi 0.142.2, numpy 2.5.3, scikit-learn 1.9.1, uvicorn 0.54.0 and uvloop 0.23.0.

The box reports one CPU, a Neoverse-V1, and the same package versions as lessons 2 to 6. These are the main commands that stop the timers and look:

  ssh -i ~/lab-data/serving/keys/ai-research-course-load.pem "${SSHO[@]}" ubuntu@"$IP" \
    'for t in $(systemctl list-units --type=timer --state=active --no-legend --plain | cut -d" " -f1); do
       sudo systemctl stop "$t"; done; sudo systemctl stop unattended-upgrades.service
     systemctl list-timers --no-pager | tail -2'
  ssh -i ~/lab-data/serving/keys/ai-research-course-load.pem "${SSHO[@]}" ubuntu@"$IP" 'lscpu | head -12; nproc; uptime'

Step 7: start the schedule, then leave. The schedule starts in the background with setsid nohup, so it keeps running after SSH disconnects. Then I log out, and nobody logs in until it ends.

A real terminal recording of step 7, with the address replaced by a placeholder. Over ssh: cd load, then setsid nohup bash load_schedule.sh 5 30 120 with its output sent to raw/schedule.log, in the background; two seconds later the log shows schedule start K=5 seconds=30 long=120.

The log's first line appeared two seconds after the start. The three numbers are the rounds, the seconds of a normal run, and the seconds of a long run. This is the command:

  ssh -i ~/lab-data/serving/keys/ai-research-course-load.pem "${SSHO[@]}" ubuntu@"$IP" \
    "cd load && (setsid nohup bash load_schedule.sh $K $SEC $LONG > raw/schedule.log 2>&1 < /dev/null &); \
     sleep 2; cat raw/schedule.log"

The 260 files came to 13 MB, small because the tester stores every run in a compact form. This is the copy command:

  rsync -a -e "ssh -i $KEY ${SSHO[*]}" ubuntu@"$IP":load/raw/ ~/lab-data/serving/load/raw/

Step 13: delete the box, the key and the door rule, and check. A box you forget keeps costing money. The script deletes the box, waits until AWS says "terminated", then deletes the key pair and the security group, and asks AWS again.

A real terminal recording of step 13, with the instance and group ids replaced by placeholders. terminate-instances prints shutting-down; wait instance-terminated; delete-key-pair prints true; delete-security-group prints true. Then the checks: describe-instances prints terminated; the count of key pairs named ai-research-course-load is 0; the count of security groups of that name is 0; the count of volumes tagged Lesson=load is 0; the count of course instances not terminated is 0. Last, the key file is removed.

Every count printed 0 and the box printed terminated, so nothing from this lesson was left in the account. These are all the teardown commands:

  aws ec2 terminate-instances --instance-ids "$ID" --query "TerminatingInstances[0].CurrentState.Name" --output text
  aws ec2 wait instance-terminated --instance-ids "$ID"
  aws ec2 delete-key-pair --key-name ai-research-course-load --query Return
  until aws ec2 delete-security-group --group-id "$SG" --query Return 2>/dev/null; do sleep 10; done
  aws ec2 describe-instances --instance-ids "$ID" --query "Reservations[0].Instances[0].State.Name" --output text
  aws ec2 describe-key-pairs --filters "Name=key-name,Values=ai-research-course-load" --query "length(KeyPairs)"
  aws ec2 describe-security-groups --filters "Name=group-name,Values=ai-research-course-load" \
    --query "length(SecurityGroups)"
  aws ec2 describe-volumes --filters "Name=tag:Lesson,Values=load" --query "length(Volumes)"
  aws ec2 describe-instances --filters "Name=tag:Project,Values=ai-research-course" \
    "Name=instance-state-name,Values=pending,running,stopping,stopped,shutting-down" \
    --query "length(Reservations[].Instances[])"
  rm -f ~/lab-data/serving/keys/ai-research-course-load.pem

The security group can only be deleted once the box is fully gone, so the script retries that command every ten seconds until AWS accepts it.

Other common tools differ in their loop. ApacheBench, ab, keeps a fixed number of requests in flight: its -c is the "Number of multiple requests to perform at a time", a closed loop. wrk2 is the opposite. Its README, at commit 44a94c1, says it is "wrk modifed to produce a constant throughput load" and measures "from the time the transmission should have occurred". It is not on pip. I ran neither ab nor wrk2; both descriptions come from their documentation. My own open loop, from lesson 4, works the way wrk2 describes.

"G10. corrected c1-on p99 between 0.5 and 1.0 x oc-on p99 in every round (the correction does not add the queue that builds after the stall)." Wrong: 0.03 to 0.04, because the open loop at that rate never emptied its queue.

  • "G11. c64-on p99 above oc-on p99 in every round; c16-on p99 below 50 ms in every round." Wrong in its first half, for the same reason; c16-on was 15.6 ms.

  • "G12. o5-on mean below 0.25 x its p99, and above its p50, in every round." Right.

  • "G13. first-10 s p99 of o5-on-long spans at least 3 x across the 5 rounds (max / min); its 120 s p99 at most 1.3 x. Stall off (o5-off-long): first-10 s p99 within 1.5 x of the 120 s p99 in every round." Right: 17.1 times, 1.24 times, and 0.98 to 1.06.

  • "G14. o5-off-nowarm's max above o5-off's in at least 4 of 5 rounds; their p99 cannot be told apart." Wrong: 3 of 5.

  • "G15. the highest stall-off rate with p99 at most 10 ms in all 5 rounds is 0.5 x S; c16-off reports at least 1.8 x that rate in answers per second; o9-off keeps up but its median p99 is above 20 ms." Right.

  • "G16. my testers' CPU at most 0.15 of the run in every run; locust's CPU per request at least 2 x my paced client's, told apart." Right: at most 0.094, and 4.44 to 4.55 times.

  • "G17. o5-off send lag p99 below 1 ms in every round." Right: 0.56 to 0.63 ms.

  • "G18. every score equals the reference to the bit." Right: 2,729,222 of 2,729,222.

  • r"""Is your load test lying? Two testers, one stalling service, on YOUR machine.
    
    Lesson 7 of 'Serving and Inference Basics'. It uses lesson 2's Python environment, made once inside this folder
    (Python 3.13 is what the lab used):
        python3 -m venv venv-sv
        source venv-sv/bin/activate              # Windows: venv-sv\Scripts\activate
        pip install -r lat_requirements.txt      # numpy, pandas, pyarrow, scikit-learn, fastapi, uvicorn, ...
    It also needs the shop data from the features chapter: run python ../../features/fetch_data.py once first. Then:
        python load_demo.py                  # train, start a small service, run both testers, print the table
        python load_demo.py --save out.json  # and save the numbers
    If a package is missing, it prints one line saying why and stops.
    The timings it prints come from YOUR machine, as it is right now: other programs move them. Read them for what the
    two testers say about the SAME service on your machine, not as numbers to set beside the lab's.
    It starts one helper process (the service, on 127.0.0.1 only) and stops it at the end. It writes nothing except a
    temporary model file, deleted at the end, and out.json if you ask.
    
    Design, written 2026-10-08 after the lab's design (load_testing.py) and before this file first ran:
      1. Train the features chapter's model (it must score test AP 0.5450), as lesson 4's bat_demo.py does.
      2. Start a small FastAPI service in a second process (this same file, run with --serve): POST /score takes one
         row's six features and answers predict_proba. It STALLS on purpose: every 2 s its event loop sleeps 100 ms,
         so nothing is answered for that tenth of a second (the lab's stall is the same size, at random times).
      3. Tester A, a CLOSED loop: one connection; it sends the next request the moment the last answer arrives, for
         10 s. Latency = answer minus send, as most load testers record it.
      4. Tester B, an OPEN loop: requests arrive at random (Poisson) at HALF of A's measured rate, for 10 s, and wait
         in a queue for one of 32 sender threads. Latency = answer minus the SCHEDULED arrival, so waiting counts.
      5. Print, for both: requests, p50, p99, max and mean; the share of requests that waited behind a stall; and A's
         p99 after the HdrHistogram correction (for each recorded latency v and expected interval I = A's median, also
         count v - I, v - 2I, ... while they are at least I).
    
    Author: Roni Das
    Created: 2026-10-08
    """
    import importlib.util
    import json
    import os
    import socket
    import subprocess
    import sys
    import tempfile
    import threading
    import time
    from pathlib import Path
    from queue import Queue
    from time import perf_counter
    
    NEEDED = ("numpy", "pandas", "pyarrow", "sklearn", "fastapi", "uvicorn")
    missing = [m for m in NEEDED if importlib.util.find_spec(m) is None]
    if missing:
        sys.exit(f"load_demo.py needs {', '.join(missing)}: make the environment in its docstring "
                 f"(pip install -r lat_requirements.txt), then run it with that environment's python.")
    
    import numpy as np  # noqa: E402
    
    SECONDS = 10.0
    STALL_EVERY_S, STALL_MS = 2.0, 100.0
    
    
    # ───────────────────────────── the service (python load_demo.py --serve PORT MODEL) ─────────────────────────────
    def serve(port: int, model_path: str) -> None:
        import asyncio
        import contextlib
        import pickle
    
        import uvicorn
        from fastapi import FastAPI, Request
        from starlette.responses import Response
    
        os.environ["OMP_NUM_THREADS"] = "1"
        model = pickle.loads(Path(model_path).read_bytes())
    
        def stall() -> None:
            time.sleep(STALL_MS / 1000)                    # the loop is blocked: nothing is read or answered
            asyncio.get_running_loop().call_later(STALL_EVERY_S, stall)
    
        @contextlib.asynccontextmanager
        async def lifespan(_app):
            asyncio.get_running_loop().call_later(STALL_EVERY_S, stall)
            yield
    
        app = FastAPI(lifespan=lifespan)
    
        @app.post("/score")
        async def score(request: Request) -> Response:
            d = json.loads(await request.body())
            p = float(model.predict_proba(np.array([d["x"]], dtype=np.float64))[0, 1])
            return Response(content=json.dumps({"id": d["id"], "score": p}).encode(), media_type="application/json")
    
        @app.get("/health")
        async def health() -> Response:
            return Response(content=b"ok")
    
        uvicorn.run(app, host="127.0.0.1", port=port, log_level="warning")
    
    
    # ───────────────────────────── the model ─────────────────────────────
    def build():
        """The features chapter's model and its 26,851 test rows (lesson 4's bat_demo.py build step)."""
        from sklearn.metrics import average_precision_score
        here = Path(__file__).resolve().parent
        sys.path.insert(0, str(here.parents[1] / "features"))
        import task
        from what_a_feature_is import HAND_COLS, hgb, joined
    
        ev = task.load_events()
        lab_tr, _, lab_te = task.splits(ev)
        tr = joined(ev, lab_tr, task.TRAIN_CUTOFFS)
        te = joined(ev, lab_te, task.TEST_CUTOFFS)
        model = hgb(0).fit(tr[HAND_COLS].to_numpy(float), tr["label"].to_numpy())
        X = np.ascontiguousarray(te[HAND_COLS].to_numpy(np.float64))
        p = model.predict_proba(X)[:, 1]
        y, cut = te["label"].to_numpy(), te["cutoff"].to_numpy()
        ap = float(np.mean([average_precision_score(y[cut == c], p[cut == c]) for c in np.unique(cut)]))
        assert round(ap, 4) == 0.5450, ap
        return model, X, ap
    
    
    # ───────────────────────────── the two testers ─────────────────────────────
    def connect(port: int):
        import http.client
        c = http.client.HTTPConnection("127.0.0.1", port, timeout=30)
        c.connect()
        c.sock.setsockopt(socket.IPPROTO_TCP, socket.TCP_NODELAY, 1)
        return c
    
    
    def ask(c, j: int, x: list[float]) -> None:
        c.request("POST", "/score", body=json.dumps({"id": j, "x": x}), headers={"Content-Type": "application/json"})
        r = c.getresponse()
        d = json.loads(r.read())
        if d["id"] != j:
            raise SystemExit(f"answer {j} carries id {d['id']}")
    
    
    def closed_loop(port: int, rows: list[list[float]]) -> tuple[list[float], list[float]]:
        c = connect(port)
        for j in range(50):                                # warm-up, not counted
            ask(c, -1 - j, rows[j])
        sends, lats = [], []
        t0 = perf_counter()
        j = 0
        while perf_counter() - t0 < SECONDS:
            a = perf_counter()
            ask(c, j, rows[j % len(rows)])
            sends.append(a - t0)
            lats.append(perf_counter() - a)
            j += 1
        c.close()
        return sends, lats
    
    
    def open_loop(port: int, rows: list[list[float]], rate: float) -> tuple[list[float], list[float]]:
        rng = np.random.default_rng(0)
        at = np.cumsum(rng.exponential(1 / rate, size=int(rate * SECONDS * 1.5) + 50))
        at = at[at < SECONDS]
        lats = [0.0] * len(at)
        q: Queue = Queue()
        t0 = 0.0                                           # set just before the first arrival, below
    
        def sender() -> None:
            c = connect(port)
            ask(c, -1, rows[0])                            # warm-up of this connection, not counted
            while True:
                j = q.get()
                if j is None:
                    break
                ask(c, j, rows[j % len(rows)])
                lats[j] = perf_counter() - (t0 + at[j])    # from the SCHEDULED arrival
            c.close()
    
        pool = [threading.Thread(target=sender, daemon=True) for _ in range(32)]
        for th in pool:
            th.start()
        time.sleep(0.5)
        t0 = perf_counter() + 0.05
        for j, a in enumerate(at):
            d = t0 + a - perf_counter()
            if d > 0:
                time.sleep(d)
            q.put(j)
        for _ in pool:
            q.put(None)
        for th in pool:
            th.join()
        return list(at), lats
    
    
    def corrected(lats_ms: np.ndarray, interval_ms: float) -> np.ndarray:
        """HdrHistogram's recordValueWithExpectedInterval: also record v - I, v - 2I, ... while at least I."""
        extra = [np.arange(v - interval_ms, interval_ms - 1e-12, -interval_ms) for v in lats_ms if v > interval_ms]
        return np.concatenate([lats_ms] + extra) if extra else lats_ms
    
    
    def row(name: str, lat_ms: np.ndarray, n_sent: int, stalled: float) -> str:
        return (f"   {name:<22}{n_sent:9,d}{np.percentile(lat_ms, 50):9.2f}{np.percentile(lat_ms, 99):9.2f}"
                f"{lat_ms.max():9.1f}{lat_ms.mean():9.2f}{stalled:10.2%}")
    
    
    def main() -> None:
        try:
            load = f"{os.getloadavg()[0]:.2f}"
        except (AttributeError, OSError):
            load = "not available on this system"
        print(f"Every time below was measured on THIS machine ({os.cpu_count()} cores, 1-minute load {load}).")
        model, X, ap = build()
        print(f"model trained: test AP {ap:.4f}; {len(X):,} test rows")
        rows = X.tolist()
        import pickle
        with tempfile.TemporaryDirectory() as tmp:
            mp = Path(tmp) / "model.pkl"
            mp.write_bytes(pickle.dumps(model))
            s = socket.socket()
            s.bind(("127.0.0.1", 0))
            port = s.getsockname()[1]
            s.close()
            svc = subprocess.Popen([sys.executable, __file__, "--serve", str(port), str(mp)],
                                   env={**os.environ, "OMP_NUM_THREADS": "1"})
            try:
                for _ in range(300):
                    try:
                        connect(port).close()
                        break
                    except OSError:
                        time.sleep(0.1)
                print(f"service started; it stalls {STALL_MS:.0f} ms every {STALL_EVERY_S:.0f} s; each tester runs "
                      f"{SECONDS:.0f} s\n")
                a_send, a_lat = closed_loop(port, rows)
                rate = len(a_lat) / SECONDS / 2
                b_at, b_lat = open_loop(port, rows, rate)
            finally:
                svc.terminate()
                svc.wait()
        a = np.array(a_lat) * 1000
        b = np.array(b_lat) * 1000
        over = lambda v: float(np.mean(v >= STALL_MS / 2))                                  # noqa: E731
        med = float(np.median(a))
        ac = corrected(a, med)
        print(f"A sent {len(a) / SECONDS:,.0f} requests a second; B's random arrivals came at {rate:,.0f} a second")
        print("                            sent   p50 ms   p99 ms   max ms  mean ms  over 50 ms")
        print(row("A closed loop", a, len(a), over(a)))
        print(row("B open loop, half rate", b, len(b), over(b)))
        print(f"\nA after the HdrHistogram correction (I = A's median, {med:.3f} ms): "
              f"{len(ac) - len(a):,} values added, p99 {np.percentile(ac, 99):.2f} ms")
        if "--save" in sys.argv:
            out = {"load": load, "cores": os.cpu_count(), "test_ap": ap, "rate_b": rate,
                   "a": {"n": len(a), "p50": float(np.percentile(a, 50)), "p99": float(np.percentile(a, 99)),
                         "max": float(a.max()), "mean": float(a.mean()), "over50": over(a)},
                   "b": {"n": len(b), "p50": float(np.percentile(b, 50)), "p99": float(np.percentile(b, 99)),
                         "max": float(b.max()), "mean": float(b.mean()), "over50": over(b)},
                   "a_corrected": {"interval_ms": med, "added": len(ac) - len(a), "p99": float(np.percentile(ac, 99))}}
            Path(sys.argv[sys.argv.index("--save") + 1]).write_text(json.dumps(out, indent=1))
    
    
    if __name__ == "__main__":
        if len(sys.argv) > 1 and sys.argv[1] == "--serve":
            serve(int(sys.argv[2]), sys.argv[3])
        else:
            main()
    

    You saw the demo running on the lab's box in step 11. That exact run is stored in results/load-demo-run.txt. On the box, the closed loop sent 903 requests a second and reported a p99 of 1.14 ms. The open loop, with arrivals at 451 a second, reported 91.71 ms. The correction moved the closed loop's p99 to 77.59 ms.

    And here is the same demo in VS Code on my laptop, run with python load_demo.py in the examples folder with venv-sv active.

    A real screenshot of VS Code's terminal on my laptop, in the venv-sv environment, after python load_demo.py. It prints: measured on THIS machine, 10 cores, 1-minute load 5.16; model trained, test AP 0.5450, 26,851 test rows; the service stalls 100 ms every 2 s, and each tester runs 10 s. A sent 3,210 requests a second, and B's random arrivals came at 1,605 a second. A, the closed loop: sent 32,095, p50 0.29 ms, p99 0.35 ms, max 110.6 ms, mean 0.31 ms, 0.01% over 50 ms. B, the open loop at half rate: sent 16,178, p50 1.09 ms, p99 107.30 ms, max 124.2 ms, mean 6.32 ms, 5.16% over 50 ms. A after the HdrHistogram correction, with I = A's median, 0.293 ms: 1,465 values added, p99 83.03 ms.

    This run is from my laptop, a different machine from the box. It has 10 cores, and its 1-minute load was 5.16 because other programs were busy on it. So you cannot set its numbers beside the box's numbers one by one. Read the two lines against each other instead.

    The closed loop A sent 3,210 requests a second and reported a p99 of 0.35 ms, close to its p50 of 0.29 ms. Its max, 110.6 ms, shows that it did meet the stalls, once each. The open loop B, with random arrivals at half that rate, 1,605 a second, reported a p99 of 107.30 ms. And 5.16% of its requests took more than 50 ms, against 0.01% for A. The correction added 1,465 values to A and moved its p99 to 83.03 ms, still below B's.

    One more difference: the demo's service stalls every 2 s on a fixed beat, so a 10-second run holds about five stalls. The lab's service stalled at random moments, on average every 4 s. So the demo's numbers are not the lab's numbers, even on the same machine.

    load_files/load_locustfile.py is the locust test, with a hook that keeps every raw response time. load_files/load_run_once.sh starts a fresh service, waits for an idle core, records the load and the system log, and runs one tester. load_files/load_schedule.sh runs the calibration and then five rounds of twenty arms, each round in its own shuffled order.

    load_report.py does not import any of the above. It reads the compact copy of the raw files with its own code and recomputes every number. It checks the guesses, the demo's stored run, the fact-check, the teardown and the redaction, and it writes the playground.