Serving And Inference

A Serving Checklist: The Whole Chapter as One Service

0 of 35 complete

0%

Contents

Back|Serving And InferenceA Serving Checklist: The Whole Chapter as One Service
1/35
95 min left
  1. Home
  2. AI Engineering: Data, RAG and Agents
  3. Serving and Inference Basics
  4. A Serving Checklist: The Whole Chapter as One Service
Prerequisites
Cost per Prediction: What Do 1,000 Answers Really Cost?requiredTimeouts and Fallbacks: What Should a Service Return When Its Lookup Is Slow?requiredBatching Requests: Does It Help a Small CPU Model?requiredCold Start: The First Request Waited for Python to Import Its LibrariesrequiredLoad Testing Honestly: Is Your Load Test Lying?requiredWorkers and Threads: Processes, Threads, or Both on One Core?required
Related Topics
Model Signatures: The Right Numbers in the Wrong Shape, and What a Schema Check CatchesPackaging, Registry and VersioningMissing at Serving: No Fill Gives Back a Row That Is GoneFeatures and Feature Stores
1 of 35
Previous lessonCost per Prediction: What Do 1,000 Answers Really Cost?

System Design

  • Foundation
  • Intermediate
  • Advanced
  • Capstone

AI Engineering

  • Foundation
  • Data, RAG and Agents
  • Evaluation, LLM Ops and Security

systemdesign.academy

  • Home
  • Glossary
  • Interview prep
  • Reviews
  • About
  • Privacy
  • Terms

The Pilot's Walk Around the Plane

Before a small plane takes off, the pilot walks around it with a list on a clipboard. She looks at the propeller, the fuel, the tyres and the flaps on the wings, one line at a time. She has done this hundreds of times, and she still reads the list.

A flat illustration of a sunny airfield. A pilot in a uniform jacket with striped cuffs stands beside the nose of a small single-propeller plane. She holds a clipboard with a list of empty tick boxes in one hand and points at the propeller with the other. A toolbox sits on the grass and a windsock flies in the distance. Below the picture: every line on her list was written after something went wrong once.

She does not read the list because she forgets things. She reads it because each line was written after something went wrong once. Reading it takes two minutes; skipping a line can cost the whole flight.

This chapter has built a long list of that kind. Each lesson took one question about running a model for real users, and measured the answer on a rented computer. In this last lesson I put the answers together into one service and race it against a plain service that ignores them. Both run on one computer, under one stream of traffic, with one slow helper behind them.

Where This Lesson Starts

This is lesson 12, the last lesson of the chapter on serving a model. Serving a model means running it as a small program that answers questions while people wait. That program is called a service.

The model is the one from the features chapter. It reads six numbers about one customer of a real online shop, such as how many days have passed since their last order. It gives back a score between 0 and 1: how likely that customer is to buy again in the next 30 days.

Lessons 1 to 11 each changed one thing and measured what happened. None of them changed everything at once. That leaves a fair question, one a real team asks on its first day. If I build the service with every one of these choices, what do I get? And is it better than the simple version?

So the question of this lesson is this: what does a service built with the chapter's measured choices look like? And how does it compare with a plain one, measured in one way, on one machine?

The Words You Need First

Please read this slide slowly if a word is new. Each word comes with an everyday picture.

A hand-drawn, already solved crossword. Across, 1: CHECKED; 2: WARMUP; 3: NAIVE. Down, 1: COLD, sharing its first letter with CHECKED; 4: STANDIN, crossing WARMUP, CHECKED and NAIVE. The clues below: 1 across, checked, the service built with every choice the chapter measured; 2 across, warm-up, one practice request, done before the service opens its door; 3 across, naive, the service as first written, with none of those choices; 1 down, cold, a program that has only just started, so every first step is slow; 4 down, stand-in, a ready-made answer, used when the real one comes too late.

I call the service as it was first written in lesson 2 the naive service. Naive does not mean stupid. It is the version most people write first, because it works on a laptop. I call the same program with the chapter's useful choices added the checked service.

A worker is one copy of the service program. A busy shop can run several copies side by side, like several cashiers. A thread is one line of work inside a program. With one thread, the model does its sums one after another, like one cook in a kitchen. Both services here ran one worker with one thread.

A program is cold when it has only just started. Its first steps are slow, because Python is still loading its tools. A warm-up is one practice request the service sends to itself before it lets real users in. Then the first real user does not pay for those slow first steps.

A live score is one the model works out now, for this request. A stand-in is a ready-made answer the service gives when the live one would come too late. Here it is the customer's score from one month earlier, as in lesson 9. It is like a waiter who brings the soup of the day when the kitchen is too slow for what you ordered.

A model call is one time the program asks the model for scores. A batch is several rows of numbers sent in one call.

The time a request waits for its answer is called its . Waiting times are in milliseconds (ms): a millisecond is a thousandth of a second, so 25 ms is about a quarter of a blink. Users notice the slowest answers most, so I line up every request's wait from shortest to longest. The p99 is the time that 99 of every 100 requests beat. I will mostly call it "the slowest 1 in 100". The max is the single slowest answer of all.

My target is lesson 11's: the slowest 1 in 100 answers within 25 ms. I add one more condition, explained on the next slides: at least 95 of every 100 answers must be live scores.

The Headline: What Was Built In, and What Was Measured

Here is the main result first, as two departure boards at an airport. Each row is one rate, the number of requests sent each second. A row is on time if the slowest 1 in 100 came within 25 ms and at least 95 in 100 answers were live.

Two split-flap departure boards hung from the ceiling on rods, one headed NAIVE and one headed CHECKED, with columns for the rate, the slowest 1 in 100 in ms, and a status flap. Naive: 150, 33.0; 300, 39.9; 450, 34-1,099, a range because the rounds split; 600, 271; 700, 461; 800, 808; 900, 2,064; 1,000, 4,105; 1,100, 6,218; 1,200, 7,833; every status LATE TIME. Checked: 150, 23.8, ON TIME; 300, 23.8, ON TIME; 450, 24.0, 4 OF 5; 600, 24.7, 1 OF 5; 700, 24.9, LATE LIVE; 800, 25.0, LATE BOTH; 900, 25.2; 1,000, 25.3; 1,100, 25.5 and 1,200, 25.7, each LATE BOTH.

The naive service was never on time, and that was built into the test. With this helper, about 1.6 lookups in 100 take longer than 25 ms by themselves: about 1 that hangs, and about 0.6 that are only slow. On the naive service's own clock, 1.2% to 1.9% of lookups took over 23 ms even at the gentlest rate, 150 a second, in every round. A service that waits for every lookup cannot keep its slowest 1 in 100 under 25 ms with a helper like this. So its failure is not a discovery. It is what the helper was made to do.

The checked service met the target up to 450 requests a second in the middle round. Its highest passing rate was 450, 600, 300, 450 and 450 a second in the five rounds, which is about $0.0224 per million answers at lesson 11's price. Its time limit makes its fast answers partly built in too: it stops waiting at 23 ms, so its slowest answers had to sit near 25 ms.

What set the checked service's limit was the helper, not the processor. The helper can work on 16 lookups at once, and above about 600 a second it filled up. More and more answers became stand-ins: 11.1% at 600 a second and 90.2% at 1,200, in the middle rounds. From 800 a second the slowest 1 in 100 also crept just past 25 ms, so those runs failed on both counts.

The checks cost almost nothing. Memory and processor use were nearly equal at 450 a second, and both services gave their first answer about 1.7 seconds after starting, too close to call.

Every live score matched, to the last digit. On 228,096 rows that both services answered with a live score, not one score differed.

Where Each Difference Lives

A request walks through five stops. It leaves the tester, reaches the service, waits for a lookup of the customer's six numbers, gets scored by the model, and its answer travels back. The tester is the program that plays the users: it sends requests and times every answer.

A hand-drawn journey from left to right in five stops, made with the Excalidraw library's own drawings, joined by arrows: 1 start, a server; 2 a request, an envelope; 3 the lookup, a database drum; 4 the model, a small network of dots; 5 the answer, a notepad. Notes under each stop: warm-up first (lesson 6), with 1 worker, 1 thread: both; planned arrival times (lesson 7); 25 ms limit, then last month's score (lesson 9); up to 32 rows in one call (lesson 4); timed from the planned arrival (lesson 7). Below: stops 2 and 5 belong to the tester, for both services; the naive service has none of the notes at 1, 3 and 4.

Stops 2 and 5 belong to the tester, and they work alike for both services. The tester plans every request's arrival time before the run starts, at random, like customers walking into a shop. It sends each request at its planned time, even if earlier answers are late, and it times each answer from the planned moment. Lesson 7 showed why: a tester that waits politely for each answer before sending the next one hides the slow moments.

On the box, the two services differed at three stops only. At the start, the checked service warms up. At the lookup, it gives up after its time limit and answers a stand-in. At the model, it scores every waiting row in one call. One worker and one thread were true of both services, on this box.

A sequence diagram with three columns, tester, checked service and feature store, and five numbered arrows. 1, tester to service: request arrives. 2, service to store: look up 6 numbers. 3, a loop on the service: 23 ms pass, with the note stop waiting. 4, service to tester: last month's score. 5, a dashed arrow from the store to the service: 1 to 2 s later, with the note nobody is waiting. Below: the tester gets an answer in about 24 ms; the store keeps its slot busy until the hang ends.

This little sequence shows one stand-in answer from start to finish. The lookup is one of the 1 in 100 that hang. After 23 ms the checked service stops waiting and sends last month's score, so the tester has an answer in about 24 ms. The helper does not know that nobody is waiting. It keeps working on the lookup for 1 to 2 seconds and only then frees its slot.

How the Lab Was Built

Before any run, I wrote the lab's plan at the top of its program file, scripts/labs/serving/serving_checklist.py. That note at the top of a Python file is called a docstring. It names every setting, the lesson each one comes from, and the rule for "clearly different". It also lists twelve guesses, written down first so they could be checked later, and so they could be wrong.

Everything ran on one rented c7g.medium from Amazon Web Services (AWS), the kind of machine lessons 2 to 11 used. It has one core: one worker inside the chip that runs one step of a program at a time. The tester, the service and the slow helper all shared that one core, as in every lesson of this chapter.

The slow helper is lesson 9's pretend : a separate small program that holds the customers' numbers. It can work on 16 lookups at once, like a kitchen with 16 cooks. Most lookups take about 2 ms. But 1 lookup in 100 hangs: it freezes for a time between 1 and 2 seconds, like a cook who disappears into the back room. Lesson 9 chose the 16 slots so that the store would be busy at 300 requests a second and near its limit at 450.

Each service is one arm of the test, like one recipe in a cooking contest. The naive arm is lesson 2's service as first written. It needed one change so that it could talk to the slow store: it asks the store over the network, instead of reading a small database inside itself.

Two short assembly lines side by side, each a conveyor belt with three stations on it: a door, a lookup drum and a model gear. Above both lines, inside one dashed outline, a person shape labelled both: 1 worker, and a wavy line labelled both: 1 thread. The left line, naive service, has nothing else. The right line, checked service, has three extra objects above its stations: an envelope with an arrow down to the door, labelled warm-up; a stopwatch with a plate under it, labelled time limit, stand-in, above the lookup; and a stack of four rows on a tray, labelled batch tray, above the model.

The two shared parts sit once, across both lines: one worker and one thread were true of both services on this box. The naive line has nothing else. The checked line carries three extra objects. A practice request drops in at the door: that is the warm-up. A stopwatch with a stand-in plate sits at the lookup: that is the time limit. A tray of rows waits at the model: that is the batcher.

The Traffic, the Rounds and the Rule

Every run began cold. The tester itself started the service as a fresh program and wrote down that moment. Then it knocked on the service's door, the port, every thousandth of a second until the door opened. Next it opened 512 connections, which are like 512 phone lines kept open to the service, and started the traffic 2 ms later. The tester sent no warm-up of its own, so the first requests met the service as it woke up. A run's first answer time runs from the moment the program started to the moment its first answer came back.

Each run sent requests for 15 seconds at one rate: 150, 300, 450, 600, 700, 800, 900, 1,000, 1,100 or 1,200 a second. Both arms ran every rate.

The runs were paired. In each round, both arms at one rate got identical planned arrival times. They got the customers in one order, and the store gave them identical slow lookups. It is like two runners on one track in one wind. So when the two arms differ, the difference comes from the service, not from luck in the traffic.

A round is one pass through all 20 runs (2 arms times 10 rates). Each round used a new random order, so the time of day could not favour one arm. I ran 5 rounds: 100 counted runs.

A run passes if three things hold. The slowest 1 in 100 took at most 25 ms. Every request got an answer. And at least 95 of every 100 answers were live scores. I added that last condition before the runs, for a simple reason: a stand-in comes back quickly by design. Without the condition, a service that answered everyone with a stand-in would pass while telling nobody anything new.

In each round, an arm's max rate is the highest rate that passed while every lower rate also passed, as in lesson 11.

For two arms, I call a difference clearly different only if one arm beat the other in all 5 rounds. Even its worst round must also beat the other's best. Otherwise it is too close to call.

The Machine, and Proof It Was Quiet

The lab ran on one c7g.medium in AWS's us-east-1 region. It has one core on a Graviton3, a chip AWS designs itself, and 2 GiB of memory, about 2.1 billion bytes. Its own report of the chip, from the command lscpu, said Neoverse-V1: the name of the design Arm made, which Graviton3 is built from. Python was 3.13.15, with the library versions of lessons 2 to 11.

Before every timed run, the core had to sit at least 95% idle for two seconds. All 100 counted runs met that, and the lowest reading was 99.01% idle. Steal, the time the bigger real computer underneath lends to other customers' rented machines, was 0.00% in every run. Nobody was logged in during any counted run. Before and after each run, the list of installed packages was turned into a short code. This fingerprint changes if anything in the list changes, like a seal on a jar. All 200 fingerprints matched.

A hand-drawn architecture sketch with real product logos, inside a dashed frame labelled one c7g.medium: 1 core. A Tester box with the Python logo, plans each arrival first, sends a request to a Service box with the FastAPI logo, one worker, one thread. From the service, an arrow labelled look up 6 numbers goes to a Feature store drum with the SQLite logo, 16 slots, 1 in 100 hangs; an arrow labelled rows that came back goes to a Model box with the scikit-learn logo, up to 32 rows a call; and a dashed arrow labelled time ran out goes to a box marked last month's score, the stand-in. Marker notes: warm-up before the door opens, at the service; 25 ms per request, counted from when it starts on it, at the stand-in. The bottom strip says both services have these parts, and the checked one adds the warm-up, the time limit with its stand-in, and the batcher.

This hand-drawn sketch shows what sat on the box. The tester plays the users. The service asks the store for six numbers. If they come back in time, the model scores them, up to 32 rows in one call. If time runs out, the answer is last month's score. All of it shares one core, so a busy tester or a busy store slows the service down too.

One Box's Life on One Line

The whole life of the rented machine fits on one time line. The top bar is drawn to scale; the flags below show every step in order, each with its real start time.

A time line of the rented box in UTC. At the top, drawn to scale: a short setup bar, a long hatched bar labelled first schedule stopped, nothing counted, 12:16 to 12:55, a long solid bar labelled 100 counted runs, nobody logged in, 12:57 to 13:33, and a short end bar. Below, the steps in order, not to scale, each with its start time and, where it applies, a real mark: 1 free?, 12:11, AWS; 2 key, door, 12:11, AWS; 3 rent, 12:11, AWS; 4 wait, 12:11; 5 copy, 12:14, Python; 6 quiet, 12:15, Ubuntu; 7 start, 12:16; 7b restart, 12:55; 8 watch, 13:00, AWS; 11 demo, 13:34, Python; 12 fetch, 13:34; 13 delete, 13:35, AWS.

Steps 1 to 7 rent the machine, set it up and start the schedule. The hatched stretch is the first schedule, which stopped on its very first counted run: 39 minutes in which nothing was counted. Step 7b is the restart, and the next slides explain the fault and the fix. Step 8 is how I watched without logging in. The long dark bar holds the 100 counted runs, steps 9 and 10.

In those runs the schedule started a fresh store and a fresh service each time, and the tester sent its timetable of requests. Steps 11 to 13 come after: the student demo, the copy of the results to my laptop, and the deletion of everything.

Build the Lab Machine Yourself

The rented machine is part of the lab, so I show every step. You can rent this kind of machine and run this schedule yourself.

You need an AWS account, and the AWS command line tool (aws) set up with your keys. You also need ssh, which logs in to another computer, scp and rsync, which copy files to and from it, and curl, which downloads a web page.

Lesson 2's model and data files must sit in ~/lab-data/serving/lat, which lesson 2's lab makes. Lesson 9's two stand-in files must be in place too, and python serving_checklist.py prepare copies them. All the steps live in one file, scripts/labs/serving/chk_files/chk_aws_steps.sh. You run it from the scripts/labs/serving folder, one step at a time, such as bash chk_files/chk_aws_steps.sh check, then access, and so on. Every code block below is copied word for word from that file.

The blocks use a few names that the file sets at its top. If you paste the blocks by hand, run these lines first, from the scripts/labs/serving folder. All of them come from the file, except HERE=$PWD. The file works out that folder from its own location, which a pasted line cannot do.

Steps 1 to 3: Is the Account Free? Lock the Door, Rent the Machine

My account allows one core at a time, and other projects share that limit. So step 1 counts every machine in the account that is starting, running or stopping, and it must print 0.

Step 2 makes a key and a locked door. A key pair is how you prove to the machine that you may come in. AWS keeps one half, and you keep the other half in a file only you can read. A security group is a door rule. Mine opens only port 22, the door that SSH (the secure remote login) uses, and only to my own address. Both get tags, small labels that say which project and lesson they belong to.

Two real terminal recordings. Step 1: aws ec2 describe-instances, counting every instance that is pending, running, stopping or shutting down, prints 0. Step 2, with the address, network and group ids replaced by placeholders: create-key-pair, a kind of key AWS calls ed25519, for ai-research-course-chk with the tags Project and Lesson writes the key straight to its file; chmod 600; curl reads my address; describe-vpcs finds the default network; create-security-group makes the door rule with its tags; authorize-security-group-ingress prints a small table, from my address /32, port 22; last, ls shows the key file, readable only by its owner.

Step 1 printed 0, so the account was free. In step 2 the key went straight into its file and was never printed, and the door rule names only my own address, written with /32 after it. Step 1 is a single command:

  aws ec2 describe-instances \
    --filters "Name=instance-state-name,Values=pending,running,stopping,shutting-down" \
    --query "length(Reservations[].Instances[])"

Step 2 runs these:

  aws ec2 create-key-pair --key-name ai-research-course-chk --key-type ed25519 \
    --tag-specifications "ResourceType=key-pair,Tags=[{Key=Project,Value=ai-research-course},{Key=Lesson,Value=chk}]" \
    --query KeyMaterial --output text > ~/lab-data/serving/keys/ai-research-course-chk.pem
  chmod 600 ~/lab-data/serving/keys/ai-research-course-chk.pem
  MYIP=$(curl -sf https://checkip.amazonaws.com)
  VPC=$(aws ec2 describe-vpcs --filters Name=isDefault,Values=true --query "Vpcs[0].VpcId" --output text)
  SG=$(aws ec2 create-security-group --group-name ai-research-course-chk --vpc-id "$VPC" \
    --description "lesson 12 serving checklist, ssh from one address" \
    --tag-specifications "ResourceType=security-group,Tags=[{Key=Project,Value=ai-research-course},{Key=Lesson,Value=chk}]" \
    --query GroupId --output text)
  aws ec2 authorize-security-group-ingress --group-id "$SG" --protocol tcp --port 22 --cidr "$MYIP/32" \
    --query "SecurityGroupRules[].{port:FromPort,from:CidrIpv4}" --output table
  ls -l ~/lab-data/serving/keys/ai-research-course-chk.pem
  echo "$SG" > "$STATE/sg"

Steps 4 to 6: Wait, Set Up, Look

In step 4 the AWS tool waits until the machine is running. Then the script reads its public address and tries SSH every five seconds until the machine answers.

A real terminal recording of step 4, with the instance id and address replaced by placeholders: aws ec2 wait instance-running; the public address read into IP; describe-instances prints running, c7g.medium, us-east-1a. Then ssh with a 5-second timeout runs true, fails twice with sleep 5 after each, and answers on the third try. The last line prints: ssh works: aarch64 Ubuntu 24.04.5 LTS.

SSH failed twice while the machine was still waking up, and worked on the third try, 32 seconds after launch:

  aws ec2 wait instance-running --instance-ids "$ID"
  IP=$(aws ec2 describe-instances --instance-ids "$ID" \
    --query "Reservations[0].Instances[0].PublicIpAddress" --output text)
  aws ec2 describe-instances --instance-ids "$ID" \
    --query "Reservations[0].Instances[0].[State.Name,InstanceType,Placement.AvailabilityZone]" --output text
  until ssh -i ~/lab-data/serving/keys/ai-research-course-chk.pem "${SSHO[@]}" -o ConnectTimeout=5 \
    ubuntu@"$IP" true 2>/dev/null; do sleep 5; done
  ssh -i ~/lab-data/serving/keys/ai-research-course-chk.pem "${SSHO[@]}" ubuntu@"$IP" \
    'echo "ssh works:" $(uname -m) $(lsb_release -ds)'
  date -u +%Y-%m-%dT%H:%M:%S+00:00 > "$STATE/ssh_ready"
  echo "$IP" > "$STATE/ip"
  aws ec2 describe-instances --instance-ids "$ID" --query "Reservations[0].Instances[0].LaunchTime" \
    --output text > "$STATE/launched"

Step 5 installs Python 3.13.15 and the library versions of lessons 2 to 11, through uv, a fast installer for Python. A venv is a private toolbox folder for one project's Python, with the exact tool versions written down. Then scp sends lesson 2's model and data, lesson 2's service, lesson 9's store, lesson 9's two stand-in files and this lesson's files. Last, the machine builds its small database and works out its own reference score for every request row. That is a score worked out once, one row at a time, to check every later answer against.

A real terminal recording of step 5, with the address replaced by a placeholder and my home folder shortened to a tilde. Over ssh: make the folders; install uv, Python 3.13.15 and a venv; copy lat_requirements.txt and install it. Then scp copies lesson 2's model.pkl, store.parquet and requests.parquet, lesson 2's service.py, build_store.py and box_info.sh, lesson 9's fb_store.py, this lesson's chk files and lesson 9's prev.parquet and fb_default.json, then the features chapter's files, the shop data and the demo. Last: the store is built with 26851 rows and sqlite 3.53.1; chk_ref.py prints rows 26851, ref_equal_lesson2_expected 26818, batch32_equal_ref 26851, batch7_equal_ref 26851, cache_rows 26127 and default_rows 724; then three sha256 fingerprints for model.pkl, store.parquet and requests.parquet.

Steps 7, 7b and 8: Start, Fix, Watch From Outside

Step 7 starts the schedule in the background with setsid nohup, a command that keeps a program running after I log out. Nobody logs in again until it ends.

A real terminal recording of step 7, with the address replaced by a placeholder: over ssh, cd chk, then setsid nohup venv/bin/python chk_schedule.py 5 with its output sent to raw/schedule.log, in the background; two seconds later the log shows CHK-PROGRESS 12:16:16 schedule start K=5.

The number 5 is the count of rounds:

  ssh -i ~/lab-data/serving/keys/ai-research-course-chk.pem "${SSHO[@]}" ubuntu@"$IP" \
    "cd chk && (setsid nohup venv/bin/python chk_schedule.py $K > raw/schedule.log 2>&1 < /dev/null &); \
     sleep 2; cat raw/schedule.log"

Step 7b was not in the plan. The two priming runs worked, and so did the first counted run, but the script around it could not write its record. After each run, the store keeps working on lookups that nobody waits for any more. When the service stops, each of those answers fails to send and leaves an error line in the store's log. That log was thousands of lines long. The script handed the whole log to Python inside the command itself. But Linux allows only so much text in one piece of a command: 32 pages of memory, about 128 KiB. The schedule then stopped.

I saw no progress for 35 minutes, so I logged in to look, while nothing was running. I found the cause and fixed the script: now the log travels in a file, and only its first and last 20 lines are kept. The restart copies the two fixed files and moves the first try's files aside. It checks the fixed script with one short run that is not counted. Then it starts the schedule again. Every counted result in this lesson comes from the restarted schedule.

A real terminal recording of step 7b, with the address replaced by a placeholder. scp copies the fixed chk_run.sh and chk_schedule.py to the box. Over ssh, the first attempt's files are moved into raw/first-attempt, and ls counts 13 of them. Then one short uncounted check run, chk_run.sh test-checked at 1,100 a second for 3 s, prints one line ending in a return code of 0 and is moved aside too. Last, the schedule starts again in the background and its log shows CHK-PROGRESS 12:55:26 schedule start K=5.

The short check run printed one line with a return code of 0, and the new schedule started at 12:55:26:

  # Added 2026-10-11 after the first schedule stopped on its first counted run (chk_run.sh passed the store's log as
  # one command-line argument, and it went over Linux's limit). Copy the two fixed files, move the first attempt's
  # files aside, check the fixed run script with one short uncounted run, then start the schedule again.
  scp -q -i ~/lab-data/serving/keys/ai-research-course-chk.pem "${SSHO[@]}" \
    "$HERE/chk_files/chk_run.sh" "$HERE/chk_files/chk_schedule.py" ubuntu@"$IP":chk/
  ssh -i ~/lab-data/serving/keys/ai-research-course-chk.pem "${SSHO[@]}" ubuntu@"$IP" \
    'cd chk && mkdir -p raw/first-attempt && mv raw/schedule.log raw/runs.log raw/plan.json raw/prime-* \
     raw/checked-1100-r1* raw/load-at-start.txt raw/first-attempt/ && ls raw/first-attempt | wc -l'
  ssh -i ~/lab-data/serving/keys/ai-research-course-chk.pem "${SSHO[@]}" ubuntu@"$IP" \
    'cd chk && bash chk_run.sh test-checked checked 1100 3 9 300 | tail -1 && mv raw/test-checked* raw/first-attempt/'
  ssh -i ~/lab-data/serving/keys/ai-research-course-chk.pem "${SSHO[@]}" ubuntu@"$IP" \
    "cd chk && (setsid nohup venv/bin/python chk_schedule.py $K > raw/schedule.log 2>&1 < /dev/null &); \
     sleep 2; cat raw/schedule.log"

Steps 11 to 13: The Demo, Bring It Home, Delete

After the schedule, step 11 ran the demo you will meet later, so you can see what it prints on one core.

A real terminal recording of step 11, with the address replaced by a placeholder: python chk_demo.py on the rented box, its output also written to chk-demo-run.txt. It prints: every time below was measured on THIS machine, 1 core, 1-minute load 0.42; model trained, test AP 0.5450, the ranking score that proves it is the chapter's model, 26,851 test rows, stand-in score 0.2342; requests at 200 a second for 10 s, a pretend store with 16 slots, 1 lookup in 100 hangs for 1 to 2 s. A table in ms: A naive, first answer 1864, first request 12.60, middle 4.27, slowest 1 in 100 1125.05, slowest 1994.6, stand-ins 0.00%; B checked, 1715, 5.38, 4.08, 24.09, 25.5, 1.88%. Last line: rows answered live by both 1,983; scores that differ between them, to the bit, 0.

This exact run is stored in results/chk-demo-run.txt:

  ssh -i ~/lab-data/serving/keys/ai-research-course-chk.pem "${SSHO[@]}" ubuntu@"$IP" \
    'cd chk/serving/examples && ../../venv/bin/python chk_demo.py --save ~/chk/raw/chk-demo.json \
     | tee ~/chk/raw/chk-demo-run.txt'

Step 12 brings the results home before anything is deleted. A whole box of results was lost once in this chapter, because it was deleted before anyone copied them.

A real terminal recording of step 12, with the address replaced by a placeholder. Over ssh the box records its info after the schedule and prints the size of its raw folder, 17M; rsync copies the box's chk/raw folder to my laptop; ls and wc count 319 files; du prints 16M for the copy.

The box's raw folder held 17 MB, and 319 files arrived:

  ssh -i ~/lab-data/serving/keys/ai-research-course-chk.pem "${SSHO[@]}" ubuntu@"$IP" \
    'cd chk && bash chk_box_info.sh after > raw/box-info-after.txt 2>&1; du -sh raw'
  rsync -a -e "ssh -i $KEY ${SSHO[*]}" ubuntu@"$IP":chk/raw/ "$STATE/raw/"
  ls "$STATE/raw/" | wc -l
  du -sh "$STATE/raw/"

Step 13 deletes the machine, the key and the door rule, and checks. A machine you forget keeps costing money. The script deletes the machine, waits until AWS says "terminated", then deletes the key pair and the security group, and asks AWS again.

A real terminal recording of step 13, with the instance and group ids replaced by placeholders. terminate-instances prints shutting-down; wait instance-terminated; delete-key-pair prints true; delete-security-group prints true. Then the checks: describe-instances prints terminated; the count of key pairs named ai-research-course-chk is 0; the count of security groups of that name is 0; the count of volumes tagged Lesson=chk is 0; the count of course instances not terminated is 0. Last, the key file is removed.

The Lab's Report, Running

The report script, chk_report.py, ran on my laptop, but it measures nothing. It reads the raw files the box wrote and does all the arithmetic again with its own code.

A terminal recording of chk_report.py in eleven numbered sections. 1: c7g.medium, 1.4142 hours, $0.0609, terminated, 0 left; 100 counted runs, 100 quiet, idle at least 0.9901, steal at most 0.0000, ssh sessions at most 0. 2: every run rebuilt; naive passes 0 of 5 at every rate; checked 5 of 5 at 150 and 300, 4 of 5 at 450, 1 of 5 at 600, 0 above. 3: max rate by round, naive 0 in all five; checked 450, 600, 300, 450, 450, with dollars per million 0.0224, 0.0168, 0.0336, 0.0224, 0.0224. 4: at every rate, the slowest 1 in 100, the slowest answer and the stand-in share are clearly different. 5: 0 live answers unequal to the reference in either service; 228,096 rows live in both, 0 different. 6: the twelve guesses graded. 7: the demo's box run, and 8 of 8 quotes found in their sources. 8: leak scan, 0 of each kind. 9: one pair of runs at 450 a second, naive 105 of 6,729 requests over 25 ms, checked 5. 10: the store's ceiling, a lookup that did not hang 3.95 ms on average, mean hold 18.9 ms, at most about 846 lookups a second. 11: the playground's real numbers. Last line: all 699 checks agree with the stored lab.

It does not import the lab. For every run it works out the slowest 1 in 100 again from the stored waiting time of every request. It redoes each pass or fail, each max rate, each cost per million, each "clearly different" verdict and each guess. Then it compares every one with the lab's results file. If any number disagreed, it would stop with an error. I tried that: I changed one stored wait by a thousandth of a millisecond, the report printed "1 MISMATCHES" and stopped, and I put the file back.

Section 8 scans every stored file, even the text inside the compressed .npz data files and the terminal recordings. It looks for my account number, IP addresses, the machine's network name, AWS ids, key fingerprints (the short codes that identify keys), request ids and my home folder. It found 0 of each. Section 10 works out the store's ceiling, which a later slide uses.

The Slowest 1 in 100, Rate by Rate

This chart puts every rate side by side. Each step up the side is ten times the last: 1, 10, 100, 1,000 ms.

A chart with requests a second along the bottom, 0 to 1,200, and the slowest 1 in 100 up the side on a scale where each step is ten times the last, 1 to 10,000 ms, with a dashed line at 25 ms and a dotted upright line at 800 marked checked over 25 ms from 800/s. The checked service's line runs almost flat just under the 25 ms line, crossing it at 800, with its rounds drawn as small rings beside it. The naive service's line starts near 33 ms, jumps to about 870 ms at 450 a second, dips to about 270 at 600, then climbs to nearly 8,000 ms at 1,200; its rounds scatter between about 27 ms and 1,200 ms at the low rates. Handwritten notes: naive, 1 lookup in 100 hangs, and it waits; at 450/s the naive rounds split, 34 to 1,099 ms.

The checked service's line lies almost flat near 25 ms. In the middle rounds it went from 23.8 ms at 150 requests a second to 24.9 ms at 700. Most of that flatness is built in. The service gives up on a lookup when its 25 ms, minus 2 ms for scoring, runs out. So its slowest answers had to sit just under 25 ms.

From 800 a second, though, it crept just over the line: 25.05 ms at 800 and 25.7 ms at 1,200. That part was not fixed by the design. The limit starts when the service starts on a request, and at high rates requests waited a little before the service got to them. Python's own manual also warns that a timed wait can run a little past its limit.

The naive service's line jumps around, and then climbs. At 150 and 300 a second its slowest 1 in 100 was mostly between 27 and 47 ms, but in some rounds it passed 1 second. The reason is the helper. About 1 lookup in 100 hangs, and about 0.6 more in 100 are only slow, so about 1.6 answers in 100 are slow by design. The slowest 1 in 100 sits right on the edge of that group, and in some runs it landed inside the hangs.

At 450 a second the middle round's 869 ms is not a real wait. In that round the slowest 1 in 100 fell between a 67 ms request and a 1,011 ms one, and the arithmetic drew a straight line between them. Two rounds at 450 gave 34 and 43 ms, and three gave 869 to 1,099 ms.

At 600 a second the middle dips to 271 ms, and this time the wait is real. The store's line had begun to form: in 4 of 5 rounds, between 10.6% and 16.9% of naive lookups took over 23 ms. Above that the naive line climbs because the one core filled up. It was about 90% busy, with the service alone using about 73%. At 1,200 a second the slowest 1 in 100 took 7,833 ms, nearly 8 seconds.

Every paired comparison here was clearly different. But below 600 a second that is mostly built in: the checked service stops waiting at 23 ms, and the naive one cannot stop.

Fifteen Seconds, Request by Request

The chart above shows one number per run. This one shows every request of one pair of runs, at 450 requests a second, in round 1. Both services met identical planned arrivals and identical slow lookups.

Two strip charts, one above the other, of the same 15 seconds at 450 requests a second, round 1, with seconds along the bottom and each request's wait up the side on a scale where each step is ten times the last, with a dashed line at 25 ms. Top, naive: a band of dots between about 1 and 25 ms, a scatter of dots between 25 and about 60 ms, and a separate row of dots between 1,000 and 2,000 ms spread across the whole run; under it, naive: 105 of 6,729 requests waited more than 25 ms. Bottom, checked: the same low band of dots, everything stopping at the 25 ms line, and nothing above it; under it, checked: 5 of 6,729 requests waited more than 25 ms.

Each dot is one request, placed at the moment it was planned, at the height of its wait. I kept every request that waited more than 25 ms and 1 in 10 of the rest, so the picture is not a solid smear.

The top picture has a ceiling of dots between 1,000 and 2,000 ms. Those are the hung lookups, spread across the whole run. Below them sits a scatter between 25 and about 60 ms; the slowest of those took 58.8 ms. Most of these are lookups that were slow but did not hang: the store takes longer than usual now and then. In that run 110 naive lookups took over 23 ms by themselves, 70 of them hangs. The naive service had 105 of its 6,729 requests over 25 ms.

The bottom picture has no ceiling. The checked service never waited for a hung lookup; it gave up at its limit and answered last month's score. Only 5 of its 6,729 requests went over 25 ms, all of them just over the line.

Where the Checked Service Ran Out

If the checked service answers quickly, why did it stop passing above 450 a second? Because more and more of its answers were not live.

A bar chart for the checked service, requests a second along the bottom and stand-in answers in percent up the side, 0 to 100, with a dashed line at 5%. The bars are low and flat, under 2%, at 150, 300 and 450, then rise in steps: about 11% at 600, 23% at 700, 50% at 800, 81% at 900 and about 87% to 90% from 1,000 to 1,200. Dots show each round's value around each bar.

At 150 to 450 requests a second, about 1.5% to 1.7% of answers were stand-ins (middle rounds). That matches the naive service's own count of lookups over 23 ms, almost run for run. At 150 a second it was 33 stand-ins against 32 slow lookups, 29 against 27, 42 against 41, and so on. Below the store's limit, then, the stand-in share is simply the share of lookups the store makes slow. Lesson 9 measured 1.62% at 300 a second, and here it was 1.71%.

At 600 a second the stand-ins jumped to 11.1%, at 800 to 49.8%, and at 1,200 to 90.2%. One round at 450 a second also tipped: round 3 reached 5.84%, just over my 5% line, which is why that round's max rate was 300.

The store has a ceiling you can work out. A lookup that did not hang took 3.95 ms on average, measured on the naive service's clock at 150 a second. With 1 lookup in 100 holding its slot for about 1.5 seconds, the average lookup holds a slot about 18.9 ms. Sixteen slots divided by 18.9 ms is about 850 lookups a second. With random bunching, the slots start to fill well before that, and here the stand-ins took off between 600 and 800 a second.

Lesson 9 tested one more choice for exactly this case: passing the deadline down to the store, so the store drops work nobody waits for. That arm was not in this lesson's list, so I did not test it here. It is the first thing I would try next.

How High the Checked Service Climbed

Put the pass or fail results together and you get each service's max rate: the highest rate that passed, with every lower rate passing too.

A round cockpit dial marked from 0 to 1,200 requests a second, with a band along its edge from about 850 to 1,200. Thick needles point at 300, at 450, which is thickest because three rounds stopped there, and at 600, for the checked service's max rate in each round; a dashed needle rests at 0 for the naive service. Below: thick needles, the checked service's max rate in each round, 450, 600, 300, 450, 450; dashed needle at 0, the naive service, no rate passed in any round; the band, the store's ceiling, worked out at about 850 lookups a second, 16 slots each held about 19 ms on average.

The checked service reached 450, 600, 300, 450 and 450 requests a second in rounds 1 to 5. The naive service reached none: no rate passed in any round, not even 150 a second. By my rule that counts as clearly different, but it is built in: with this store the naive service could not pass at any rate. The finding is the checked service's own number, 300 to 600 a second.

That number belongs to this pretend store, not to the service. The core was still 28% to 36% idle at every rate from 600 up, so the processor was not the limit. The store was. Lesson 9 chose the store's 16 slots so that it would be near its limit at 450 a second. And 450 is where the checked service stopped in the middle round. With a different helper, the number would be different.

Here is a rough comparison across two lessons, not a paired one. Lesson 11 measured 914 a second on this kind of machine, with no slow store, no live condition and no batcher. The slow store roughly halved what one core could carry.

The First Answer Waits for Python

Every run began cold. How long did each service take to give its first answer?

Two runways seen from above, one per service, drawn to one scale, each a long strip with a dashed centre line, a door line near the far end and a small plane at the end. Naive: taxiing, Python loads its tools, 1,689 ms; door opens, lines 20 ms, first answer at 1,731 ms. Checked: taxiing, 1,702 ms; door opens, lines 20 ms, first answer at 1,735 ms. Below: the checked service's warm-up took about 6 ms, a sliver just before its door, thinner than the door line at this scale.

Both took about 1.73 seconds. Across all 50 runs of each service, the middle first answer was 1,735 ms for the checked service and 1,731 ms for the naive one. Starting up, until the port opened, was about 1,700 ms of that. At every rate the five rounds went both ways, so by my rule the two were never clearly different.

That matches lesson 6: almost all of a cold start is Python loading its tools. The checked service's warm-up ran before its port opened and took about 6 ms. So the warm-up did not make the first answer earlier. It moved a few milliseconds of work to before the door opened, so that the first real request would not pay for them.

Did the first real request gain? In the middle rounds it waited less with the checked service, at every rate. It was 10.9 ms against 22.6 ms at 300 a second, and 6.2 against 11.7 ms at 150. By my rule, though, the gap was clearly different at only one rate of ten, 1,100 a second. A single request is noisy: on one core it can land behind the tester or the store. So I can say the first request usually waited less, and not more than that.

With the naive service, the slowest of the first 100 requests took 47 to 182 ms in the middle rounds. That is slow lookups and the cold start, with no limit. With the checked service it was 24 to 27 ms, which its 23 ms budget makes nearly certain. A budget here is the 25 ms the service allows itself for one request.

What the Checks Cost the Core

Did the checks make the service work harder? This picture shows who had the one core at 450 requests a second.

Two stacks of flat rings drawn in 3D, one per service, each stack one core's time at 450 requests a second, each ring as thick as one share, with a key. Naive service: service 45%, store 7%, tester 5%, everything else and idle 43%. Checked service: service 43%, store 8%, tester 5%, everything else and idle 44%. Below: each stack's four rings add up to the whole core; the two stacks are nearly alike.

Each stack is one core, and each ring is as thick as one program's share of its time. The stacks are nearly alike. The naive service used 45% of the core and the checked service 43%. The store used 7% and 8%, and the tester 5% in both. The rest, about 44%, was the system and idle time.

At higher rates the checked service used less of the core per request than the naive one. More of its answers were stand-ins, which need no model call, and its batches grew. Those two effects are mixed together, so I give no number for batching alone.

Did the Batcher Do Anything?

The checked service scores every waiting row together, up to 32 in one model call. How many rows rode together?

A hand-drawn bar chart for the checked service, average rows in one model call for live answers, middle of 5 rounds, by requests a second: 1.2 at 150, 1.4 at 300, 1.6 at 450, 1.9 at 600, 2.1 at 700, 2.3 at 800, 2.6 at 900, 3.1 at 1,000, 3.3 at 1,100 and 4.0 at 1,200.

Not many. At 150 requests a second a call carried 1.22 rows on average (live answers, middle rounds). At 450 it carried 1.63, and at 1,200 it carried 4.0. A batch forms only when several requests are ready at the same moment, and at low rates that is rare. Lesson 11 measured 12.7 rows a call, but at 2,737 requests a second, a rate this lab never reached.

So at the rates where the checked service passed, the batcher was mostly scoring one row at a time. Lesson 4 warned that a batcher which never groups anything is pure cost. This lab changed three settings at once, so it cannot say what the batcher cost on its own here.

What the Checks Cost in Memory

The last cost to check is memory. At the end of each run the tester read how much memory the service was holding.

A drawn balance scale with its beam level and two pans hanging straight down from the ends. A tag on top of the post reads too close to call. Under the left pan: naive, 273 MB. Under the right pan: checked, 266 MB. Below: 1 MB is a million bytes; one byte holds one letter; the beam stays level because by the 5-round rule the gap is too close to call.

At 450 requests a second, the naive service held 273 MB and the checked service 266 MB (middle rounds). By my rule those are too close to call, so the beam is drawn level. Both services load one Python, one set of libraries and one model, and the batcher's queue and the stand-in table are small. From 800 a second up the checked service held less in all 5 rounds, by 7 to 13 MB in the middle rounds. My guess, which I did not test, is that the naive service had thousands of requests waiting inside it at those rates.

Two Services, One Set of Answers

A faster service is worthless if it gives different answers, so I checked every live score.

A drawn zipper whose two rows of teeth meet along their whole length. The upper row is labelled naive service's live scores and the lower row checked service's live scores. Under it in large type: 228,096 rows live in both, and 0 different, to the last digit. Below: every live answer of the checked service had a twin in the naive run; built in, one model, one row, one number.

Every live answer in every run was compared with the machine's own reference score, to the last digit. That was 540,003 answers from the naive service and 228,096 from the checked one. Not one differed. The naive service never gives a stand-in, so every row the checked service answered live was answered live by both: 228,096 rows, 0 different.

That much is built in. One model scoring six numbers gives one score, whether it scores one row or a batch of 32, which chk_ref.py also checked on all 26,851 rows. The check still matters, because a bug in a batcher, such as a row matched to the wrong answer, would show up here.

The Price of a Million Answers

Lesson 11 turned a max rate into a cost: the price of one hour, divided by the answers in that hour.

Two torn ticket stubs side by side. The checked service stub: fare $0.0363 an hour; max rate 450 a second; $0.0224 per million answers, middle round; a row of five punched holes sized by each round's cost, labelled $0.0224, $0.0168, $0.0336, $0.0224, $0.0224. The naive service stub: no rate met the target, a tilted stamp reading NO BOARDING, and the line built into the store: waiting for every lookup cannot pass.

The c7g.medium costs $0.0363 an hour. At the checked service's max rate, a million answers cost $0.0224 in the middle round, and between $0.0168 and $0.0336 across the rounds. The naive service has no price at the target, because no rate met it, and that, again, is built into the store.

At one fixed rate, the two services cost exactly alike per answer: one machine, one price, one number of answers. The real difference is how much traffic one machine can carry while keeping the promise.

The Chapter in Twelve Stops

Before the checklist, here is the whole chapter in one picture. Each station is one lesson, with the one number it measured. Every number on it is read from that lesson's own results file.

A metro map: one line that snakes across the page in three rows, with twelve numbered stations. 1 batch or online: 99.1% of nightly scores never read. 2 one request: scoring, 54% of one request. 3 tail latency: slowest, the 1st request, 7.7 x the middle. 4 batching: slowest 1 in 100, 5.26 to 2.50 ms. 5 workers: 1 worker, 1 thread, 1,071 a second. 6 cold start: first answer after 1,499 ms. 7 load testing: closed test 0.97 ms, open 85.0 ms. 8 caching: 43% hits in 30 days, all stale. 9 timeouts: slowest 1 in 100, 1,030 to 25.7 ms. 10 many models: 4 workers, 1,080 to 385 MB. 11 cost: $0.0110 per million answers. 12 this lesson: every choice, in one service, a filled station at the end of the line.

And here is the journey in words, one finding per lesson, with a link back to each.

  1. Batch or online. Scoring every customer once a night could not be told apart from scoring each visit live, in ranking. But 99.1% of the nightly scores were never read.

  2. Latency anatomy. A typical request took 1,075.9 microseconds (millionths of a second), and the model's own scoring was the biggest step: 54% of it.

  3. Tail latency. On a quiet machine the slowest 1 in 100 was only 1.08 times the middle request. The single slowest was 7.7 times the middle, and it was always the first request of the run.

  4. Batching requests. Scoring up to 32 waiting requests in one call, with no waiting on purpose, helped. At about 550 requests a second, the slowest 1 in 100 fell from 5.26 ms to 2.50 ms.

The Checklist: Take This Page to Work

This is the page I would pin above my desk. Each line is one decision, the lesson that measured it, that lesson's number, and when the line does not apply. Skipping a line that does not fit your service is part of using the list well.

A drawn cockpit panel with screws in its corners. The top two rows are toggle switches, each with its lesson above it and its decision below it. Switched up, with a lit lamp: lesson 5, one worker per core; lessons 3 and 5, one thread in the model; lesson 6, warm up before opening; lesson 9, one time limit, then a stand-in; lesson 4, batch the rows that wait. Switched down under a guard cover: lesson 10, load before forking workers, off here; lesson 8, cache the answers, off here. Below a dashed line, five round lamps: lesson 1, count scores never read; lesson 2, time every step; lesson 3, watch the slowest 1 in 100; lesson 7, test with planned arrivals; lesson 11, price each answer. Below the panel: up means on in the checked service, and the first two were true of the naive one too, on this box.

DecisionLessonWhat it measuredWhen this does NOT apply
Decide online or nightly by counting how many scores are read1a month-old score could not be told apart from a live one in ranking, and 99.1% of nightly scores were never readwhen the answer changes within hours, or a new customer needs a score before tonight
Time every step of one request before changing anything2

Does This Item Apply to Me?

A checklist is only useful if you know which lines to skip. This little decision tree asks the three questions that decided it here.

A flowchart with three questions in six-sided boxes, top to bottom, and each answer drawn in its own shape. More than one worker? Yes leads to a double-sided box, load before copying, then gc.freeze (lesson 10), which joins the no path to the next question. Can a lookup stall? Yes leads to a rounded box, one time limit, then a stand-in (lesson 9), which joins the no path. Are requests often ready together? Yes leads to a drum, batch up to 32, never wait (lesson 4); no leads to a circle, one row at a time.

Start with how many workers you run. With more than one, lesson 10's recipe saves memory: load the model once, then make the copies, which share it, then call gc.freeze. That call tells Python's memory cleaner to leave the shared memory alone, so it stays shared. With one worker, skip it.

Next, ask whether a lookup can stall. Your service may wait for something outside itself, such as a database, a or another service. If so, give the whole request one time limit and a stand-in answer, and watch the stand-in share. Lesson 9 and this lesson both found that without a limit, the slowest answers are as slow as the slowest helper.

Last, ask whether requests are often ready together. If several requests often arrive at the same moment, batch them. If not, a batcher does little, as this lesson's rows-per-call numbers showed.

A few lines apply almost everywhere. Warm up before the door opens. Test with planned arrivals, timed from the plan, as lesson 7 did. And on a computer with more cores than workers, set one thread per worker, because the library may otherwise pick many.

My Twelve Guesses Before the Run, Checked

I wrote twelve guesses into the lab before it ran. Each is quoted word for word from the docstring. Seven were right, two were partly right and three were wrong.

  1. "G1. Naive at 300 a second: slowest 1 in 100 above 500 ms in at least 4 of 5 rounds (lesson 9's no-limit arm: 1,030 ms, but 29.6 ms in 1 round of 5). Max above 1,000 ms in every round." Partly right. The max was above 1,000 ms in every round, 1,961 to 2,003 ms. But the slowest 1 in 100 was above 500 ms in only 2 of 5 rounds; in the other 3 it was 29.9 to 39.9 ms.

  2. "G2. Checked at 150, 300 and 450 a second: slowest 1 in 100 at most 25 ms in every round (close to built in)." Right: 23.5 to 24.5 ms. As the guess says, the limit builds most of this in.

  3. "G3. Checked: max rate between 600 and 900 a second in every round, set by the store's 16 slots filling up (stand-ins above 5%), not by the core." Wrong on the numbers: 450, 600, 300, 450 and 450. The reason was right: every run that set a max rate failed on the live condition alone, with its slowest 1 in 100 still under 25 ms. From 800 a second the runs also failed on time.

  4. "G4. Naive: no rate passes in at least 4 of 5 rounds (max rate "none"); naive at 150 a second may pass in a round where the hangs land just above the 1-in-100 mark." Right. No rate passed in any round. The second half could not happen: with this store, about 1.6 lookups in 100 are slow, not 1.

  5. "G5. Launch to first answer between 1,200 and 2,500 ms in both arms; checked minus naive within plus or minus 100 ms and too close to call (the warm-up moves work before the port opens; it does not remove it)." Right: middle values 1,716 to 1,752 ms, never clearly different.

  6. "G6. The first request's own wait: checked below naive in all 5 rounds at every rate, clearly different." Wrong. The middle values were lower for the checked service at every rate, but the gap was clear at only 1 rate of 10.

  7. "G7. Checked at 300 a second: stand-in share between 1.0% and 2.5% in every round (lesson 9: 1.62%)." Right: 1.26% to 2.09%.

Try It Yourself

The full lab needs a rented machine for about an hour and a half. The demo, chk_demo.py, runs on your own computer. It starts the naive service and then the checked one, each from cold. It sends both one stream of random traffic, with lesson 9's pretend slow store inside, and prints a small table.

I wrote the demo's design into its docstring after the lab's design and before the demo first ran.

A real screenshot of VS Code with chk_demo.py open at the top of the file, showing its docstring: what it needs, how to run it, and the design written before it first ran, with two services, A naive and B checked.

Before you run this lab. The demo uses lesson 2's Python environment, with the library versions of the lab's machine. If you made it for lessons 2 to 11, use it again. If not, make it once, inside the scripts/labs/serving/examples folder:

python3 -m venv venv-sv
source venv-sv/bin/activate        # on Windows: venv-sv\Scripts\activate
pip install -r lat_requirements.txt

I used Python 3.13. Then run python ../../features/fetch_data.py once, which downloads the shop data. Now run python chk_demo.py in the examples folder, with venv-sv active. It needs no graphics card, no cloud account and no administrator rights. If a package is missing, it prints one line saying why and stops. The service it starts listens on 127.0.0.1, an address only your own computer can reach.

The first line it prints says the numbers come from your machine, with its number of cores and its load at that moment. The demo also prints the model's AP, a score from 0 to 1 for how well the model ranks customers. It must be 0.5450, which proves it is the chapter's model. On a laptop with many cores, the naive service also uses many threads inside the model, which lesson 5 found slower. Read the two rows against each other, on your machine.

r"""A naive service and a checked one, side by side, with one slow lookup, on YOUR machine.

Lesson 12 of 'Serving and Inference Basics'. It uses lesson 2's Python environment, made once inside this folder
(Python 3.13 is what the lab used):
    python3 -m venv venv-sv
    source venv-sv/bin/activate              # Windows: venv-sv\Scripts\activate
    pip install -r lat_requirements.txt      # numpy, pandas, pyarrow, scikit-learn, fastapi, uvicorn, ...
It also needs the shop data from the features chapter: run python ../../features/fetch_data.py once first. Then:
    python chk_demo.py                  # train, then run the two services one after another, print the table
    python chk_demo.py --save out.json  # and save the numbers
If a package is missing, it prints one line saying why and stops.
The timings it prints come from YOUR machine, as it is right now: other programs move them. Read the two rows
against each other, on your machine; they are not numbers to set beside the lab's rented machine.
It starts one helper process at a time (the service, on 127.0.0.1 only, so only your own computer can reach it) and
stops it. It writes nothing except a temporary model file, deleted at the end, and out.json if you ask.

Design, written 2026-10-11 after the lab's design (serving_checklist.py) and before this file first ran:
  1. Train the features chapter's model (it must score test AP 0.5450), as lesson 9's fb_demo.py does.
  2. A small FastAPI service in a second process (this same file, run with --serve). Inside it, lesson 9's PRETEND
     feature store: 16 slots; each lookup waits about 2 ms most of the time, with a long tail, and 1 lookup in 100
     hangs for 1 to 2 s. The waits are drawn from a fixed seed, so both services meet the same waits. (The lab's
     store is a separate process; here it lives inside the service, to keep the file short.)
  3. Two services, each started fresh, so each one starts cold:
       A naive    no warm-up; no time limit on the lookup; one row per model call; the thread count left to the
                  library (on a laptop with many cores, that is many threads)
       B checked  one thread (OMP_NUM_THREADS=1); a warm-up before the port opens; one 25 ms limit for the whole
                  request, 2 ms of it kept for scoring; a stand-in score when time runs out (here the share of buyers
                  in the training rows: the demo has no "last month's score"); up to 32 waiting rows in one model call
  4. Requests arrive at random (Poisson) at 200 a second for 10 s, from the moment the port opens; latency from the
     SCHEDULED arrival (lesson 7). Request j asks for test row j. Both services get the same arrival times.
  5. Print for each: launch to first answer, the first request's wait, the middle wait, the slowest 1 in 100, the
     slowest of all, and the stand-in share. Then: rows answered live by both, and how many of them differ.

Author: Roni Das
Created: 2026-10-11
"""
import importlib.util
import json
import os
import socket
import subprocess
import sys
import tempfile
import threading
import time
from pathlib import Path
from queue import Queue
from time import perf_counter

NEEDED = ("numpy", "pandas", "pyarrow", "sklearn", "fastapi", "uvicorn")
missing = [m for m in NEEDED if importlib.util.find_spec(m) is None]
if missing:
    sys.exit(f"chk_demo.py needs {', '.join(missing)}: make the environment in its docstring "
             f"(pip install -r lat_requirements.txt), then run it with that environment's python.")

import numpy as np  # noqa: E402

RATE, SECONDS = 200, 10.0
WAYS = ("A naive", "B checked")


# ───────────────────────────── the service (python chk_demo.py --serve PORT FILE WAY) ─────────────────────────────
def serve(port: int, data_path: str, way: str) -> None:
    import asyncio
    import pickle
    from contextlib import asynccontextmanager

    import uvicorn
    from fastapi import FastAPI, Request
    from starlette.responses import Response

    checked = way == "B checked"
    d = pickle.loads(Path(data_path).read_bytes())
    model, X, stand_in = d["model"], d["X"], d["stand_in"]
    rng = np.random.default_rng(0)                    # the same waits for both services
    wait = 0.002 * np.exp(rng.standard_normal(len(X)))
    hang = rng.random(len(X)) < 0.01
    wait = np.where(hang, rng.uniform(1.0, 2.0, len(X)), wait)
    slots: list = []
    queue: list = []
    have: list = []

    async def store(row: int) -> list[float]:
        """The pretend feature store: wait for one of 16 slots, hold it for the drawn time, answer."""
        if not slots:
            slots.append(asyncio.Semaphore(16))
        async with slots[0]:
            if row >= 0:
                await asyncio.sleep(float(wait[row]))
        return X[max(row, 0)].tolist()

    async def batcher() -> None:
        """Up to 32 waiting rows in one model call; never waits on purpose for more."""
        while True:
            if not queue:
                have[0].clear()
                await have[0].wait()
            await asyncio.sleep(0)
            take = queue[:32]
            del queue[:32]
            p = model.predict_proba(np.array([f for f, _ in take], dtype=np.float64))[:, 1]
            for (_, fut), s in zip(take, p):
                fut.set_result(float(s))

    @asynccontextmanager
    async def lifespan(_app):
        if checked:                                   # the warm-up, before the port opens
            have.append(asyncio.Event())
            asyncio.get_running_loop().create_task(batcher())
            f = await store(-1)
            model.predict_proba(np.array([f], dtype=np.float64))
            model.predict_proba(np.array([f] * 32, dtype=np.float64))
        yield

    app = FastAPI(lifespan=lifespan)

    @app.post("/score")
    async def score(request: Request) -> Response:
        t_in = perf_counter()
        q = json.loads(await request.body())
        if not checked:
            feats = await store(q["row"])                                   # no time limit
            p = float(model.predict_proba(np.array([feats], dtype=np.float64))[0, 1])
            out = {"id": q["id"], "score": p, "stand_in": False}
        else:
            left = t_in + 0.025 - 0.002 - perf_counter()                    # one budget for the whole request
            try:
                feats = await asyncio.wait_for(store(q["row"]), max(left, 0.0))
            except TimeoutError:
                feats = None
            if feats is None:
                out = {"id": q["id"], "score": stand_in, "stand_in": True}
            else:
                fut = asyncio.get_running_loop().create_future()
                queue.append((feats, fut))
                have[0].set()
                out = {"id": q["id"], "score": await fut, "stand_in": False}
        return Response(content=json.dumps(out).encode(), media_type="application/json")

    uvicorn.run(app, host="127.0.0.1", port=port, log_level="warning")


# ───────────────────────────── the model ─────────────────────────────
def build():
    """The features chapter's model and its test rows (lesson 9's fb_demo.py build step)."""
    from sklearn.metrics import average_precision_score
    here = Path(__file__).resolve().parent
    sys.path.insert(0, str(here.parents[1] / "features"))
    import task
    from what_a_feature_is import HAND_COLS, hgb, joined

    ev = task.load_events()
    lab_tr, _, lab_te = task.splits(ev)
    tr = joined(ev, lab_tr, task.TRAIN_CUTOFFS)
    te = joined(ev, lab_te, task.TEST_CUTOFFS)
    model = hgb(0).fit(tr[HAND_COLS].to_numpy(float), tr["label"].to_numpy())
    X = np.ascontiguousarray(te[HAND_COLS].to_numpy(np.float64))
    p = model.predict_proba(X)[:, 1]
    y, cut = te["label"].to_numpy(), te["cutoff"].to_numpy()
    ap = float(np.mean([average_precision_score(y[cut == c], p[cut == c]) for c in np.unique(cut)]))
    assert round(ap, 4) == 0.5450, ap
    return model, X, ap, float(tr["label"].mean())


# ───────────────────────────── launch, then the open-loop tester (lesson 7's) ─────────────────────────────
def connect(port: int):
    import http.client
    c = http.client.HTTPConnection("127.0.0.1", port, timeout=30)
    c.connect()
    c.sock.setsockopt(socket.IPPROTO_TCP, socket.TCP_NODELAY, 1)
    return c


def run_way(way: str, data_path: str, n_rows: int) -> dict:
    s = socket.socket()
    s.bind(("127.0.0.1", 0))
    port = s.getsockname()[1]
    s.close()
    env = {k: v for k, v in os.environ.items() if k != "OMP_NUM_THREADS"}
    if way == "B checked":
        env["OMP_NUM_THREADS"] = "1"
    rng = np.random.default_rng(1)                    # the same arrivals for both services
    at = np.cumsum(rng.exponential(1 / RATE, size=int(RATE * SECONDS * 1.5) + 50))
    at = at[at < SECONDS]
    lats = np.zeros(len(at))
    answers: list = [None] * len(at)
    done_at = np.zeros(len(at))
    q: Queue = Queue()
    start = [0.0]
    t_launch = perf_counter()
    svc = subprocess.Popen([sys.executable, __file__, "--serve", str(port), data_path, way], env=env)
    try:
        while True:
            try:
                socket.create_connection(("127.0.0.1", port), timeout=1).close()
                break
            except OSError:
                time.sleep(0.001)

        def sender() -> None:
            c = connect(port)
            while True:
                j = q.get()
                if j is None:
                    break
                c.request("POST", "/score", body=json.dumps({"id": j, "row": j % n_rows}),
                          headers={"Content-Type": "application/json"})
                answers[j] = json.loads(c.getresponse().read())
                done_at[j] = perf_counter()
                lats[j] = done_at[j] - (start[0] + at[j])          # from the SCHEDULED arrival
            c.close()

        pool = [threading.Thread(target=sender, daemon=True) for _ in range(64)]
        for th in pool:
            th.start()
        start[0] = perf_counter() + 0.002
        for j, a in enumerate(at):
            d = start[0] + a - perf_counter()
            if d > 0:
                time.sleep(d)
            q.put(j)
        for _ in pool:
            q.put(None)
        for th in pool:
            th.join()
    finally:
        svc.terminate()
        svc.wait()
    lat = lats * 1000
    return {"n": len(lat), "first_answer_ms": (done_at.min() - t_launch) * 1000, "first_request_ms": lat[0],
            "p50": float(np.percentile(lat, 50)), "p99": float(np.percentile(lat, 99)), "max": float(lat.max()),
            "stand_in_share": float(np.mean([a["stand_in"] for a in answers])),
            "scores": [None if a["stand_in"] else a["score"] for a in answers]}


def main() -> None:
    try:
        load = f"{os.getloadavg()[0]:.2f}"
    except (AttributeError, OSError):
        load = "not available on this system"
    n = os.cpu_count()
    print(f"Every time below was measured on THIS machine ({n} {'core' if n == 1 else 'cores'}, 1-minute load {load}).")
    model, X, ap, stand_in = build()
    print(f"model trained: test AP {ap:.4f}; {len(X):,} test rows; stand-in score {stand_in:.4f}")
    print(f"requests arrive at random, {RATE} a second for {SECONDS:.0f} s, from the moment the port opens;")
    print("the pretend store has 16 slots and 1 lookup in 100 hangs for 1 to 2 s\n")
    import pickle
    out = {"load": load, "cores": n, "test_ap": ap, "ways": {}}
    print("all times in ms   first answer  first request   middle  slowest 1 in 100  slowest   stand-ins")
    res = {}
    with tempfile.TemporaryDirectory() as tmp:
        dp = Path(tmp) / "data.pkl"
        dp.write_bytes(pickle.dumps({"model": model, "X": X, "stand_in": stand_in}))
        for way in WAYS:
            r = run_way(way, str(dp), len(X))
            res[way] = r
            out["ways"][way] = {k: v for k, v in r.items() if k != "scores"}
            print(f"   {way:<12}{r['first_answer_ms']:13.0f}{r['first_request_ms']:15.2f}{r['p50']:9.2f}"
                  f"{r['p99']:18.2f}{r['max']:9.1f}{r['stand_in_share']:11.2%}")
    a, b = res["A naive"]["scores"], res["B checked"]["scores"]
    both = [(x, y) for x, y in zip(a, b) if x is not None and y is not None]
    diff = sum(x != y for x, y in both)
    print(f"\nrows answered live by both: {len(both):,}; scores that differ between them, to the bit: {diff}")
    out["live_both"], out["live_differ"] = len(both), diff
    if "--save" in sys.argv:
        Path(sys.argv[sys.argv.index("--save") + 1]).write_text(json.dumps(out, indent=1))


if __name__ == "__main__":
    if len(sys.argv) > 1 and sys.argv[1] == "--serve":
        serve(int(sys.argv[2]), sys.argv[3], " ".join(sys.argv[4:]))
    else:
        main()

Flip the Switches Yourself

This box runs in your browser and needs nothing but Python. It is a make-believe service with a pretend clock: a pretend store with 16 places, 1 lookup in 100 hanging, and one core. The time each step takes comes from lessons 2, 3 and 4, with its source written next to it. Everything else is made up.

Press Run. Then switch the items off one at a time. With TIME_LIMIT = False the slowest answers jump to about a second. With WARM_UP = False only the first request's wait changes. With BATCHING = False only the rows per call change, because at 300 requests a second few requests are ready together. Then set RATE = 1000 with the time limit on, and watch the stand-ins take over, as they did on the real box.

Every random draw has its own list of dice rolls, written down first. So switching an item on or off meets exactly the arrivals and the lookups it met before. A seed is the number that picks which lists of dice rolls you get, and one seed is luck, so it runs five.

The report script writes the box's real numbers at 300 requests a second into the last line, so you can compare. With all three items on, the pretend service's slowest 1 in 100 is exactly 23.5 ms in every seed. That is 23 ms of limit plus one request's work, so it is built in. About 1.7% to 2.0% of its answers are stand-ins; the box measured 23.75 ms and 1.71%.

With TIME_LIMIT = False the pretend service prints slowest-1-in-100 times above a second in all five seeds, while the box's middle round at 300 a second was only 39.9 ms. The box simply got fewer slow lookups than average in 3 of its 5 rounds, so its slowest 1 in 100 landed just below the hangs. The pretend service is also simpler than the real one. Its time limit starts at the request's arrival. It charges the core a fixed time for each request. And its store never slows down from its own work.

The Lab's Code, Piece by Piece

The lab has one file that runs on my laptop, small files that run on the rented machine, and the step script you saw above.

serving_checklist.py, on my laptop, holds the design and my twelve guesses in its docstring, written before any run, and a labelled note added after the runs. Its commands are check, which checks every file is there, and prepare, which copies lesson 9's two stand-in files. status reads the machine's console. cost works out the bill and checks that everything was deleted. The others are collect, which copies the raw files into results/chk-raw/ through chk_redact.py and does the arithmetic, and factcheck, which finds each outside claim in its source.

chk_files/chk_service.py is the service, both arms in one file. CHK_ARM is a setting the program reads when it starts, and it picks naive or checked. The file imports lesson 2's service unchanged, for the model and the shape of a request. The checked arm's warm-up runs in a start-up step that the web server, uvicorn, finishes before it opens the door. The batcher is lesson 4's, with no timed wait.

chk_files/chk_client.py is the tester. It starts the service itself, so the cold start and the traffic are read on one clock. Then it runs lesson 7's timetable of requests and checks every answer against the machine's own reference score, to the last digit.

When to Use the Checklist, and When Not To

Use it as a list of questions, not a list of answers. Each line says what to measure, and why. The numbers belong to one small model on one small machine. Your model, your helper and your traffic will give different numbers, and maybe a different answer.

A time limit with a stand-in belongs wherever a request waits for something outside the service. But watch the stand-in share as closely as the waiting times. A service can look fast while it quietly stops answering with the model, as the checked service did above 600 requests a second. The limit hides a full helper; it does not fix one.

Do not copy a setting that has no reason to apply. Loading the model before copying workers saves memory only when there are several workers. A cache helps only when the same question comes back and an old answer is still a good answer. A batcher helps only when requests are ready together.

Do not judge a service by its first answer. Both services here took about 1.7 seconds to give their first answer, almost all of it Python starting. If a service must answer fast right after it starts, the fix is to keep a warm copy running, not to warm it faster.

And do not trust a load test that you did not run with planned arrivals. Every number in this lesson was timed from the planned moment. A polite tester would have hidden the hung lookups, as lesson 7 showed.

What This Lab Cannot Tell You

A hand-drawn chart with two axes: along the bottom, more requests a second, on more cores; up the side, the helper stalls more. A curved line falls from the top left to the bottom right, and outside it a note says nobody has flown. Inside the line, near the bottom left, a solid oval reads tested here, 1 core, 1 stall in 100. Two hatched blobs sit beside it: to the right, more cores and workers, not tested; above, a helper that stalls more often, or for everyone at once, not tested.

Everything shared one core, the tester included. On a bigger machine with one worker per core, the numbers would change, and lesson 5's advice was never tested beyond one core.

The slow helper was lesson 9's pretend store, on the same box, with slow lookups that use no processor time and are independent of each other. A real store across a network can be slow for everyone at once. Its size was chosen by hand in lesson 9, and that size set the checked service's limit here.

The model is small and runs on the processor. A big neural network on a graphics card would spend its time very differently, and batching would matter much more. That case sits on neither axis of the picture above: it is a different kind of lab.

The traffic was steady and random, for 15 seconds a run. Real traffic rises and falls, and a cold start can land in the busy hour.

The first schedule stopped on its first counted run because of a fault in the run script, which step 7b explains. I fixed it and restarted. No counted run came from the first schedule, and nothing about the arms changed. I logged in four times to find the fault, while nothing was running, and the commands are shown above. The copy step was recorded twice, because the first recording lost its top lines.

One setting from lesson 9 was not tried. Passing the deadline down to the store, so it drops work nobody waits for, was not on this lesson's list. On this box it is the most likely way to push the checked service past 450 requests a second.

Knowledge Check

Knowledge Check

4 questions - Score 80% to pass

Q1

The naive service never met the 25 ms target, not even at 150 requests a second. What was the main reason?

Q2

Above 600 requests a second, the checked service's slowest answers stayed near 25 ms. What failed first?

Q3

The checked service warmed up before opening its door. What did that change?

Q4

This lab used one worker. Why was lesson 10's recipe, loading the model before copying workers and then gc.freeze, left out?

SettingLessonNaive arm on this boxChecked arm on this box
One worker on the one core5: 1 worker with 1 thread answered 1,071 a second, more than any other mixone worker: no differenceone worker
One thread inside the model (OMP_NUM_THREADS=1)3 and 5: the first request got faster; more threads were slowerthe library picks 1 thread on 1 core: no difference here1 thread, set on purpose
A warm-up before the door opens6: the first request took 4.7 ms cold and 1.7 ms after one warm-upnoneone practice request
One time limit for the whole request, then last month's score9: no limit, 1,030 ms; a 25 ms limit on the lookup, 25.7 ms; one 25 ms limit for the whole request, 23.6 ms, lower partly by designnone: waits as long as it takes25 ms, 2 ms of it kept for scoring
Every waiting row in one model call, up to 324: at about 550 requests a second, half of what the core could do, the slowest 1 in 100 fell from 5.26 ms to 2.50 msone row per callup to 32 rows, never waiting for more
Load the model before copying workers, then gc.freeze10: 4 workers used 385 MB instead of 1,080 MBnot usednot used: with one worker there is nothing to share
Keep a cache of answers8: every real cache hit was stalenot usednot used

On this box the two services differed in three ways: the warm-up, the time limit with its stand-in, and the batcher. The thread setting still matters on a bigger computer, as the laptop run near the end shows. Two words in the table need a gloss. A library is a ready-made box of code that a program uses. And gc.freeze is a call that tells Python's memory cleaner to leave shared memory alone.

The time limit needs a word more. In lesson 9, a 25 ms limit on the lookup alone was not quite enough: the slowest 1 in 100 took 25.7 ms. One limit for the whole request, with 2 ms kept back for scoring, came in at 23.6 ms. Part of that is built in, because that limit gives the lookup at most 23 ms, so it had to come in lower.

Lesson 9's clock also started at the planned arrival. A real service cannot see that moment, so here the clock starts when the service starts working on the request. Time spent waiting before that is not covered, and the tester still counts it. The results show what that costs.

The cache was left out on purpose. A cache is a shelf of answers already given, so a repeated question gets the old answer at once; a question the shelf can answer is a hit. Lesson 8 found that every hit on the shop's real visits was stale, out of date, because the customer had bought something since. It did not find a cache that was both safe and useful. In this test every request asks about a different customer and month, so a cache could never hit anyway.

export AWS_DEFAULT_REGION=us-east-1 AWS_PAGER="" NAME=ai-research-course-chk TYPE=c7g.medium ARCH=arm64 HERE=$PWD STATE=${CHK_STATE:-$HOME/lab-data/serving/chk} KEY=$HOME/lab-data/serving/keys/$NAME.pem K=${K:-5} B=/home/ubuntu/chk LAT=$HOME/lab-data/serving/lat FEAT=$HERE/../features mkdir -p "$STATE" "$HOME/lab-data/serving/keys" SSHO=(-o LogLevel=ERROR -o StrictHostKeyChecking=accept-new -o "UserKnownHostsFile=$STATE/known_hosts") [ -f "$STATE/id" ] && ID=$(cat "$STATE/id") [ -f "$STATE/ip" ] && IP=$(cat "$STATE/ip") [ -f "$STATE/sg" ] && SG=$(cat "$STATE/sg")

The last three lines read back small state files. Earlier steps save the machine's id, its address and its door rule's id in them, so later steps can find them.

The machine was billed by the second at its on-demand price, the price with no contract, $0.0363 an hour, which lesson 11 read from AWS's price list. Its public internet address cost $0.005 an hour, and its 16 GiB disk $0.08 per GB per month. It existed for 1.4142 hours, from launch to deleted, and cost $0.0609 in all. Those numbers are in results/chk-cost.json.

Every line of the terminal recordings passed through a small filter, chk_redact.py, a copy of lesson 11's. It blacks out my account number and every IP address (a computer's number on the internet, like a phone number). It also hides every id AWS gives a machine, a disk or a door rule, and the machine's own name on AWS's network. My user name, my home folder and anything from a key file are hidden too.

Where you see <ip> or i-<id>, your screen shows the real value. Lines that start with + are the shell printing each command just before it runs it. Only these AWS recordings are filtered: the two VS Code screenshots near the end show my laptop's own prompt.

Keep the key file outside any git folder. Mine lives in ~/lab-data/serving/keys/, and the last step deletes it.

Step 3 rents the machine. First the script asks AWS's parameter store, a noticeboard of current settings, for the newest Ubuntu 24.04 image for Arm chips. An image is a ready-made copy of a computer's software that a new machine starts from. Then run-instances asks for one c7g.medium with a 16 GiB disk that is deleted with the machine.

A real terminal recording of step 3, with the image, group and instance ids replaced by placeholders: aws ssm get-parameter reads the Ubuntu 24.04 arm64 image id; aws ec2 run-instances with that image, type c7g.medium, the key ai-research-course-chk, a 16 GiB disk of AWS's standard gp3 kind deleted on termination and the tags Project, Lesson and Name; describe-instances prints the id, c7g.medium, pending and the image id.

The machine started in the "pending" state, which is normal for its first few seconds.

  AMI=$(aws ssm get-parameter --query Parameter.Value --output text \
    --name "/aws/service/canonical/ubuntu/server/24.04/stable/current/$ARCH/hvm/ebs-gp3/ami-id")
  ID=$(aws ec2 run-instances --image-id "$AMI" --instance-type "$TYPE" --count 1 \
    --key-name ai-research-course-chk --security-group-ids "$SG" \
    --block-device-mappings '[{"DeviceName":"/dev/sda1","Ebs":{"VolumeSize":16,"VolumeType":"gp3","DeleteOnTermination":true}}]' \
    --tag-specifications \
    "ResourceType=instance,Tags=[{Key=Project,Value=ai-research-course},{Key=Lesson,Value=chk},{Key=Name,Value=ai-research-course-chk}]" \
    "ResourceType=volume,Tags=[{Key=Project,Value=ai-research-course},{Key=Lesson,Value=chk}]" \
    --query "Instances[0].InstanceId" --output text)
  aws ec2 describe-instances --instance-ids "$ID" \
    --query "Reservations[0].Instances[0].[InstanceId,InstanceType,State.Name,ImageId]" --output text
  echo "$ID" > "$STATE/id"

The line in braces is the reference check. All 26,851 rows were scored one at a time, then again in groups of 32 and in groups of 7, and every group matched the one-at-a-time scores exactly. 26,818 of the rows also matched the scores lesson 2 made on my laptop. In the other 33, the last digit or two came out different, because two kinds of chip round tiny numbers differently; lesson 11 saw this too. The three fingerprints at the bottom match lesson 2's files.

I recorded this step twice. The first recording was too short for its window and lost its top lines. So I ran the step again in a taller window, and nothing was timed in it. The whole step, run after step 4 has set IP:

  S=(ssh -i "$KEY" "${SSHO[@]}" ubuntu@"$IP")
  C=(scp -q -i "$KEY" "${SSHO[@]}")
  "${S[@]}" "mkdir -p $B/raw $B/features/results $B/serving/examples ~/lab-data/features"
  "${S[@]}" "curl -LsSf https://astral.sh/uv/install.sh | sh > /dev/null 2>&1 && \
    ~/.local/bin/uv python install 3.13.15 > /dev/null 2>&1 && \
    ~/.local/bin/uv venv --allow-existing --python 3.13.15 $B/venv > /dev/null 2>&1"
  "${C[@]}" "$HERE/examples/lat_requirements.txt" ubuntu@"$IP":$B/
  "${S[@]}" "~/.local/bin/uv pip install --python $B/venv/bin/python -q -r $B/lat_requirements.txt"
  "${C[@]}" "$LAT/model.pkl" "$LAT/store.parquet" "$LAT/requests.parquet" \
    "$HERE/lat_files/service.py" "$HERE/lat_files/build_store.py" "$HERE/lat_files/box_info.sh" \
    "$HERE/fb_files/fb_store.py" "$HERE"/chk_files/chk_*.py "$HERE"/chk_files/chk_*.sh \
    "$STATE/prev.parquet" "$STATE/fb_default.json" ubuntu@"$IP":$B/
  "${C[@]}" "$FEAT/task.py" "$FEAT/what_a_feature_is.py" ubuntu@"$IP":$B/features/
  "${C[@]}" "$FEAT/results/data-manifest.json" ubuntu@"$IP":$B/features/results/
  "${C[@]}" "$HOME/lab-data/features/retail.parquet" ubuntu@"$IP":lab-data/features/
  "${C[@]}" "$HERE/examples/chk_demo.py" "$HERE/examples/lat_requirements.txt" ubuntu@"$IP":$B/serving/examples/
  "${S[@]}" "cd $B && OMP_NUM_THREADS=1 venv/bin/python build_store.py store.parquet store.sqlite \
    && OMP_NUM_THREADS=1 venv/bin/python chk_ref.py | tee raw/ref-build.json && sha256sum model.pkl store.parquet requests.parquet"

Step 6 quiets the machine. Ubuntu runs small jobs on a timer, such as checking for updates, and one of those starting in the middle of a run would land in the results. So the script stops every timer first, then prints the chip, the load and every installed package with its version.

A real terminal recording of step 6, with the address replaced by a placeholder. Over ssh: every active timer and the unattended-upgrades service are stopped, after which systemctl prints 0 timers listed; lscpu prints architecture aarch64, 1 CPU, vendor ARM, model name Neoverse-V1, 1 thread per core, 1 core per socket; nproc prints 1; uptime shows a load average of 0.66. Then Python 3.13.15 and the full pip freeze, a list of every installed package with its version, from annotated-doc to uvloop.

The machine reports one CPU, a Neoverse-V1, and 0 timers left running:

  ssh -i ~/lab-data/serving/keys/ai-research-course-chk.pem "${SSHO[@]}" ubuntu@"$IP" \
    'for t in $(systemctl list-units --type=timer --state=active --no-legend --plain | cut -d" " -f1); do
       sudo systemctl stop "$t"; done; sudo systemctl stop unattended-upgrades.service
     systemctl list-timers --no-pager | tail -2'
  ssh -i ~/lab-data/serving/keys/ai-research-course-chk.pem "${SSHO[@]}" ubuntu@"$IP" 'lscpu | head -14; nproc; uptime'
  ssh -i ~/lab-data/serving/keys/ai-research-course-chk.pem "${SSHO[@]}" ubuntu@"$IP" \
    'cd chk && bash chk_box_info.sh before > raw/box-info-before.txt 2>&1; venv/bin/python --version; \
     ~/.local/bin/uv pip freeze --python venv/bin/python'

The four log-ins I used to find the fault were not recorded. These are the commands I ran in them, all read-only. I have shortened the long ssh options and the longest Python lines to ...:

# not recorded: looking for the fault, while the first schedule had stopped
ssh ... ubuntu@$IP 'cd chk; date -u; ps -eo pid,etime,pcpu,args --sort=-pcpu | head -12; ls raw | head -30; tail -5 raw/runs.log; cat raw/schedule.log | tail -5'
ssh ... ubuntu@$IP 'cd chk; cat raw/schedule.log; head -60 raw/runs.log; python3 -c "... read raw/checked-1100-r1-env.json ..."'
ssh ... ubuntu@$IP 'cd chk; ls -la raw/checked-1100-r1-env.json raw/*.log; sudo journalctl --since "12:18:30" --until "12:20:00" --no-pager -q | wc -c; ...'
ssh ... ubuntu@$IP 'cd chk; venv/bin/python -c "... print the size of the run record inside raw/checked-1100-r1.npz ..."'

The first try's 13 files came home with the rest in step 12, inside raw/first-attempt/, but they are not in results/chk-raw/ and nothing in this lesson uses them.

Step 8 watches without touching. The schedule writes one progress line to the machine's serial console after every run. The serial console is a log that AWS keeps for each machine, and you can read it through AWS without logging in.

A real terminal recording of step 8, with the instance id replaced by a placeholder: aws ec2 get-console-output, piped through grep CHK-PROGRESS and tail -12, prints the two priming runs, not counted, then one line per run, such as checked-1100-r1 first=1760.6 p99=25.553 standin=0.89 FAIL, naive-300-r1 p99=32.49 FAIL, checked-600-r1 p99=24.67 standin=0.137 FAIL, and checked-150-r1 first=1714.2 p99=23.81 standin=0.0142 pass.

Each line names a run, its first-answer time, its slowest 1 in 100, its share of stand-ins, and pass or FAIL. I drew this one recording from its saved session 102 characters wide, not 100. That way a long decimal number in it does not break across two lines.

  aws ec2 get-console-output --instance-id "$ID" --latest --output text | grep CHK-PROGRESS | tail -12

Every count printed 0 and the machine printed terminated. The security group can only be deleted once the machine is fully gone, so the script retries that command every ten seconds until AWS accepts it.

  aws ec2 terminate-instances --instance-ids "$ID" --query "TerminatingInstances[0].CurrentState.Name" --output text
  aws ec2 wait instance-terminated --instance-ids "$ID"
  aws ec2 delete-key-pair --key-name ai-research-course-chk --query Return
  until aws ec2 delete-security-group --group-id "$SG" --query Return 2>/dev/null; do sleep 10; done
  date -u +%Y-%m-%dT%H:%M:%S+00:00 > "$STATE/terminated"
  aws ec2 describe-instances --instance-ids "$ID" --query "Reservations[0].Instances[0].State.Name" --output text
  aws ec2 describe-key-pairs --filters "Name=key-name,Values=ai-research-course-chk" --query "length(KeyPairs)"
  aws ec2 describe-security-groups --filters "Name=group-name,Values=ai-research-course-chk" \
    --query "length(SecurityGroups)"
  aws ec2 describe-volumes --filters "Name=tag:Lesson,Values=chk" --query "length(Volumes)"
  aws ec2 describe-instances --filters "Name=tag:Project,Values=ai-research-course" \
    "Name=instance-state-name,Values=pending,running,stopping,stopped,shutting-down" \
    --query "length(Reservations[].Instances[])"
  rm -f ~/lab-data/serving/keys/ai-research-course-chk.pem

Workers and threads. On one core, one worker with one thread answered the most: 1,071 requests a second. Every other mix gave fewer.

  • Cold start. A fresh service gave its first answer 1,499 ms after it started, and 96.9% of the time before its door opened went to Python loading its tools.

  • Load testing honestly. One service with random stalls read 0.97 ms for the slowest 1 in 100 under a polite tester. Under a tester that kept to its timetable, it read 84.95 ms.

  • Caching predictions. A 30-day cache keyed by customer would have answered 43.0% of real visits from the shelf, and every one of those answers was stale.

  • Timeouts and fallbacks. With 1 lookup in 100 hanging, the slowest 1 in 100 took 1,030 ms with no limit. With a 25 ms limit and a stand-in score, it took 25.7 ms.

  • One box, many models. Four workers that each loaded their own models used 1,080 MB; four that shared models loaded once before copying used 385 MB.

  • Cost per prediction. On this machine type, a million answers one at a time cost $0.0110 at the limit, and $0.0037 with batching.

  • This lesson: the checked service met the target up to 450 requests a second in the middle round, set by the helper, not the core.

  • scoring was 54% of a 1,075.9-microsecond request
    never skip it, but do not copy the numbers: they belong to one small model on one quiet core
    Watch the slowest 1 in 100, and the first request on its own3the slowest request was 7.7 times the middle one, and it was the firsta nightly job, where nobody waits for one answer
    Batch the rows that are waiting, up to 32, never waiting for more4the slowest 1 in 100 fell from 5.26 to 2.50 ms at about 550 requests a secondwhen requests rarely overlap: here a call carried 1.2 to 1.6 rows at the rates that passed
    One worker and one thread per core51 worker with 1 thread gave 1,071 answers a second, the most of 11 mixeson one core the library already picks 1 thread; with more cores, one worker per core is the idea to test, and this lab could not
    Warm up before the service opens its door6the first request took 4.7 ms cold and 1.7 ms after one warm-upwhen a readiness check already sends the service a practice request before users can reach it, or when no user arrives in the first seconds
    Load test with planned arrivals, timed from the plan7one service read 0.97 ms under a polite tester and 84.95 ms under an honest onewhen you only want the most answers a second the service can give
    Cache answers only if inputs repeat and stale answers are safe843.0% of real visits would have hit a 30-day cache, and every hit was stalewhen no request repeats, as in this lab, or when the request itself is the news
    One time limit for the whole request, then last month's score9the slowest 1 in 100 fell from 1,030 ms to 25.7 ms with a 25 ms limit on the lookup; this lesson used the whole-request forma nightly job, or a helper that never stalls
    With several workers, load the model before copying them, then gc.freeze104 workers used 385 MB instead of 1,080 MBwith one worker, as in this lab, there is nothing to share
    Price each answer at your target, not by the hour11a million answers cost $0.0110 at the limit on this machine typewhen the machine is idle most of the day: then divide by how busy it really is
    Measure the whole service together, against a plain one12, this lessonthe checked service met the target up to 450 a second in the middle round; the naive service's failure was built into the storewhen your traffic, model or helper differs: run the race again on yours

    The switch panel above holds these decisions in one picture. Switches up were on in the checked service; on this box, the first two were true of the naive service as well. The two under guard covers are lines this lab could not use: there was one worker, and no request repeated. The lamps are checks that every number must pass before you trust it.

  • "G8. Checked at 300 a second: the mean batch of live answers is at most 1.3 rows; at its max rate at least 1.5." Partly right. At 300 a second it was 1.41 rows, more than I guessed. At each round's own max rate it was at least 1.5 in 4 of 5 rounds, and 1.41 in the round that stopped at 300.

  • "G9. The service's memory (VmRSS at the end): checked within 10% of naive at every rate." Right. VmRSS is the memory the program is holding right now, and the two were always within 10%.

  • "G10. Every live score equals the reference to the bit in both arms, and every row answered live by both arms of a pair has the same score in both (a check)." Right, and built in: 0 of 540,003 for the naive service, 0 of 228,096 for the checked one, and 0 of 228,096 between them.

  • "G11. At 1,000 a second and above: naive's slowest 1 in 100 above 1,000 ms in every round; checked fails on the live condition (stand-ins above 5%) in every round." Right.

  • "G12. The checked arm's cost per million answers at its max rate between $0.011 and $0.017 (arithmetic from G3)." Wrong, because G3 was wrong: $0.0224 in the middle round. This one is only arithmetic once the max rate is known.

  • You saw the demo running on the rented box in step 11. On that one core, both services gave their first answer after about 1.7 to 1.9 seconds. In that single run it was 1,864 ms for the naive one and 1,715 ms for the checked one. The naive service's slowest 1 in 100 was 1,125.05 ms and its slowest answer 1,994.6 ms, because it waited for the hung lookups. The checked service's were 24.09 and 25.5 ms, with 1.88% stand-ins.

    The 1,983 rows that both answered live had matching scores, to the last digit. These are one run's numbers on one machine, so read them as a picture of the lab, not as a second measurement of it.

    Here is the demo in VS Code on my laptop, run with python chk_demo.py in the examples folder with venv-sv active.

    A real screenshot of VS Code's terminal on my laptop, in the venv-sv environment, after python chk_demo.py. The first line says this machine has 10 cores and a 1-minute load of 8.16, so it was very busy. Model trained: test AP 0.5450, 26,851 test rows, stand-in score 0.2342. Requests arrive at random, 200 a second for 10 s, from the moment the port opens; the pretend store has 16 slots and 1 lookup in 100 hangs for 1 to 2 s. A table, all times in ms: A naive, first answer 2601, first request 23.71, middle 2009.36, slowest 1 in 100 4434.40, slowest 5423.9, stand-ins 0.00%; B checked, 2199, 4.86, 4.12, 25.01, 28.9, 1.88%. Last line: rows answered live by both 1,983; scores that differ between them, to the bit, 0. The terminal was narrow, so the stand-ins column wraps onto the next line.

    This run is from my laptop, not from the rented box. It is a different computer: it has 10 cores, not 1, and it was very busy. Its 1-minute load was 8.16, which means that about 8 programs were waiting for a turn, on average. So do not set these numbers beside the box's numbers one by one. Compare the two rows of this one run.

    The checked service behaved as it did on the box. Its middle answer took 4.12 ms. Its slowest 1 in 100 took 25.01 ms, just over the line. 1.88% of its answers were stand-ins, the share it gave on the box, because the pretend store rolls identical dice on both computers. The 1,983 rows that both services answered with a live score had matching scores, to the last digit.

    The naive service did much worse here. Its middle answer took about 2 seconds, not 4 ms, and its slowest took 5.4 seconds. That means it fell behind. A request came about every 5 ms. If one answer takes longer than 5 ms, requests pile up, like cars behind a slow toll gate, and the line grows for the whole 10 seconds.

    Why was it so slow? The naive service lets the model choose how many helpers, called threads, to use. On a laptop with 10 cores it uses many, and lesson 5 found that more threads made this small model slower. On a busy laptop the threads also wait for other programs.

    When I tried it later on the same laptop, one model call took 1.83 ms with the model's own choice of threads and 0.16 ms with one thread. That test was after the run and at a quieter moment, so it is my likely explanation, not a measurement of this run. The checked service uses one thread, and it kept up.

    The first answers, 2.6 and 2.2 seconds, came later than on the box, because the busy laptop was slower to start Python.

    chk_files/chk_run.sh wraps one run. It starts a fresh store and waits for an idle core. It records the load and the package list's fingerprint before and after the run. Then it runs the tester and keeps the system's own log lines. chk_files/chk_schedule.py runs the 5 rounds, each in its own shuffled order. chk_files/chk_ref.py builds the reference scores and checks that scoring in groups of 32 or 7 gives the one-at-a-time numbers.

    The slow store is lesson 9's fb_files/fb_store.py, copied unchanged. chk_report.py does not import any of the above; it reads the stored raw files with its own code and works out every number again.