Serving And Inference

Timeouts and Fallbacks: What Should a Service Return When Its Lookup Is Slow?

0 of 33 complete

0%

Contents

Back|Serving And InferenceTimeouts and Fallbacks: What Should a Service Return When Its Lookup Is Slow?
1/33
90 min left
  1. Home
  2. AI Engineering: Data, RAG and Agents
  3. Serving and Inference Basics
  4. Timeouts and Fallbacks: What Should a Service Return When Its Lookup Is Slow?
Prerequisites
Load Testing Honestly: Is Your Load Test Lying?requiredCaching Predictions: Every Real Hit Was Stale, and Dropping Old Entries Did Not HelprequiredLatency Anatomy: Where the Time in One Prediction Request GoesrequiredTail Latency: What Sits in the Slowest 1% of Requestsrequired
Related Topics
Model Signatures: The Right Numbers in the Wrong Shape, and What a Schema Check CatchesPackaging, Registry and Versioning
1 of 33
Previous lessonCaching Predictions: Every Real Hit Was Stale, and Dropping Old Entries Did Not Help

System Design

  • Foundation
  • Intermediate
  • Advanced
  • Capstone

AI Engineering

  • Foundation
  • Data, RAG and Agents
  • Evaluation, LLM Ops and Security

systemdesign.academy

  • Home
  • Glossary
  • Interview prep
  • Reviews
  • About
  • Privacy
  • Terms

The Late Order

Let me start in a small cafe.

You order a meal at the counter. The waiter takes your order to the kitchen. Most days it comes back in two minutes. Today the kitchen is stuck, and your plate does not come.

What should the waiter do?

He could do nothing and let you stand there. You might wait twenty minutes. The people behind you wait too, because the waiter is busy with you.

He could come back after five minutes and say "sorry, no food today". You leave hungry, but at least you know.

He could come back after five minutes with something that is ready now: a sandwich from the fridge, or the same meal you had last week. It is not exactly what you asked for, but it is food, and it is on time.

He could also run back into the kitchen and shout the order again. If the kitchen is stuck because it is too busy, shouting every order twice makes it busier still.

A flat illustration of a small cafe: a customer in a mustard sweater stands at a wooden counter and looks up at a round wall clock, while a smiling waiter in a navy apron places a ready plate of food on the counter in front of her, with the kitchen door open behind him. Below the picture: the customer watches the clock; the waiter does not keep her waiting for the kitchen, he brings a plate that is ready now.

A prediction service has the same problem. Before it can score a customer, it must fetch that customer's numbers from another system. Most of the time that takes two thousandths of a second. Sometimes it takes two whole seconds. In this lesson I measure what each of the waiter's choices does, on a real rented computer.

Where This Lesson Starts

This is lesson 9 of the chapter on serving a model. It uses the same model and the same small web service as the lessons before it.

The model is the features chapter's model. It looks at six numbers about one customer of a real online shop, such as the days since their last order. It gives a score between 0 and 1: how likely they are to buy again in the next 30 days.

How good is it? I measure that with AP, average precision. Think of a teacher ranking students for a prize: AP is high when the real winners are near the top of the list. On the months I kept aside for checking, this model's AP is 0.5450. Higher is better.

The service is the small web service from lesson 2, latency anatomy. A request comes in. The service looks up the customer's six numbers in a small database. The model scores them, and the answer goes back.

Lesson 7, load testing honestly, built the tester I use here: the program that sends requests and times each answer. Lesson 8 of the features chapter, missing at serving, measured what a model loses when a customer's numbers are missing. I lean on both and do not repeat them.

The machine has one core. A core is one worker inside a computer's chip, and it runs one program step at a time. My AWS account may rent only one core at a time, so, as in lessons 2 to 8, I rent a small computer with just one. The tester, the service and the slow database all share it.

The question: when the database the service depends on is slow, what should the service send back?

The Words You Need First

Please read this slide slowly if any word is new. Every slide after it uses these words, and each one comes with a picture from the cafe first.

A hand-drawn cafe with nine red marker notes, each pinned to a part of the scene. A yellow clock: timeout, stop waiting after 25 ms. A customer in a blue shirt holds a ticket reading 25 ms, with an arrow from the note deadline, one budget for the whole visit. A waiter in an apron stands at the same counter, and a black arrow runs from him to a big stove: retry, shout the order again. The stove has 16 burners, 13 red and 3 white: dependency, the kitchen (the feature store), and an arrow to one burner: slot, one of 16 burners (cooks). Under the counter, a dotted line runs from the customer to the waiter: latency, the wait, this dotted line. An arrow points from the note fallback, a plate that is ready now (a stand-in), to a fridge with a door and two handles. A small second stove with 4 burners: hedge, ask a second kitchen too. An hourglass: backoff and jitter, a short, random pause before each retry. Caption: 1 second = 1,000 milliseconds (ms).

is how long you wait for an answer. In the cafe, it is the time from ordering to getting your plate. Here I measure it in thousandths of a second, called milliseconds (ms). One second is 1,000 ms.

A dependency is another system your service needs before it can answer. The kitchen is the waiter's dependency. Here the dependency is the : the small database that holds each customer's six numbers.

A timeout is a time limit on waiting. "If the kitchen has not sent the plate in five minutes, stop waiting." In code it is one number, such as 25 ms.

A fallback is what you give back when the real answer did not come in time. In the cafe it is the sandwich from the fridge. Here it is a stand-in score: a score that is not the model's fresh answer. I try three kinds.

  • The last-known score is the score the same customer got one month earlier, kept in memory.
  • The is one number for everyone: the share of buyers in the training data, 23.4%.

Four Answers to a Slow Lookup

When the lookup is slow, the service has four basic choices. They are the waiter's four choices from the cafe.

Four cards in a row, each with a line icon. Wait: no limit, the slowest 1 in 100 took 1,030 ms. Give up: stop at 25 ms; 1.6% errors. Stand-in: stop at 25 ms; 1.6% get last month's score. Ask again: calm at 300 a second; a storm at 450. Under the cards, two bands: each needs a time limit, chosen early; each costs time, ranking, or store load. Caption: measured on the box at 300 requests a second while 1 lookup in 100 hung; median of 5 rounds.

Wait means no time limit at all. Give up means a time limit, and then an error message instead of a score. Stand-in means a time limit, and then a score that is ready now. Ask again means a retry.

The two bands under the cards are the two things I want you to remember from this slide. Every choice needs a time limit that you pick before anything goes wrong. And every choice costs something: the user's time, the quality of the ranking, or extra work for the store.

The Headline: A Time Limit Moves the Needle

Here is the main result first. Each dial shows the slowest 1 in 100. Its scale runs from 1 ms on the left to 3 seconds on the right, and each mark is 10 times the one before it. One lookup in every 100 froze for 1 to 2 seconds, and requests arrived at 300 a second.

Four dials, each a half circle from 1 ms to 1 s and more, where each mark is 10 times the one before, with a needle. Wait: the needle sits near 1 s; slowest 1 in 100: 1,030 ms. 25 ms, stand-in: the needle sits just past 10; slowest 1 in 100: 25.7 ms; a shaded band starts at 25 ms. 25 ms, retry twice: 29.1 ms. Ask twice at 10 ms: 16.6 ms. Caption: the healthy store, with no hangs at all: 7.3 ms.

With no time limit, the slowest 1 in 100 took 1,030 ms. About one request in a hundred hit a frozen lookup and waited a second or more. The other 99 were fast.

With a 25 ms limit and a stand-in score, it was 25.7 ms, forty times shorter. The price was that 1.62% of answers were stand-ins, not the model's fresh score. (This is the middle of the 5 rounds; all the rounds together give 1.66%.)

Trying again, up to twice, brought the stand-ins to zero at this rate. That is close to expected: each try draws its own wait, and a second try rarely freezes too. The measured part is the price: 1.7% more calls to the store, and a slowest 1 in 100 of 29.1 ms.

Asking twice after 10 ms gave the shortest wait of all: 16.6 ms, for 6.6% more calls. This held here because each second try drew its own random wait from a store that needs no work from the core. If the whole store is slow, both tries are slow.

One result turned all of this on its head. At 450 requests a second, trying again turned one bad second into a flood that never ended, in 3 of 5 rounds. I come to that later. It is the most important finding of the lesson.

How the Lab Was Built

I wrote the lab's design in the note at the top of the program file scripts/labs/serving/timeouts_fallbacks.py on 2026-10-11, before any timed run. It names every way of running the service, the two speeds, my rule for "clearly different", and fifteen guesses.

Before writing it, I also ran a small make-believe version on my laptop: a pretend store and pretend requests, with no real service. I used it only to choose two settings: the store's 16 slots and the two speeds. I wanted the store busy at 300 requests a second and near its limit at 450. I say so in the note, and no number from it is in this lesson.

The store. In lesson 2, the service read each customer's six numbers from a small database inside its own program. Here I moved that lookup into a separate small program, the pretend , on the same computer. The service asks it over a network connection, the way a real service asks a real feature store. The store holds lesson 2's table and runs lesson 2's lookup. Before it answers, it waits.

The wait comes in three kinds.

  • Fixed: every lookup waits 2 ms. This is a healthy store.
  • Long tail: most lookups wait about 2 ms, but a few take 10, 20 or 40 ms. It is like a queue where most people take 2 minutes and a few take 20.
  • Freezes: the long tail, but 1 lookup in 100 instead freezes for 1 to 2 seconds.

The wait of every lookup is decided in advance from a seed, a number that fixes a random draw so it can be repeated exactly. It is like writing down the dice rolls first, so every way of running the service faces the same rolls.

The ways. I call one way of running the service an arm, like one recipe in a cooking test. There are 21 of them. Some wait. Some give up at 10, 25, 50 or 100 ms, or use a 25 ms limit with a stand-in. Others use a deadline, pass the deadline to the store, try again with or without a pause, or ask twice.

The Store, and Why Slowness Spreads

Before the results, one picture of how a slow store hurts everyone, not only the unlucky request.

An isometric drawing of the pretend store as a base with 16 slot cylinders in two rows. 13 are tall and dark, marked in red marker: 13 of 16 slots busy (average over that second); 3 are short and pale, marked 3 free. Caption: tall cylinders are busy slots; the short ones were free that second.

Think of the kitchen with 16 cooks. A normal order takes a cook two minutes. A frozen order keeps a cook busy for an hour. If several frozen orders arrive close together, most cooks are stuck. Then every new order waits in line, even the easy ones.

That is what this drawing shows, from a real second of a real run at 450 requests a second, with no retries. Between 2 and 3 seconds into the run, 13 of the store's 16 slots were busy on average. In that second, 183 of 475 requests got a stand-in. A slow dependency is like a pipe with something stuck in it: everything behind the blockage slows down too.

The Machine

Like lessons 2 to 8, every timing behind a finding comes from a small computer I rented from Amazon Web Services.

Two isometric chips, each with three blocks standing on it, one per program, each block's height its share of the one core, with a logo and a percentage above it. At 300 requests a second: tester (Python logo) 3%, service (FastAPI logo) 32%, store (SQLite logo) 5%. At 450 requests a second: tester 5%, service 45%, store 7%. Top right, the Amazon EC2 and Ubuntu logos, labelled an Amazon EC2 machine, Ubuntu 24.04. Caption: every timing behind a finding comes from this box.

The computer is an EC2 c7g.medium. Amazon's own records list it with 1 core and 2 GB of memory, and its speed does not drop after a busy minute. The tester, the service and the store all take turns on that one core. The blocks show how much of it each one used. The service used the most: about a third of the core at 300 requests a second, and almost half at 450.

Before every run, a fresh store and a fresh service were started. Then the core had to be doing nothing else for at least 95 of every 100 moments, over two seconds, before the tester began. In all 110 runs it was idle at least 99.5 of every 100 moments. Steal, time that the big computer underneath lent to other customers' machines, was zero. Nobody was logged in at the start or end of any run.

The machine's diary, its system log, had 64 lines during the runs, in 12 of the 110 runs. Some came from its timetable program. Others came from an Amazon helper program that failed to phone home, because my machine had no permission to talk to it. The same kinds of lines appeared in lesson 7.

How busy do the two speeds make the core? I first measured the most answers a second the service could give, with the store at no wait at all: 934.1. So 300 a second used about a third of that, and 450 about half.

The Whole Lab in One Picture

Here is the whole lab as one numbered sequence. The numbers on the arrows match the numbers on the recordings that follow.

A sequence diagram with five columns, one per computer or program: my laptop, AWS, box: tester, box: service and box: store, and thirteen numbered arrows. 1, my laptop asks AWS: any machine running? 2, key, door rule. 3, image, rent one. 4, laptop to the tester: wait, then log in. 5, Python and files. 6, stop timers, look. 7, start, log out. 8, laptop to AWS: read the console. 9, tester to service: requests. 10, service to store: lookups, time limit. 11, laptop to tester: the demo. 12, tester to laptop: raw files. 13, laptop to AWS: delete everything. Caption: steps 1 to 8 and 11 to 13 are the recordings that follow, numbered the same; all on the box that measured the numbers.

Steps 1 to 7 happen once: check, rent the computer, set it up, and start the schedule. Step 8 is how I watched the schedule without touching the computer. Arrows 9 and 10 are the timed work: requests from the tester to the service, and lookups from the service to the store. Steps 11 to 13 come after the schedule: the student demo, the copy of the results to my laptop, and the deletion of everything.

Build the Lab Computer Yourself

The rented computer is part of the lab, so I show every step of it. You can rent the same kind of computer and run the same schedule.

Which computer you see. Every recording comes from the computer that measured this lesson's numbers. No recording and no login happened during a timed run. Steps 5 and 6 were recorded twice: the first recording of each was too short for the screen, and its top lines scrolled away. Both steps are safe to repeat, so I ran each again on the same computer. Before step 7, I also ran two short test runs of three seconds, to check the run script worked. I deleted their files, and no number in this lesson comes from them.

What you need first. An AWS account, the AWS command line tool (aws) set up with your keys, and ssh, scp, rsync and curl. All the steps live in one file, scripts/labs/serving/fb_files/fb_aws_steps.sh. You run it from the scripts/labs/serving folder, one step at a time: bash fb_files/fb_aws_steps.sh check, then access, and so on. Every code block below is copied word for word from that file.

The blocks use a few names that the file sets at its top. If you paste the blocks by hand, run these lines first, from the scripts/labs/serving folder. All of them are copied from the file, except HERE=$PWD. The file works out that folder from its own location, which a pasted line cannot do.

Steps 1 to 3: Check, Lock the Door, Rent the Computer

Step 1: is any other course computer running? My account allows one computer of this kind at a time, and another lesson may be using it. So the first command counts the course's computers that are not yet deleted. It must print 0.

A real terminal recording of step 1. The shell prints the command aws ec2 describe-instances with filters for the tag Project=ai-research-course and the states pending, running, stopping, stopped and shutting-down, asking for the number of instances; the answer is 0.

The recording shows the command and its answer, 0, so no other course computer was alive. This is the command:

  aws ec2 describe-instances --filters "Name=tag:Project,Values=ai-research-course" \
    "Name=instance-state-name,Values=pending,running,stopping,stopped,shutting-down" \
    --query "length(Reservations[].Instances[])"

Step 2: a key and a locked door. A key pair is how you prove to the computer that you are allowed in. Amazon keeps one half, and you keep the other half in a file only you can read. A security group is a door rule. Mine opens only door 22, the one that SSH (the safe remote login) uses, and only to my own address. Both get tags, small labels that say which project and lesson they belong to.

A real terminal recording of step 2, with the address, network and group ids replaced by placeholders. create-key-pair for ai-research-course-fb with the tags Project and Lesson writes the key to the key file without printing it; chmod 600 on the key file; curl reads my address; describe-vpcs finds the default network; create-security-group makes the group ai-research-course-fb with its tags; authorize-security-group-ingress prints a small table: from my address /32, port 22. Last line: the key file, readable only by its owner, 388 bytes.

The key went straight into the file and was never printed, and the door rule names only my own address. These are the commands:

  aws ec2 create-key-pair --key-name ai-research-course-fb --key-type ed25519 \
    --tag-specifications "ResourceType=key-pair,Tags=[{Key=Project,Value=ai-research-course},{Key=Lesson,Value=fb}]" \
    --query KeyMaterial --output text > ~/lab-data/serving/keys/ai-research-course-fb.pem
  chmod 600 ~/lab-data/serving/keys/ai-research-course-fb.pem
  MYIP=$(curl -sf https://checkip.amazonaws.com)
  VPC=$(aws ec2 describe-vpcs --filters Name=isDefault,Values=true --query "Vpcs[0].VpcId" --output text)
  SG=$(aws ec2 create-security-group --group-name ai-research-course-fb --vpc-id "$VPC" \
    --description "lesson 9 timeouts and fallbacks, ssh from one address" \
    --tag-specifications "ResourceType=security-group,Tags=[{Key=Project,Value=ai-research-course},{Key=Lesson,Value=fb}]" \
    --query GroupId --output text)
  aws ec2 authorize-security-group-ingress --group-id "$SG" --protocol tcp --port 22 --cidr "$MYIP/32" \
    --query "SecurityGroupRules[].{port:FromPort,from:CidrIpv4}" --output table
  echo "$SG" > "$STATE/sg"

Steps 4 to 7: Wait, Set Up, Look, Start

Step 4: wait until it runs. The AWS tool waits until the computer is running. Then the script reads its public address into IP, and tries SSH every five seconds until the computer answers.

A real terminal recording of step 4, with the instance id and address replaced by placeholders. aws ec2 wait instance-running; describe-instances reads the public address into IP; describe-instances prints running, c7g.medium, us-east-1a. Then ssh with ConnectTimeout=5 runs true and fails once, followed by sleep 5; the next ssh answers, and the last prints: ssh works: aarch64 Ubuntu 24.04.5 LTS.

The first SSH try failed because the computer was still starting, and the second answered. These are the main commands. The last line keeps the start time, which the cost step needs:

  aws ec2 wait instance-running --instance-ids "$ID"
  IP=$(aws ec2 describe-instances --instance-ids "$ID" \
    --query "Reservations[0].Instances[0].PublicIpAddress" --output text)
  until ssh -i ~/lab-data/serving/keys/ai-research-course-fb.pem "${SSHO[@]}" -o ConnectTimeout=5 \
    ubuntu@"$IP" true 2>/dev/null; do sleep 5; done
  echo "$IP" > "$STATE/ip"
  aws ec2 describe-instances --instance-ids "$ID" --query "Reservations[0].Instances[0].LaunchTime" \
    --output text > "$STATE/launched"

Step 5: Python and the lab files. The computer gets the same Python, 3.13.15, and the same tool versions as lessons 2 to 8. They go into a venv, a private toolbox folder for Python, with every tool's exact version written down. uv is a fast installer that fills it.

Then scp, a copy command that works through SSH, sends the files. The last command builds the database of customer numbers. It works out the reference score for every request (the score this computer gives that row, which every answer is checked against) and the last-known score. Then it prints the files' sha256 fingerprints, which change if even one byte of a file changes.

A real terminal recording of step 5, with the address replaced by a placeholder and my home folder shortened to a tilde. Over ssh: make the folders; install uv, Python 3.13.15 and a venv; install the pinned packages. Then scp copies lesson 2's model, data and service files, this lesson's fb_files, the month-earlier numbers and the default score, the features chapter's files and data, and the demo. Last: store: 26851 rows, sqlite 3.53.1; raw/ref.npy: 26851 single-row scores, equal to lesson 2's expected column 26818 of 26851; raw/cache.npy: last-known score for 26127 of 26851 rows; the sha256 of model.pkl starts bec4d11e, store.parquet 62aadc30, requests.parquet 407d62e6.

Steps 8 and 11 to 13: Watch, Demo, Bring It Home, Delete

Step 8: watch without touching. While it runs, the schedule writes one progress line to the computer's serial console after every run. The serial console is a log that Amazon keeps for the computer, and you can read it without logging in.

A real terminal recording of step 8, with the instance id replaced by a placeholder. aws ec2 get-console-output, piped through grep FB-PROGRESS and tail -12, prints progress lines with times: schedule start K=5 seconds=30 at 02:49:23; load before the schedule 0.09 0.16 0.09 after waiting 75s; cal-r1 to cal-r5 done between 02:50:52 and 02:51:44; then H3-t10-r1, F3-none-r1 and H4-dldown-r1 done by 02:53:29.

These lines came from Amazon's copy of the console while the runs went on, so the computer was never touched. This is the command:

  aws ec2 get-console-output --instance-id "$ID" --latest --output text | grep FB-PROGRESS | tail -12

Step 11: the student demo, on the same computer. After the schedule, I ran the demo you will meet later, so you can see what it prints on one core.

A real terminal recording of step 11, with the address replaced by a placeholder: python fb_demo.py on the box, its output also written to fb-demo-run.txt. It prints: measured on THIS machine, 1 cores, load 0.08; test AP 0.5450, stand-in score 0.2342; 200 a second for 10 s, 16 slots, 1 lookup in 100 hangs. A wait: middle 3.82, 99 in 100 beat 44.44, slowest 2001.0 ms, 0.00% stand-ins, 1.000 lookups per request. B timeout: 3.88, 26.06, 28.0, 1.58%, 1.000. C retry: 3.98, 28.35, 42.7, 0.00%, 1.016. D hedge: 4.00, 15.98, 27.0, 0.05%, 1.070. B's cost in ranking on 2,021 rows: AP 0.6150 live, 0.6113 with 32 stand-ins, -0.0037.

This run happened after the last timed run had finished, so it could not disturb any timing. Its first line says "1 cores", which is how my demo prints the count. This is the command:

  ssh -i ~/lab-data/serving/keys/ai-research-course-fb.pem "${SSHO[@]}" ubuntu@"$IP" \
    'cd fb/serving/examples && ../../venv/bin/python fb_demo.py --save ~/fb/raw/fb-demo.json \
     | tee ~/fb/raw/fb-demo-run.txt'

Step 12: bring the results home first. A whole box of results was lost once in this chapter, because it was deleted before anyone copied them. So the raw files come home before anything is deleted.

A real terminal recording of step 12, with the address replaced by a placeholder. Over ssh, the box records its info after the schedule and prints the size of its raw folder, 56M; then rsync copies the box's fb/raw folder to my laptop; ls and wc count 456 files; du prints 55M for the copy.

The Lab's Report, Running

This is a real recording of the report program, fb_report.py. It ran on my laptop, but it times nothing. It reads the raw files the rented computer measured and does all the arithmetic again with its own code.

A terminal recording of fb_report.py in eight numbered sections. 1: c7g.medium, 1 core, 2048 MiB of memory; 1.234 h at $0.0363 an hour = $0.0448 plus disk $0.0022 = $0.047; terminated; nothing left behind; 110 runs, core at least 99.5% idle, 0 ssh sessions, steal 0.00%. 2: top speed 934.1 answers a second, so 300 and 450 are 0.32 and 0.48 of it; 105 counted runs, 1,073,583 answers, 966,653 live, 0 scores not equal; a table of the middle wait, the slowest 1 in 100, the slowest, stand-ins, lookups per request and cooks busy per request for all 21 ways. 3: 23 named comparisons, each clearly different or too close to call. 4: the stand-in score's cost on AP for four ways. 5: 8 guesses right, 7 wrong. 6: the demo's box run and 15 of 15 fact-check quotes found. 7: 457 stored texts scanned, 0 account ids, addresses, host names, fingerprints, resource ids, uuids or home paths. 8: the playground's real numbers and its first seed. Last line: all 1,219 checks agree with the stored lab.

The report does not use the lab's own code. It works out every wait, share and verdict again, and compares each one with the lab's results file. If any number disagreed, it would stop with an error. I checked that too. I changed one stored wait by 1 microsecond (a millionth of a second), and the report printed "1 MISMATCHES" and stopped with an error. Then I put the file back.

Section 7 scans every stored file, even the compressed data files and the terminal recordings. It looks for my account number, addresses, Amazon's names for what I rented, and any secret-looking codes. It found none.

Waiting With No Limit

The first arm is the waiter who never comes back to say anything. The service waits for the store as long as it takes.

A chart where both scales grow by 10 times at each mark: wait in ms from 1 to 3,000 along the bottom, and the percent of requests that waited longer, from 0.001 to 100, up the side; a dotted line at 25 ms. Healthy store, no limit: the curve falls steeply and ends near 16 ms. Hangs, no limit: the curve falls until about 30 ms, then runs flat at about 1% until about 1,000 ms, then falls. Hangs, 25 ms then stand-in: it follows the same curve until 25 ms, then drops straight down by about 30 ms. Caption: most requests never notice the hangs; the few that do wait hundreds of times longer.

Read each line as "how many requests waited longer than this?". Both scales grow by 10 times at each mark. For the healthy store, almost no request waited longer than 10 ms. With freezes and no limit, the line goes flat at about 1 in 100 and stays there until a full second. Those are the requests that met a frozen lookup: 0.96% of them waited one second or more. The slowest 1 in 100 took 1,030 ms (the middle of 5 rounds), and the single slowest took 1,996 ms.

The middle request hardly changed. The middle wait was 4.28 ms with freezes and 3.91 ms without. The damage is all in the slowest 1 in 100. That is why an average or a middle value can look fine while real customers wait two seconds. With a 25 ms limit and a stand-in, nobody waited that long: the slowest 1 in 100 took 25.7 ms.

Round 4 is an honest warning. In four rounds the slowest 1 in 100 took over 1,000 ms. In round 4 it took only 29.6 ms, because that round happened to draw only 66 freezes in 8,922 requests, fewer than 1 in 100. When rare events happen about 1 time in 100, the slowest-1-in-100 number can jump a lot between runs. I had guessed it would be above 1,000 ms in every round, and that guess was wrong for this reason.

A Time Limit Is a Gauge

The next four arms put a time limit on the lookup: 10, 25, 50 or 100 ms. When the limit runs out, the service answers with an error message that says "took too long" ( 504), and no score.

A dot chart of the percent of requests given an error, 0 to 10, against the time limit: 10, 25, 50 and 100 ms, with five dots per limit, one per round. At 10 ms the dots sit near 8%; at 25 ms near 1.6%; at 50 and 100 ms near 1%, one dot each a little lower. A dashed line at 1% is labelled 1 in 100 really hung. Above each column, its slowest 1 in 100: 11.9, 25.7, 50.2 and 100.0 ms. Caption: set the limit above the normal slow answers, and below the hangs.

The slowest 1 in 100 sat just above whatever limit I chose. 10 ms gave 11.9 ms, 25 ms gave 25.7 ms, and so on. That is true by design: a limit of 25 ms cannot let a lookup wait 2 seconds. So the figure prints it as a small label, not as the main result.

The main result is how many errors each limit gave. A 10 ms limit cut off 7.61% of requests. Only about 1 in 100 really froze, so a 10 ms limit also threw away about 6 in 100 lookups that were merely slow. A 25 ms limit cut off 1.64%. Limits of 50 and 100 ms cut off about 1.1% and 1.0%, close to the 1 in 100 that really froze.

How do you pick the number? Amazon's engineering articles (the AWS Builders' Library) describe their method: "we choose an acceptable rate of false timeouts (such as 0.1%)". A false timeout is a timeout on a lookup that would have answered. Then they look at how long the slowest 1 in 1,000 normal lookups take, and set the limit there. Here, 25 ms sits above the slow-but-normal lookups and far below the freezes.

The Service Gave Up; the Store Did Not

A timeout protects the person waiting. But what happens to the lookup itself? Here is one real frozen lookup from round 1 of the 25 ms stand-in arm, followed through the service and the store.

A to-scale timeline in two lanes, with the time axis broken between 30 and 1,140 ms. Service lane: a bar from 0.5 to about 25.5 ms, and a mark labelled stand-in sent at 26.5 ms. Store lane: a long bar from 1.4 ms that runs past the break and ends at a mark labelled slot free at 1,157 ms; below it, in marker, nobody is waiting now. The axis reads 0, 10, 20, 30, then 1,140, 1,150, 1,160 milliseconds after the request was meant to arrive. Caption: a time limit protects the person waiting, not the store.

The service sent the lookup half a millisecond after the request arrived. At 25 ms it stopped waiting and answered with the customer's last-known score, 26.5 ms after arrival. But the store did not know that. It kept that lookup in a slot for 1,155 ms, then sent an answer that nobody wanted.

This is the hidden cost of a timeout. My service uses Python's stopwatch tool, asyncio.wait_for. It stops my program from waiting. Python's documentation says: "If a timeout occurs, it cancels aw and raises TimeoutError." Here aw is the thing being waited for. But it only cancels the waiting on my side. The store is a separate program, and it keeps working unless I tell it to stop. That run had 90 frozen lookups, and with a time limit alone, the store finished every one of them. The next slides are about telling it.

A Fallback Is a Fork in the Path

A timeout with an error is like the waiter saying "sorry, no food". A fallback is the sandwich from the fridge.

A hand-drawn footpath on one calm ground, with grass tufts on both sides; the width of each part grows with the square root of the number of requests. A wide path labelled 44,743 requests reaches a signpost with a stopwatch marked 25 ms, where it forks. The wide branch climbs to the upper right and ends under the label live score: 44,001. A thin branch goes down and splits again: one end has a small lunchbox, labelled last month's: 720, and a thinner end has a sandwich, labelled default: 22. Caption: both branches end on time; only the wide one is the model's real answer.

Every request reaches the stopwatch at the fork. If the lookup answered within 25 ms, the request took the wide path to the model's live score: 44,001 of 44,743 requests, all 5 rounds together. The other 742 took the thin path to a stand-in.

On the machine, the stand-in was the last-known score: the model's score on the same customer's six numbers one month earlier. If a customer was new a month earlier, it sent the default score instead. Only 22 of the 742 needed the default.

A check, not a finding. The slowest 1 in 100 was 25.7 ms with an error and 25.7 ms with a stand-in. Both arms stop at the same 25 ms, so they had to match. Reading one number from memory adds no time you can see.

Off the machine, I also scored a third choice, "no score". The next slide asks what each stand-in costs the ranking.

What a Stand-In Costs the Ranking

A stand-in answers on time, but it is not the model's fresh answer. The model's job is to rank customers, so I measured the cost as a change in AP.

This part times nothing, so it ran on my laptop. I took the requests that got a stand-in on the machine, among the first 8,922 requests of each round. I gave them each kind of stand-in in turn.

I built the model 20 times, each time with different dice, and averaged, so luck in building it evens out. Each month, the model looks at customers as they were on the 1st and guesses who buys in the next 30 days. That date is the cutoff. As in the features chapter, I worked out AP for each monthly cutoff and took the average.

How sure is each number? I redid the calculation 1,000 times. Each time I picked customers at random, some twice and some not at all, the same picks for every kind of stand-in. That shows how much the answer wobbles by luck. This is called a paired bootstrap. From it I get a range that the answer falls in 95 times out of 100. If that range includes zero, the cost might really be nothing.

Two dumbbell charts of AP, each with three stand-ins: month-old (last month's score), default and none. For each, a dark dot shows AP with every answer live, a red dot AP with the stand-ins, and a red whisker the range the change falls in 95 times in 100. Left, 300 a second with a 25 ms limit, AP from 0.528 to 0.540: the month-old red dot sits on the live dot near 0.5365; default drops to about 0.5325; none to about 0.531. Right, 450 a second with retries (storm), AP from 0.36 to 0.56: the month-old red dot sits on the live dot near 0.54; default and none both drop to about 0.38. Caption: note the two scales: the left panel is zoomed in about 17 times.

Last month's score kept the ranking. On the left, 1.66% of answers were stand-ins, and last month's score changed AP by -0.0002. Its range includes zero, so the cost is too close to zero to call. On the right, retries at 450 a second gave stand-ins to 47.5% of answers. That 47.5% is an average of 3 flooded rounds and 2 calm rounds with none. Even there, last month's score changed AP by only -0.0018. That is still too close to zero to call, but a loss as large as 0.006 cannot be ruled out.

This surprised me. One possible reason is that these six numbers change slowly for most customers in one month. The features chapter's lesson 3 on feature freshness found something similar: 30-day-old numbers could not be told from fresh ones.

One Deadline for the Whole Request

A timeout limits one step. A deadline limits the whole trip. Think of a car with a full fuel tank. Every part of the journey burns from the same tank. When it is empty, the trip is over, wherever you are.

In this lab, the tester gave each request a deadline of 25 ms from the moment it was meant to arrive. The service checked how much time was left before the lookup. It kept 2 ms back for scoring and answering, and gave the lookup only the rest.

Two horizontal bars drawn to scale in ms, with a thick line at 25 ms. 25 ms limit on the lookup only, 26.4 ms: a thin piece for waiting to start, a long piece for the lookup, and a red-outlined part past the line marked 1.4 ms over the budget. One 25 ms deadline for the whole request, 24.1 ms: the same pieces, ending before the line, with a hatched piece from 23 to 25 ms marked 2 ms kept back for scoring. A key: waiting to start, the lookup, scoring and answering, the way back. Caption: the tank is filled once, when the request arrives; every step burns from it.

Part of this is built in. By design, the deadline arm gave the lookup at most 23 ms (25 ms, minus the time already spent, minus 2 ms kept back). So its slowest 1 in 100 had to be lower than the plain 25 ms limit's: 23.6 ms against 25.7 ms. A fair contest would need a plain 23 ms limit, which I did not run.

The check is that the budget held. For the slowest 1 in 100 requests, a limit on the lookup alone added up to 26.4 ms on average, 1.4 ms over the 25 ms budget. With one deadline, the same requests took 24.1 ms. The single slowest request still took 27.5 ms, so the budget is a target the code aims at, not a wall. I did not find out why that one ran over. One possible reason is that the service was busy with other requests when its time ran out, so it noticed a little late.

The price was a few more stand-ins: 1.86% against 1.62%, because the lookup got slightly less time. Google's book on running big websites (the SRE book) gives the rule plainly: "Rather than inventing a deadline when sending RPCs to backends, servers should employ deadline propagation." An RPC is a call from one program to another, like our lookup. Deadline propagation means passing the deadline along.

Passing the Deadline Down Unclogs the Store

In one more arm, the deadline also travelled with the lookup to the store. The store then stopped waiting the moment the deadline passed, and freed the cook for the next order.

Two cards side by side. The left card, only the service knows, shows 18.2 ms per order and an isometric duct with 18 round blobs on top. The right card, the store knows too, shows 3.6 ms per order and the same duct with only 4 blobs. One blob is 1 ms that the store's cooks stayed busy per order. Caption: the store stops on the deadline, so frozen lookups stop clogging it.

The store did five times less work. I measured how long the store's cooks stayed busy per order. It was 18.2 ms when only the service knew the deadline. It was 3.6 ms when the store knew it too. That was clearly different, in every round. Almost all of the saving is the freezes: without the deadline, the store kept each one for 1 to 2 seconds, for nobody. 727 lookups were cut short across the 5 rounds.

The user saw little difference here. At 300 a second, the slowest 1 in 100 and the share of stand-ins were too close to call. At 450 a second, the stand-ins were 1.79% to 2.24% with the deadline passed down. With a plain 25 ms limit they were 1.49% to 4.92%. That is too close to call again. With 16 slots and these speeds, the store had room even when freezes held some cooks for nobody.

I guessed that passing the deadline down would cut stand-ins at 450 a second, and that guess was wrong. Its benefit here was the store's work, and the next slides show why the store's work matters so much.

Retries: Calm at 300, a Flood at 450

A retry asks the store again after a timeout. At 300 requests a second, it worked well.

With up to two retries straight away, no request needed a stand-in in any round, and the slowest 1 in 100 took 29.1 ms. That is close to expected, because a second try draws its own wait and rarely freezes too. The measured part is the price: 1.7% more calls to the store.

Amazon and Google both recommend waiting a short random time before trying again, called backoff with jitter. Here it gave the same zero stand-ins, but a longer wait for the slowest: 34.1 ms against 29.1 ms, clearly different in all five rounds. At this calm speed, the pause only made slow requests slower.

Then I ran the same arms at 450 requests a second.

Five line charts, one per round, of calls reaching the store each second over 30 seconds at 450 requests a second, with one key at the top: no retry and retry twice. Round 1, stand-ins 4.05% with no retry and 94.35% with retries: the retry line jumps from about 430 to about 1,400 at 2 s, marked storm at 2 s, and stays there; the no-retry line stays near 450. Round 2, 3.45% and 83.70%: the jump comes at 5 s. Round 3, 4.92% and 59.92%: at 12 s. Rounds 4 and 5, 1.49% and 1.67% with no retry and 0.00% with retries: the two lines lie on top of each other near 450 the whole time. Caption: by my 5-round rule this is too close to call; the danger is the 3 rounds that tipped, not an average.

One bad second became a flood that did not end, in 3 of 5 rounds. In round 1, between 2 and 3 seconds in, several freezes arrived close together and filled the store's slots. Without retries, the store got through it in about a second: 4.05% stand-ins over the whole run. With retries, every lookup that timed out came back for a second and third try. The calls jumped from about 450 to about 1,400 a second, and stayed there. 94.3% of answers were stand-ins.

Rounds 2 and 3 tipped the same way, at 5 and 12 seconds, with 83.7% and 59.9% stand-ins. Each time, the flood began in the very second that the arm without retries had its worst moment.

By my own rule, this is too close to call. In 3 rounds retries were far worse (60% to 94% stand-ins, against 3% to 5%), and in 2 rounds slightly better (0%, against 1.5% to 1.7%). The danger is the 3 rounds that tipped, not an average.

The pause did not help. I tried immediate retries, a fixed pause and a random pause. All three tipped in the same 3 rounds, with shares of stand-ins within one request of each other. I had guessed that a random pause would stop the flood, and that guess was wrong.

Why the Flood Could Not End

The arithmetic is simple, and it explains both the flood and why the pause could not stop it.

Two bathtubs, each with a 16-slot drain grate at the bottom and an arrow pouring in. No retry, 450 calls a second: the water is a little under half way up, labelled 6.8 of 16 slots: it drains. Storm, 1,300 calls a second: the water is heaped above the rim and spills over the right side, down into a row of waiting squares labelled 8,590 calls waiting in line; the tub is labelled 19.5 of 16: it overflows. Caption: a pause before a retry changes when the water arrives, not how much.

Every call has a 1 in 100 chance to freeze, and a frozen call keeps a slot busy for 1.5 seconds on average. At 450 calls a second, the freezes alone keep about 6.8 slots busy. That leaves room in a 16-slot store, so a bad second drains away.

In the flood, the calls tripled to about 1,300 a second. Now the freezes alone need about 19.5 slots, more than the store has. The line of waiting calls can only grow. The line reached 8,590 calls. The core was not the problem: during the floods, the service used only 25% to 33% of it.

A flowchart that loops. A few hangs arrive close together; the 16 slots are full; new lookups wait in line; the wait passes 25 ms: the service gives up. From there, with retries: it asks again, twice, and an arrow labelled more calls goes back to new lookups wait in line. With no retry: the line drains within a second. Caption: the loop only stops if something cuts it: fewer retries, or a store that drops work nobody waits for.

So the flood is a loop. The slots fill, new lookups wait, the wait passes 25 ms, the service asks again, and that adds more lookups to the line. A pause of up to 10 ms, then up to 20 ms, changes when a retry is sent, not how many are sent. So it cannot shrink a load that is too big.

Amazon's engineering articles warn about exactly this: "When failures are caused by overload, retries that increase load can make matters significantly worse." Google's SRE book suggests a retry budget, a cap such as "no more than 60 second tries a minute" in one program: "Consider having a server-wide retry budget". I did not test a retry budget, or retries together with a deadline passed down. Both are the obvious next arms.

Asking Twice on Purpose

A hedged request sends a second lookup if the first has not answered after a short wait, and uses whichever answer comes first. Jeff Dean and Luiz Barroso describe it in "The Tail at Scale": a client "falls back on sending a secondary request after some brief delay". They suggest waiting for "the 95th-percentile expected before issuing the hedged request", which means the wait that 95 in 100 normal lookups beat. Here that is about 10 ms.

At the top, a sketch of one request, labelled not a measurement: a first try bar still waiting, and a second try bar that starts at 10 ms and answers first: used. Below, a step chart of the percent of requests slower than each wait, 0 to 35%, against wait in ms, 5 to 40, with a dotted line at 10 ms labelled second try starts. One try, 25 ms limit: falls slowly and reaches zero near 28 ms. Ask twice: falls faster after 10 ms and is near zero by about 20 ms. Caption: this held because each try drew its own random wait from a store that needs no work from the core; if the whole store is slow, both tries are slow.

It cut the slowest waits for a small price. I compare it with the plain 25 ms limit, because the hedge also had a 25 ms limit. With the long tail and no freezes, the slowest 1 in 100 fell from 23.0 to 16.0 ms, clearly different, for 5.7% more calls. With freezes, it fell from 25.7 to 16.6 ms, and stand-ins fell from 1.62% to 0.04%. Asking twice worked on one core because waiting lookups need no work from the core.

Be careful with this result. In my store, the second lookup always drew its own, separate wait. In real life the second request goes to another copy of the store. That helps only if the two copies are not slow at the same time. If the whole store is overloaded, asking twice adds load just like a retry. I did not run hedging at 450 a second.

My Fifteen Guesses Before the Run, Checked

I wrote fifteen guesses in the note at the top of the lab's file before it ran. Each is quoted here word for word. Eight were right and seven were wrong.

The guesses use short codes for the arms. The first letter is the kind of wait: H for freezes ("hangs"), L for the long tail, F for the fixed 2 ms. The number is the speed: 3 for 300 requests a second, 4 for 450.

After the dash comes the way. none is no limit, t25 is a 25 ms limit with an error, and fb25 adds a stand-in. dl is the deadline, and dldown passes it down. r2 is two retries, r2b adds a fixed pause and r2bj a random pause. hedge is asking twice.

S is the most answers a second the service could give. Inside the quotes, p99 means the slowest 1 in 100, max means the single slowest, and "told apart" is my rule for clearly different.

  1. "G1. S between 600 and 1,000 answers a second." Right: 934.1.

  2. "G2. H3-none p99 above 1,000 ms and max above 1,500 ms in every round; F3-none p99 below 10 ms in every round." (Freezes and no limit, and the healthy store.) Wrong: round 4 drew only 66 freezes in 8,922 requests, and its slowest 1 in 100 took 29.6 ms.

  3. "G3. H3-tT p99 between T and T + 10 ms in every round, for T = 10, 25, 50, 100." (A limit of T ms with an error.) Wrong, for the same round 4: with 50 and 100 ms limits it took 29.3 and 29.2 ms, below the limit.

  4. "G4. H3-t25 and H3-fb25 p99 cannot be told apart (the fallback costs no time)." (An error against a stand-in.) Right, but true by design: both stop at the same 25 ms.

  5. "G5. H3-fb25 fallback share between 1.0% and 2.5% in every round." (The stand-in arm.) Right: 1.33% to 1.91%.

  6. "G6. H3-dldown holds the store's slots for at most a third of H3-dl's slot-seconds per request, in every round." (Deadline passed down against kept by the service.) Right: 3.5 to 3.6 ms against 14.3 to 18.8 ms.

Try It Yourself

The full lab needs a rented computer and about an hour and a quarter. The demo, fb_demo.py, runs on your own computer. It starts a small service with a slow pretend store inside it. Then it tries four ways to handle the slow lookups: wait, give up at 25 ms with a stand-in, retry, and ask twice.

I wrote the demo's design in the note at the top of its file, after the lab's design and before the demo first ran.

A real screenshot of VS Code with fb_demo.py open at the top of the file, showing the note at its top: what it needs, how to run it, and the design written before it first ran.

The note lists what the demo needs and the order it works in. It trains the model, starts the service with the pretend store, runs the four ways one after another, and prints a table. It also says plainly that the times it prints come from your own computer.

Before you run this lab. The demo uses lesson 2's Python toolbox folder, with the same tool versions as the lab's computer. If you made it for lessons 2 to 8, use it again. If not, make it once, inside the scripts/labs/serving/examples folder:

python3 -m venv venv-sv
source venv-sv/bin/activate        # on Windows: venv-sv\Scripts\activate
pip install -r lat_requirements.txt

I used Python 3.13. Then run python ../../features/fetch_data.py once, which downloads the shop data. Now run python fb_demo.py. It needs no graphics card and no cloud account. Its service runs only on your own computer, never on the internet, and stops at the end. If a tool is missing, it prints one line saying why and stops.

The first line it prints says the times come from your computer, with its number of cores and how busy it is at that moment. Read the four rows against each other, on your computer. Do not set them beside the lab's numbers. One difference from the lab: in the demo the store lives inside the service, so giving up on a lookup also frees its slot.

r"""What should a service answer when its feature lookup is slow? Four answers, one slow store, on YOUR machine.

Lesson 9 of 'Serving and Inference Basics'. It uses lesson 2's Python environment, made once inside this folder
(Python 3.13 is what the lab used):
    python3 -m venv venv-sv
    source venv-sv/bin/activate              # Windows: venv-sv\Scripts\activate
    pip install -r lat_requirements.txt      # numpy, pandas, pyarrow, scikit-learn, fastapi, uvicorn, ...
It also needs the shop data from the features chapter: run python ../../features/fetch_data.py once first. Then:
    python fb_demo.py                  # train, then run the four ways one after another, print the table
    python fb_demo.py --save out.json  # and save the numbers
If a package is missing, it prints one line saying why and stops.
The timings it prints come from YOUR machine, as it is right now: other programs move them. Read the four rows
against each other, on your machine; they are not numbers to set beside the lab's.
It starts one helper process at a time (the service, on 127.0.0.1 only) and stops it. It writes nothing except a
temporary model file, deleted at the end, and out.json if you ask.

Design, written 2026-10-11 after the lab's design (timeouts_fallbacks.py) and before this file first ran:
  1. Train the features chapter's model (it must score test AP 0.5450), as lesson 7's load_demo.py does.
  2. A small FastAPI service in a second process (this same file, run with --serve). Inside it, a PRETEND feature
     store: it can work on 16 lookups at once (16 slots); each lookup waits a random time first, about 2 ms most
     of the time, with a long tail, and 1 lookup in 100 hangs for 1 to 2 s. The waits are drawn from a fixed seed,
     so all four ways meet the same waits. (The lab's store is a separate process reached over TCP; here it lives
     inside the service, to keep the file short. One difference follows: here, giving up on a lookup also stops
     it and frees its slot, which the lab's separate store does only when the deadline travels with the call.)
  3. Four ways to handle a slow lookup, each in a fresh service:
       A  wait     no time limit
       B  timeout  give up after 25 ms and answer a stand-in score (the share of buyers in the training rows)
       C  retry    like B, but try again at once, up to 2 more times, before the stand-in
       D  hedge    if the lookup has not answered after 10 ms, ask a second time; take whichever answers first,
                   within 25 ms; then the stand-in
  4. Requests arrive at random (Poisson) at 200 a second for 10 s; latency from the SCHEDULED arrival (lesson 7).
     Request j asks for test row j.
  5. Print for each: p50, p99 and max; the share of stand-in answers; lookups per request. Then, for B, the ranking
     cost: the model's AP on the requested rows with every answer live, and with B's stand-in answers.

Author: Roni Das
Created: 2026-10-11
"""
import importlib.util
import json
import os
import socket
import subprocess
import sys
import tempfile
import threading
import time
from pathlib import Path
from queue import Queue
from time import perf_counter

NEEDED = ("numpy", "pandas", "pyarrow", "sklearn", "fastapi", "uvicorn")
missing = [m for m in NEEDED if importlib.util.find_spec(m) is None]
if missing:
    sys.exit(f"fb_demo.py needs {', '.join(missing)}: make the environment in its docstring "
             f"(pip install -r lat_requirements.txt), then run it with that environment's python.")

import numpy as np  # noqa: E402

RATE, SECONDS = 200, 10.0
WAYS = {"A wait": {"timeout": None, "retries": 0, "hedge": None},
        "B timeout": {"timeout": 0.025, "retries": 0, "hedge": None},
        "C retry": {"timeout": 0.025, "retries": 2, "hedge": None},
        "D hedge": {"timeout": 0.025, "retries": 0, "hedge": 0.010}}


# ───────────────────────────── the service (python fb_demo.py --serve PORT FILE WAY) ─────────────────────────────
def serve(port: int, data_path: str, way: str) -> None:
    import asyncio
    import pickle

    import uvicorn
    from fastapi import FastAPI, Request
    from starlette.responses import Response

    os.environ["OMP_NUM_THREADS"] = "1"
    d = pickle.loads(Path(data_path).read_bytes())
    model, X, stand_in = d["model"], d["X"], d["stand_in"]
    w = WAYS[way]
    rng = np.random.default_rng(0)                    # the same waits for all four ways
    wait = 0.002 * np.exp(rng.standard_normal((len(X), 3)))
    hang = rng.random((len(X), 3)) < 0.01
    wait = np.where(hang, rng.uniform(1.0, 2.0, (len(X), 3)), wait)
    calls = [0]
    app = FastAPI()
    slots: list = []

    async def store(row: int, attempt: int) -> list[float]:
        """The pretend feature store: wait for one of 16 slots, hold it for the drawn time, answer."""
        if not slots:
            slots.append(asyncio.Semaphore(16))
        calls[0] += 1
        async with slots[0]:
            await asyncio.sleep(float(wait[row, attempt]))
        return X[row].tolist()

    async def lookup(row: int):
        if w["timeout"] is None:                                      # A: no limit
            return await store(row, 0)
        if w["hedge"] is not None:                                    # D: a second ask after 10 ms
            first = asyncio.ensure_future(store(row, 0))
            done, _ = await asyncio.wait({first}, timeout=w["hedge"])
            if done:
                return first.result()
            second = asyncio.ensure_future(store(row, 1))
            done, pending = await asyncio.wait({first, second}, timeout=w["timeout"] - w["hedge"],
                                               return_when=asyncio.FIRST_COMPLETED)
            for p in pending:
                p.cancel()                         # we stop waiting; the pretend store frees its slot
            return next(iter(done)).result() if done else None
        for attempt in range(1 + w["retries"]):                       # B and C
            try:
                return await asyncio.wait_for(store(row, attempt), w["timeout"])
            except TimeoutError:
                pass
        return None

    @app.post("/score")
    async def score(request: Request) -> Response:
        q = json.loads(await request.body())
        feats = await lookup(q["row"])
        if feats is None:
            out = {"id": q["id"], "score": stand_in, "stand_in": True}
        else:
            p = float(model.predict_proba(np.array([feats], dtype=np.float64))[0, 1])
            out = {"id": q["id"], "score": p, "stand_in": False}
        return Response(content=json.dumps(out).encode(), media_type="application/json")

    @app.get("/calls")
    async def get_calls() -> Response:
        return Response(content=json.dumps({"calls": calls[0]}).encode(), media_type="application/json")

    @app.get("/health")
    async def health() -> Response:
        return Response(content=b"ok")

    uvicorn.run(app, host="127.0.0.1", port=port, log_level="warning")


# ───────────────────────────── the model ─────────────────────────────
def build():
    """The features chapter's model and its test rows (lesson 7's load_demo.py build step)."""
    from sklearn.metrics import average_precision_score
    here = Path(__file__).resolve().parent
    sys.path.insert(0, str(here.parents[1] / "features"))
    import task
    from what_a_feature_is import HAND_COLS, hgb, joined

    ev = task.load_events()
    lab_tr, _, lab_te = task.splits(ev)
    tr = joined(ev, lab_tr, task.TRAIN_CUTOFFS)
    te = joined(ev, lab_te, task.TEST_CUTOFFS)
    model = hgb(0).fit(tr[HAND_COLS].to_numpy(float), tr["label"].to_numpy())
    X = np.ascontiguousarray(te[HAND_COLS].to_numpy(np.float64))
    p = model.predict_proba(X)[:, 1]
    y, cut = te["label"].to_numpy(), te["cutoff"].to_numpy()
    ap = float(np.mean([average_precision_score(y[cut == c], p[cut == c]) for c in np.unique(cut)]))
    assert round(ap, 4) == 0.5450, ap
    return model, X, y, cut, p, ap, float(tr["label"].mean())


# ───────────────────────────── the open-loop tester (lesson 7's) ─────────────────────────────
def connect(port: int):
    import http.client
    c = http.client.HTTPConnection("127.0.0.1", port, timeout=30)
    c.connect()
    c.sock.setsockopt(socket.IPPROTO_TCP, socket.TCP_NODELAY, 1)
    return c


def ask(c, j: int, row: int) -> dict:
    c.request("POST", "/score", body=json.dumps({"id": j, "row": row}), headers={"Content-Type": "application/json"})
    d = json.loads(c.getresponse().read())
    if d["id"] != j:
        raise SystemExit(f"answer {j} carries id {d['id']}")
    return d


def calls_so_far(port: int) -> int:
    c = connect(port)
    c.request("GET", "/calls")
    n = json.loads(c.getresponse().read())["calls"]
    c.close()
    return n


def open_loop(port: int, n_rows: int) -> tuple[np.ndarray, list[dict], int]:
    rng = np.random.default_rng(1)
    at = np.cumsum(rng.exponential(1 / RATE, size=int(RATE * SECONDS * 1.5) + 50))
    at = at[at < SECONDS]
    lats = np.zeros(len(at))
    answers: list = [None] * len(at)
    q: Queue = Queue()
    start = [0.0]

    def sender() -> None:
        c = connect(port)
        ask(c, -1, n_rows - 1)                         # warm-up of this connection, not counted
        while True:
            j = q.get()
            if j is None:
                break
            answers[j] = ask(c, j, j % n_rows)
            lats[j] = perf_counter() - (start[0] + at[j])      # from the SCHEDULED arrival
        c.close()

    pool = [threading.Thread(target=sender, daemon=True) for _ in range(64)]
    for th in pool:
        th.start()
    time.sleep(0.5)
    before = calls_so_far(port)                        # the warm-up asks, not counted
    start[0] = perf_counter() + 0.05
    for j, a in enumerate(at):
        d = start[0] + a - perf_counter()
        if d > 0:
            time.sleep(d)
        q.put(j)
    for _ in pool:
        q.put(None)
    for th in pool:
        th.join()
    return lats * 1000, answers, calls_so_far(port) - before


def run_way(way: str, data_path: str, n_rows: int) -> dict:
    s = socket.socket()
    s.bind(("127.0.0.1", 0))
    port = s.getsockname()[1]
    s.close()
    svc = subprocess.Popen([sys.executable, __file__, "--serve", str(port), data_path, way],
                           env={**os.environ, "OMP_NUM_THREADS": "1"})
    try:
        for _ in range(300):
            try:
                connect(port).close()
                break
            except OSError:
                time.sleep(0.1)
        lat, answers, calls = open_loop(port, n_rows)
    finally:
        svc.terminate()
        svc.wait()
    stand = np.array([a["stand_in"] for a in answers])
    return {"n": len(lat), "p50": float(np.percentile(lat, 50)), "p99": float(np.percentile(lat, 99)),
            "max": float(lat.max()), "stand_in_share": float(stand.mean()), "calls_per_request": calls / len(lat),
            "stand_in_rows": [j for j, a in enumerate(answers) if a["stand_in"]]}


def main() -> None:
    try:
        load = f"{os.getloadavg()[0]:.2f}"
    except (AttributeError, OSError):
        load = "not available on this system"
    print(f"Every time below was measured on THIS machine ({os.cpu_count()} cores, 1-minute load {load}).")
    model, X, y, cut, p, ap, stand_in = build()
    print(f"model trained: test AP {ap:.4f}; {len(X):,} test rows; stand-in score {stand_in:.4f}")
    import pickle
    out = {"load": load, "cores": os.cpu_count(), "test_ap": ap, "ways": {}}
    print(f"requests arrive at random, {RATE} a second for {SECONDS:.0f} s; the pretend store has 16 slots and "
          f"1 lookup in 100 hangs for 1 to 2 s\n")
    print("                p50 ms    p99 ms    max ms   stand-in   lookups per request")
    with tempfile.TemporaryDirectory() as tmp:
        dp = Path(tmp) / "data.pkl"
        dp.write_bytes(pickle.dumps({"model": model, "X": X, "stand_in": stand_in}))
        for way in WAYS:
            r = run_way(way, str(dp), len(X))
            out["ways"][way] = {k: v for k, v in r.items() if k != "stand_in_rows"}
            print(f"   {way:<10}{r['p50']:9.2f} {r['p99']:9.2f} {r['max']:9.1f} {r['stand_in_share']:9.2%}"
                  f"   {r['calls_per_request']:10.3f}")
            if way == "B timeout":
                rows_b = r["stand_in_rows"]
                n_req = r["n"]
    from sklearn.metrics import average_precision_score
    req = np.arange(n_req) % len(X)
    live = p[req]
    with_b = live.copy()
    with_b[rows_b] = stand_in
    yy, cc = y[req], cut[req]

    def ap_of(s: np.ndarray) -> float:
        return float(np.mean([average_precision_score(yy[cc == c], s[cc == c]) for c in np.unique(cc)]))

    a_live, a_b = ap_of(live), ap_of(with_b)
    print(f"\nB's cost in ranking, on the {n_req:,} requested rows: AP {a_live:.4f} with every answer live, "
          f"{a_b:.4f} with B's {len(rows_b)} stand-in answers ({a_b - a_live:+.4f})")
    out["ap"] = {"rows": n_req, "live": a_live, "with_b": a_b, "stand_ins": len(rows_b)}
    if "--save" in sys.argv:
        Path(sys.argv[sys.argv.index("--save") + 1]).write_text(json.dumps(out, indent=1))


if __name__ == "__main__":
    if len(sys.argv) > 1 and sys.argv[1] == "--serve":
        serve(int(sys.argv[2]), sys.argv[3], " ".join(sys.argv[4:]))
    else:
        main()

Play With a Make-Believe Store

This box runs in your browser and needs nothing but Python. It is a simulation: a make-believe version that counts time instead of measuring it. It has a pretend store with 16 slots, pretend waits and pretend requests. It has no model, no network and no service, so it shows how things work, not the machine's numbers. The machine's real numbers are printed under it.

Each lookup's wait is decided in advance for every request and every try, like writing down the dice rolls first. So when you switch retries or the random pause on and off, the store meets exactly the same rolls. It runs 5 seeds, because one seed is one roll of luck.

Press Run. Then try TIMEOUT_MS = None and watch the slowest 1 in 100 jump to about a second. Try RETRIES = 2: at RATE = 300 the stand-ins vanish. Then try RATE = 450 with RETRIES = 2, and watch some seeds flood while others stay calm, as on the machine. Turn on JITTER = True at that speed: the flooded seeds stay flooded. Finally try PASS_DEADLINE = True with the retries still on, and compare.

The report program writes the machine's real numbers into this box from the lab's results, and runs it. In this make-believe store, passing the deadline down stopped the flood at 450 a second. On the machine I did not test that combination, so treat it as a question, not an answer.

The Lab's Code, Piece by Piece

The lab has one file that runs on my laptop, a few small files that run on the rented computer, and the step script you saw above.

timeouts_fallbacks.py, on my laptop, holds the design and my guesses in the note at its top, written before any timed run. Its commands are prepare, check, status, collect, ap, cost and factcheck. prepare makes the last-known numbers and the default score. check makes sure lesson 2's files have the same fingerprints, and status reads the computer's console. collect copies the computer's raw measurement files into results/fb-raw/, passes every log line through fb_redact.py, and does all the arithmetic, with help from fb_stats.py. ap works out the stand-ins' cost on AP.

is the pretend . It is a separate program on the same computer. It holds lesson 2's table and runs lesson 2's lookup, but first it makes each lookup wait. The wait for every lookup is decided in advance from a seed, so every arm of a round meets exactly the same waits. It has 16 slots, and the rest wait in line.

How I Would Choose

Here is the order I would follow, using only what this lab measured.

Three rings around a dot labelled one request. Outer ring, one deadline: 25 ms for the whole trip, passed to the store: 23.5 ms. Middle ring, a stand-in ready: last month's score, AP -0.0002. Inner ring, a retry gate: open only when the store is healthy. Caption: decide what the user gets when things are slow, before they are slow.

Each ring protects the request inside it. The outer ring is the time budget, the middle ring is the stand-in, and the inner ring is the gate that decides whether to try again. Here are the same steps in words, each with the number from this lab behind it.

  1. Pick the longest wait a user should ever see. Here, 25 ms. Without any limit, the slowest 1 in 100 took 1,030 ms.

  2. Make it one deadline for the whole request, and pass it on. The budget held for the slowest requests (24.1 ms on average), and when the store knew it too, the store did five times less work.

  3. Choose the stand-in before you need it. Last month's score cost too little ranking quality to measure here. A default or "no score" cost a clear amount, and the cost grows with the share of stand-ins.

  4. Retry only a healthy store, and count your calls. At 300 a second, two retries (what I tested) removed every stand-in. At 450 a second they turned one bad second into a flood in 3 of 5 rounds, and a random pause did not help. Watch calls per request: it jumped from 1.00 to about 2.9 when the flood began.

  5. Ask twice when the slowness is random. Asking twice after 10 ms cut the slowest 1 in 100 by about a third, for about 6% more calls, because each try had its own luck. If the whole store is slow, both tries are slow.

When to Use Each One, and When Not To

No time limit: almost never. It is right only for work where nobody is waiting, such as a nightly batch job. For a user, one frozen lookup means a two-second wait.

A time limit with an error: right when a wrong answer is worse than no answer. A bank checking for fraud may prefer to say "try again" rather than guess. It is the wrong choice when the caller cannot cope with an error, because 1.64% of users would see one here.

A time limit with a stand-in: right when a slightly old or rough answer is still useful, like a ranking. Here last month's score kept the ranking almost intact. It is wrong when the stand-in misleads, for example a default credit limit for someone you know nothing about.

A deadline passed down: right almost always, when your code calls other code. Google's SRE book recommends it, and here it saved the store about four fifths of its work. It needs every program in the chain to read the deadline and stop on time.

Retries: right for a healthy dependency with rare random failures, at a calm speed. Wrong when the dependency is struggling: then each retry is more load. Here, retries tipped over at 450 a second in 3 of 5 rounds, with or without a pause.

Asking twice: right when slowness is random and each copy of the store is slow on its own. Wrong when the whole store is overloaded, because the second question is more load too.

What This Lab Cannot Tell You

An isometric scene split down the middle by a line. Left, this lab: slow one at a time: one chip with three program blocks on it and twelve small dice, only one of them dark; label: one chip, three programs. Right, real life: slow all together: three separate machines joined by a dashed network line, each with three dice, and every die dark; label: separate machines, a network. Dark dice are slow lookups.

A pretend store. My store's slowness was drawn at random, one lookup at a time. A real store is often slow for a reason: a busy disk, a bad network link, one large customer. Then many lookups are slow together, and retries and asking twice help less.

One core. The tester, the service and the store shared one core. At 450 a second the service used about 45% of it. A real store on its own machines would not take time from the service.

Arms I did not run. I did not test a retry budget, or retries together with a deadline passed down. I also did not run a plain 23 ms limit to compare fairly with the deadline, or asking twice at 450 a second.

Changes after the results. After the first collection, the stored raw copy was 42.5 MB, over my limit of about 30 MB. I changed only how times are stored (as differences between neighbours), not any measurement, and the results came out the same.

After looking at the figures, I added two measurements that the design did not name. One splits each slow request's time into steps. The other follows one frozen lookup from start to end. I also counted the freezes in each round to explain round 4.

Four recordings (steps 4, 6, 12 and 13) lost their top lines, because long commands wrapped across rows. I re-drew those four from their own recording files in a taller window. The text shown is exactly what the commands printed.

After a review, I rebuilt the playground so that switching the random pause on and off meets the same dice rolls. Before that change, it wrongly showed the pause stopping the flood.

What to Do on Monday

A hand-drawn service: request, service and store boxes joined by arrows, a small box under the service labelled stand-ins and a round meter by the store labelled calls per request. Five yellow sticky notes sit near the parts they belong to: 1, a deadline: one budget for the whole request; 2, a limit on every wait, from that budget; 3, retry only a healthy store; 4, a stand-in ready, and its cost on AP known; 5, watch calls per request: a jump means a storm. Caption: a planned stand-in on time beats a perfect answer that comes too late.

If you take one thing to work on Monday, find every place where your service waits for another system, and check that each one has a time limit. A call with no limit is a call that can make one customer wait forever.

Then look at your retries. How many tries does your code, your library and your (the program that shares requests among servers) each make? If three layers each try three times, one failure can become 27 calls. Make sure the system you call is not already struggling when you send it more.

The one idea to keep: decide what the user gets when things are slow, before things are slow. A stand-in on time is often worth more than the perfect answer two seconds late.

Knowledge Check

Knowledge Check

4 questions - Score 80% to pass

Q1

With no time limit and 1 lookup in 100 freezing, the slowest 1 in 100 took about 1,030 ms. Why did a 25 ms limit not also make the store do less work?

Q2

Which stand-in score cost the least ranking quality (AP) when the lookup was too slow?

Q3

In round 1 at 450 requests a second, retries turned one bad second into 94.3% stand-ins. What did a random pause before each retry change?

Q4

What did passing the 25 ms deadline down to the store change, at 300 requests a second?

default score
  • No score means the service says "I do not know", and whoever asked must cope.
  • A deadline is a time limit for the whole visit, not for one step. "This customer must be served within five minutes of walking in, whatever happens on the way." Each step can see how much of that time is left.

    A retry is trying again after a failure: the waiter shouts the order a second time. Backoff means waiting a little before you try again. Jitter means making that wait a bit random, so that many waiters do not all shout at the same moment.

    A hedged request is asking twice on purpose. If the first kitchen is slow, you also ask a second kitchen, and take whichever plate comes first.

    A slot is one cook in the kitchen. My pretend store has 16, so it can work on 16 lookups at the same time.

    When I talk about many requests, I line up all their waits from shortest to longest. The middle wait (also called the median, or p50) is the one in the middle: half of the requests were faster. The p99 is the time that 99 of every 100 requests beat. So it tells you about the slowest 1 in 100, and those are the customers who complain. From here on I mostly say "the slowest 1 in 100".

    Each arm ran for 30 seconds. A round is one full pass through all 21 arms. I did 5 rounds, each in a new random order, so the time of day could not favour one arm. The tester sent each request at its planned time, even if earlier answers were late, and timed each one from that planned time, as in lesson 7.

    My rule. I count a difference between two arms as real only if one beat the other in all 5 rounds, and even its worst round beat the other's best. Then I call them clearly different. If not, I call it too close to call.

    export AWS_DEFAULT_REGION=us-east-1 AWS_PAGER=""
    NAME=ai-research-course-fb
    HERE=$PWD
    STATE=${FB_STATE:-$HOME/lab-data/serving/fb}
    KEY=$HOME/lab-data/serving/keys/$NAME.pem
    K=${K:-5}; SEC=${SEC:-30}
    mkdir -p "$STATE" "$HOME/lab-data/serving/keys"
    SSHO=(-o LogLevel=ERROR -o StrictHostKeyChecking=accept-new -o "UserKnownHostsFile=$STATE/known_hosts")
    [ -f "$STATE/id" ] && ID=$(cat "$STATE/id")
    [ -f "$STATE/ip" ] && IP=$(cat "$STATE/ip")
    [ -f "$STATE/sg" ] && SG=$(cat "$STATE/sg")
    

    The copy step also needs two small files that my laptop makes first, from the features chapter's code. python timeouts_fallbacks.py prepare writes each customer's numbers from one month earlier, for the last-known score. It also writes the default score.

    What it costs. A c7g.medium costs $0.0363 an hour, from Amazon's price list, and Amazon bills it by the second. My computer was on for 1.234 hours, from start to deleted, which is $0.0448. Its 16 GB disk, deleted with it, added $0.0022, so $0.047 in all. These numbers are in results/fb-cost.json.

    What the recordings hide. The AWS recordings passed through fb_redact.py, a small program that blacks out my account number and addresses before anything is saved. It also hides the names Amazon makes up for each thing I rent, host names, my home folder and anything from a key file. Only the AWS recordings are filtered. The two VS Code pictures, taken on my laptop, are not, and show my own prompt. Where you see <ip> or i-<id>, your screen shows the real value. Lines that start with + are the shell printing each command just before it runs it.

    The last line keeps the group's name in a small file, so the later steps can read it back. Keep the key file outside any git folder. Mine lives in ~/lab-data/serving/keys/, and the last step deletes it.

    Step 3: rent the computer. An image is a ready-made copy of a computer's software, here Ubuntu 24.04. First the script asks Amazon's noticeboard, the parameter store, for the newest Ubuntu image for Arm chips. That way you never copy an image name that has gone out of date. Then run-instances asks for one c7g.medium with a 16 GB disk that is deleted with the computer, with the tags on both.

    A real terminal recording of step 3, with the image, group and instance ids replaced by placeholders. aws ssm get-parameter reads the Ubuntu 24.04 arm64 image id into AMI; aws ec2 run-instances with that image, type c7g.medium, the key ai-research-course-fb, a 16 GiB gp3 disk deleted on termination and the tags Project, Lesson and Name, asking only for the instance id; describe-instances prints the id, c7g.medium, pending and the image id.

    The computer started in the "pending" state, which is normal for the first few seconds. These are the commands:

      AMI=$(aws ssm get-parameter --query Parameter.Value --output text \
        --name /aws/service/canonical/ubuntu/server/24.04/stable/current/arm64/hvm/ebs-gp3/ami-id)
      ID=$(aws ec2 run-instances --image-id "$AMI" --instance-type c7g.medium --count 1 \
        --key-name ai-research-course-fb --security-group-ids "$SG" \
        --block-device-mappings '[{"DeviceName":"/dev/sda1","Ebs":{"VolumeSize":16,"VolumeType":"gp3","DeleteOnTermination":true}}]' \
        --tag-specifications \
        "ResourceType=instance,Tags=[{Key=Project,Value=ai-research-course},{Key=Lesson,Value=fb},{Key=Name,Value=ai-research-course-fb}]" \
        "ResourceType=volume,Tags=[{Key=Project,Value=ai-research-course},{Key=Lesson,Value=fb}]" \
        --query "Instances[0].InstanceId" --output text)
      echo "$ID" > "$STATE/id"
    

    The three fingerprints match the ones lesson 2 recorded, so the computer serves the same model. 26,818 of the 26,851 reference scores equal the scores lesson 2 stored. The other 33 differed in the 16th decimal place, a rounding difference between computers. Every answer here is checked against this computer's own reference. 26,127 rows have a last-known score; the other 724 customers were new a month earlier and get the default. These are the commands of the step, copied from the file, run after step 4 has set IP:

      B=/home/ubuntu/fb
      LAT=$HOME/lab-data/serving/lat
      FEAT=$HERE/../features
      S=(ssh -i "$KEY" "${SSHO[@]}" ubuntu@"$IP")
      C=(scp -q -i "$KEY" "${SSHO[@]}")
      "${S[@]}" "mkdir -p $B/raw $B/features/results $B/serving/examples ~/lab-data/features"
      "${S[@]}" "curl -LsSf https://astral.sh/uv/install.sh | sh > /dev/null 2>&1 && \
        ~/.local/bin/uv python install 3.13.15 > /dev/null 2>&1 && \
        ~/.local/bin/uv venv --allow-existing --python 3.13.15 $B/venv > /dev/null 2>&1"
      "${C[@]}" "$HERE/examples/lat_requirements.txt" ubuntu@"$IP":$B/
      "${S[@]}" "~/.local/bin/uv pip install --python $B/venv/bin/python -q -r $B/lat_requirements.txt"
      "${C[@]}" "$LAT/model.pkl" "$LAT/store.parquet" "$LAT/requests.parquet" \
        "$HERE/lat_files/service.py" "$HERE/lat_files/build_store.py" "$HERE/lat_files/box_info.sh" \
        "$HERE"/fb_files/fb_*.py "$HERE"/fb_files/fb_*.sh "$STATE/prev.parquet" "$STATE/fb_default.json" ubuntu@"$IP":$B/
      "${C[@]}" "$FEAT/task.py" "$FEAT/what_a_feature_is.py" ubuntu@"$IP":$B/features/
      "${C[@]}" "$FEAT/results/data-manifest.json" ubuntu@"$IP":$B/features/results/
      "${C[@]}" "$HOME/lab-data/features/retail.parquet" ubuntu@"$IP":lab-data/features/
      "${C[@]}" "$HERE/examples/fb_demo.py" "$HERE/examples/lat_requirements.txt" ubuntu@"$IP":$B/serving/examples/
      "${S[@]}" "cd $B && OMP_NUM_THREADS=1 venv/bin/python build_store.py store.parquet store.sqlite \
        && OMP_NUM_THREADS=1 venv/bin/python fb_ref.py && sha256sum model.pkl store.parquet requests.parquet"
    

    Step 6: stop the timers and look at the computer. Ubuntu runs small jobs on a timer, such as checking for updates. One of those starting in the middle of a run would land in the results, so the script stops every timer first. Then it prints the chip's details, the number of cores, the load, and every installed tool with its version.

    A real terminal recording of step 6, with the address replaced by a placeholder. Over ssh: stop every active timer and the unattended-upgrades service, after which systemctl prints 0 timers listed; lscpu prints architecture aarch64, 1 CPU, vendor ARM, model name Neoverse-V1, 1 thread per core, 1 core per socket; nproc prints 1; uptime prints up 5 min with load average 0.29, 0.21, 0.09. Then the box info is saved, and Python 3.13.15 and the full pip freeze are printed, including fastapi 0.142.2, numpy 2.5.3, scikit-learn 1.9.1, uvicorn 0.54.0 and uvloop 0.23.0.

    The computer reports one core, of an Arm chip called Neoverse-V1, and the same tool versions as lessons 2 to 8. These are the commands that stop the timers and look:

      ssh -i ~/lab-data/serving/keys/ai-research-course-fb.pem "${SSHO[@]}" ubuntu@"$IP" \
        'for t in $(systemctl list-units --type=timer --state=active --no-legend --plain | cut -d" " -f1); do
           sudo systemctl stop "$t"; done; sudo systemctl stop unattended-upgrades.service
         systemctl list-timers --no-pager | tail -2'
      ssh -i ~/lab-data/serving/keys/ai-research-course-fb.pem "${SSHO[@]}" ubuntu@"$IP" 'lscpu | head -12; nproc; uptime'
    

    Step 7: start the schedule, then leave. The schedule starts with setsid nohup, a command that keeps a program running after I log out. Nobody logs in again until it ends.

    A real terminal recording of step 7, with the address replaced by a placeholder. Over ssh: cd fb, then setsid nohup bash fb_schedule.sh 5 30 with its output sent to raw/schedule.log, in the background; two seconds later the log shows 02:49:23 schedule start K=5 seconds=30.

    The log's first line appeared two seconds after the start. The two numbers are the rounds and the seconds of each run. This is the command:

      ssh -i ~/lab-data/serving/keys/ai-research-course-fb.pem "${SSHO[@]}" ubuntu@"$IP" \
        "cd fb && (setsid nohup bash fb_schedule.sh $K $SEC > raw/schedule.log 2>&1 < /dev/null &); \
         sleep 2; cat raw/schedule.log"
    

    The 456 files came to 55 MB. I keep a smaller copy of 31 MB in the repo, with every time stored as a small whole number. This is the copy command:

      rsync -a -e "ssh -i $KEY ${SSHO[*]}" ubuntu@"$IP":fb/raw/ ~/lab-data/serving/fb/raw/
    

    Step 13: delete the computer, the key and the door rule, and check. A computer you forget keeps costing money. The script deletes the computer, waits until Amazon says "terminated", then deletes the key pair and the security group, and asks Amazon again.

    A real terminal recording of step 13, with the instance and group ids replaced by placeholders. terminate-instances prints shutting-down; wait instance-terminated; delete-key-pair prints true; delete-security-group prints true. Then the checks: describe-instances prints terminated; the count of key pairs named ai-research-course-fb is 0; the count of security groups of that name is 0; the count of volumes tagged Lesson=fb is 0; the count of course instances not terminated is 0. Last, the key file is removed.

    Every count printed 0 and the computer printed terminated, so nothing from this lesson was left in the account. These are the teardown commands:

      aws ec2 terminate-instances --instance-ids "$ID" --query "TerminatingInstances[0].CurrentState.Name" --output text
      aws ec2 wait instance-terminated --instance-ids "$ID"
      aws ec2 delete-key-pair --key-name ai-research-course-fb --query Return
      until aws ec2 delete-security-group --group-id "$SG" --query Return 2>/dev/null; do sleep 10; done
      aws ec2 describe-instances --instance-ids "$ID" --query "Reservations[0].Instances[0].State.Name" --output text
      aws ec2 describe-key-pairs --filters "Name=key-name,Values=ai-research-course-fb" --query "length(KeyPairs)"
      aws ec2 describe-security-groups --filters "Name=group-name,Values=ai-research-course-fb" \
        --query "length(SecurityGroups)"
      aws ec2 describe-volumes --filters "Name=tag:Lesson,Values=fb" --query "length(Volumes)"
      aws ec2 describe-instances --filters "Name=tag:Project,Values=ai-research-course" \
        "Name=instance-state-name,Values=pending,running,stopping,stopped,shutting-down" \
        --query "length(Reservations[].Instances[])"
      rm -f ~/lab-data/serving/keys/ai-research-course-fb.pem
    

    The security group can only be deleted once the computer is fully gone, so the script tries that command every ten seconds until Amazon accepts it.

    The default and "no score" cost real ranking quality. On the left, the default cost 0.0040 of AP and "no score" 0.0057, both clearly below zero. On the right, both cost about 0.16, which takes AP from 0.54 to about 0.38. A default puts every unlucky customer at the same score, so the model can no longer tell them apart. Features lesson 8, missing at serving, found the same for lookups that fail completely.

    A forest plot of the change in AP, from -0.009 to +0.001, with a dotted line at no change. Three rows. Last month's: a dot near -0.0002 with a short line crossing zero; its grey dots, the random customers, sit just to its right. Default: a dot near -0.004 with a line from about -0.0054 to -0.0027; grey dots from about -0.0053 to -0.0031. No score: a dot near -0.0057 with a line from about -0.0079 to -0.0038; grey dots from about -0.0073 to -0.0044. Caption: the waits never depended on the customer, so the slow requests behaved like random customers.

    A check against luck. I also gave stand-ins to 5 random sets of customers of the same size, to see what plain luck looks like. For the default and "no score", the machine's cost sat among those 5 random sets. For last month's score it sat just outside, below the lowest of the 5 (-0.00022 against -0.00017 to +0.00040). That is expected: the freezes were drawn without looking at the customer, so the slow requests behave like random customers.

  • "G7. H3-r2 fallback share below 0.3% and store calls per request between 1.01 and 1.05, in every round." (Two retries at 300 a second.) Right: 0.00% and 1.01 to 1.02.

  • "G8. H3-r2bj p99 above H3-r2's, told apart; both fallback shares below 0.3%." (A random pause against none.) Right.

  • "G9. H3-hedge p99 below H3-fb25's, told apart; extra calls between 4% and 10% of requests." (Asking twice against the stand-in arm.) Right: 6.3% to 6.7% extra calls.

  • "G10. L3-none p99 between 15 and 30 ms; L3-hedge p99 below 15 ms, told apart." (The long tail, waiting and asking twice.) Wrong in its second half: asking twice gave 15.9 to 16.2 ms.

  • "G11. H4-r2 (immediate retries at 450 a second) tips over in at least 3 of 5 rounds: fallback share above 20% and store calls per request above 1.5; H4-r2bj tips over in none (fallback share below 2%)." Wrong in its second half, and this is the finding of the lesson: the random pause tipped over in the same 3 rounds.

  • "G12. H4-dldown fallback share below H4-fb25's, told apart." (Deadline passed down against a plain limit, at 450 a second.) Wrong: too close to call.

  • "G13. H4-none p99 above 1,000 ms in every round." (No limit at 450 a second.) Wrong: round 4 again, 39.6 ms.

  • "G14. Every live score equals the reference to the bit; every cached and default answer equals its stored value." Right: 966,653 live answers, none different.

  • "G15. AP, H3-fb25: with CACHED the cost cannot be told apart from zero; with NONE it is measurably below zero; DEFAULT lies between them. The box's cost sits inside the spread of the 5 random controls for every stand-in." Wrong only in its last part: for last month's score, the machine's cost sat just outside the 5 random sets, below the lowest of them.

  • You saw the demo running on the lab's computer in step 11. That exact run is stored in results/fb-demo-run.txt. On the rented computer, waiting gave a slowest 1 in 100 of 44.44 ms. The 25 ms limit gave 26.06 ms with 1.58% stand-ins, retrying 28.35 ms with none, and asking twice 15.98 ms. Using the default score for 32 stand-ins changed AP on the 2,021 rows asked for from 0.6150 to 0.6113.

    Here is the same demo in VS Code on my laptop, run with python fb_demo.py in the examples folder with venv-sv active.

    A real screenshot of the VS Code terminal on my laptop after running python fb_demo.py. The first line says the laptop has 10 cores and was busy with other programs (its 1-minute load was 4.49). The model scored AP 0.5450 on 26,851 test rows, and the stand-in score was 0.2342. Requests came at 200 a second for 10 seconds, the pretend store had 16 slots, and 1 lookup in 100 froze for 1 to 2 seconds. For each way, the table lists the middle wait, the wait 99 in 100 beat, the slowest wait (all in ms), then stand-ins and lookups per request. A wait: 4.95, 45.20, 2001.3, 0.00%, 1.000. B timeout: 4.92, 26.64, 32.5, 1.63%, 1.000. C retry: 4.88, 29.82, 43.4, 0.00%, 1.016. D hedge: 5.41, 18.43, 28.3, 0.05%, 1.070. Last line: on the 2,021 rows asked for, AP was 0.6150 with every answer live and 0.6104 with B's 33 stand-ins, a change of -0.0046.

    This run is from my laptop, not the rented machine. It is a different computer: it has 10 cores, not 1, and other programs were busy on it at the same time. So do not set its numbers beside the rented machine's one by one. Compare the four rows with each other instead.

    The order is the same as on the rented machine. Asking twice was fastest for the slowest 1 in 100 (18.43 ms). Then came giving up at 25 ms (26.64 ms), then trying again (29.82 ms). Waiting with no limit was slowest (45.20 ms).

    Why only 45 ms for waiting, when the lab measured about 1,030 ms? The demo sends only about 2,000 requests. When frozen lookups are about 1 in 100, the wait that 99 in 100 beat can land just below them, as in round 4 of the lab. The slowest wait, 2,001 ms, shows that a lookup did freeze. Giving up gave 33 customers a stand-in score (1.63%), and the ranking score AP fell from 0.6150 to 0.6104.

    fb_files/fb_store.py

    fb_files/fb_service.py is the prediction service. It uses lesson 2's service file unchanged, for the model and the request format. The one change is the lookup: instead of reading the database itself, it asks the store over a network connection on the same computer. How long it waits, and what it does next, comes from one setting per run: the time limit, the stand-in, the deadline, the retries and asking twice. Each time limit is Python's asyncio.wait_for.

    fb_files/fb_client.py is the tester. It is lesson 7's, copied with credit. It draws all arrival times first. It sends each request at its planned time, even if earlier answers are late, and times each one from that planned time. It checks every live score against the computer's own single-row score, and every stand-in against its stored value.

    fb_files/fb_run_once.sh starts a fresh store and a fresh service, waits for an idle core, records the load and the system log, and runs the tester. fb_files/fb_schedule.sh runs the warm-up runs that measure top speed, and then five rounds of the 21 arms, each round in its own random order.

    fb_report.py does not use any of the above. It reads the compact copy of the raw files with its own code and works out every number again. It checks the guesses, the demo's stored run, the fact-check, the deletion and the leak scan, and it writes the playground's real numbers.