Serving And Inference

Cold Start: The First Request Waited for Python to Import Its Libraries

0 of 32 complete

0%

Contents

Back|Serving And InferenceCold Start: The First Request Waited for Python to Import Its Libraries
1/32
79 min left
  1. Home
  2. AI Engineering: Data, RAG and Agents
  3. Serving and Inference Basics
  4. Cold Start: The First Request Waited for Python to Import Its Libraries
Prerequisites
Workers and Threads: Processes, Threads, or Both on One Core?requiredTail Latency: What Sits in the Slowest 1% of RequestsrequiredLatency Anatomy: Where the Time in One Prediction Request Goesrequired
Related Topics
Model Signatures: The Right Numbers in the Wrong Shape, and What a Schema Check CatchesPackaging, Registry and VersioningMissing at Serving: No Fill Gives Back a Row That Is GoneFeatures and Feature StoresContainer Images: Most of the Size Was the Base and the Libraries, and Slimming Kept Every ScorePackaging, Registry and Versioning
1 of 32
Previous lessonWorkers and Threads: Processes, Threads, or Both on One Core?

System Design

  • Foundation
  • Intermediate
  • Advanced
  • Capstone

AI Engineering

  • Foundation
  • Data, RAG and Agents
  • Evaluation, LLM Ops and Security

systemdesign.academy

  • Home
  • Glossary
  • Interview prep
  • Reviews
  • About
  • Privacy
  • Terms

The First Customer of the Morning

Let me start at a desk in an office, early in the morning.

The desk opens at nine. The first person in the queue walks up at one minute past. The clerk has to switch on the lights and the computer, wait for it to start, open the right program and find the right drawer. Only then can she help. The second person waits only for the first one to finish.

A flat illustration of a queue of three people holding small cards at a front desk, while a clerk at the desk hands a card to an older man behind a glass window. Below the picture: when the desk has just opened, the first person in line waits for the lights, the computer and the right drawer; the people after them only wait for each other.

Nobody did anything wrong. The desk was simply not ready yet. A new copy of a prediction service has the same problem. Before it can answer anyone, it must start a program, load its tools, load the model and open its door. And the very first question it answers may still be slower than the ones after it.

In this lesson I measure that start, piece by piece, on a rented machine. I ask where the time goes, and which parts a real user would feel. Then I try four ways to hide those parts from users. I keep the files in memory, send a warm-up request, put a readiness check in front, and run the service in a container.

Where This Lesson Starts

This is lesson 6 of the chapter on serving a model. Lesson 2, latency anatomy, took one request apart. Lesson 3, tail latency, found that the slowest request of a run was always the first one, but it measured only after the service had started. Lesson 5, workers and threads, showed that one worker with one thread was the best choice on one core. So in this lesson the service always runs with one worker and one thread.

The model, the service code and the rented machine are the same as in lessons 2 to 5, so I do not explain them again. In short: the model is the features chapter's gradient boosted tree model, which scores one customer of a real online shop and reached a test AP of 0.5450. The service is a small FastAPI web service run by uvicorn. It reads the customer's six features from SQLite.

The container image follows packaging lesson 5, container images, which showed how to build a slim image and what was inside it. I do not repeat that here. This lesson asks a different question about the same kind of image: how long does it take to start?

The Words You Need First

Please read this slide slowly if any word is new. Every slide after it uses these words.

A hand-drawn grid of twelve cards, each a word with its meaning: cold start, process, import, .pyc file, unpickle, page cache, warm or cold cache, container, image, warm-up, readiness check and port open. Below: request, latency, p50, p99, open loop and microsecond (us) mean what they meant in lessons 2 to 5; 1 s = 1,000 ms = 1,000,000 us.

A cold start is everything a new copy of a service does before it can answer, plus anything extra its first request pays. A process is a running program with its own memory. Starting a new process here means starting Python again from nothing.

To import a library is to make Python read the library's files and run its set-up code. This happens once per process. A .pyc file is a module that Python has already turned into its own internal code, called bytecode. When the .pyc file exists and is up to date, the next import skips that step. To unpickle the model means to turn the bytes of the saved model file back into a model object in memory, as packaging lesson 1 showed.

The page cache is a part of memory where Linux keeps copies of files it has read recently. When a file is in it, reading the file is fast. I call that a warm cache. When it is not, Linux must read from the disk: a cold cache.

A container is a process started by from an image, which is the packed set of files it needs: Python, the libraries and the model. A warm-up is a request or a call that the service makes to itself before real users arrive. A readiness check is a question that a asks the service again and again: are you ready for traffic yet? Here it is a request to GET /ready. The port opening is the moment a network connection to the service first succeeds.

is the time from a request to its answer. The p99 is the 99th percentile: the time that 99 of every 100 requests stay under.

The Headline: Most of the Start Was Python Importing Libraries

Here is the main result first. Each lane is one way of starting the service, from the moment it was started to the moment the first answer came back. Each lane is the median of eight starts.

Six horizontal lanes, one per way of starting, each split into coloured parts: start Python, imports, model SQLite and server, warm-up and /ready, first request. A, bare with a warm cache, ends at 1,499 ms; C50, 50 HTTP warm-ups, at 1,571 ms; D2, container with a compiled standard library, at 1,845 ms; D, container, at 2,043 ms; B, bare with the cache emptied, at 4,215 ms; E, container with the cache emptied, at 5,229 ms. In every lane the imports part is by far the longest. Below: in A the imports took 1,449 ms, 96.9% of the time until the port opened; reading and unpickling the model took 2.0 ms; with the page cache emptied the first answer came 2.81 times as late; the container image, whose Python is a different build, added 539 ms (container and build not separated); precompiled files gave back 200 ms; all told apart. Caption: the model itself loaded in 2.0 ms; the libraries around it took a second and a half.

A fresh service answered for the first time after 1,499 milliseconds. That was a bare Python process on the box, with all its files already in memory.

Almost all of that was Python importing libraries. The imports took 1,449 ms, which is 96.9% of the time until the port opened. scikit-learn alone took 824 ms. Reading the model file and unpickling it took 2.0 ms.

Files that were not in memory made the start 2.81 times as long: 4,215 ms. The container image, whose Python is a different build, added 539 ms (container and build not separated). A container whose image held compiled copies of Python's standard library gave 200 ms of that back.

The first request itself was slower than the second: 4,680 microseconds against 1,849. One warm-up request, sent by the service to itself before it said "ready", brought the first real request down to 1,655 microseconds.

But under 539 random arrivals a second, I could not tell a warmed service from a cold one: the run-to-run spread was larger than any difference between them.

How the Lab Was Built

I wrote the lab's design into the docstring of scripts/labs/serving/cold_start.py on 2026-10-04, before any timed run. It names every way of starting, every comparison, the rule for "told apart from noise", and eighteen guesses.

A page in five labelled zones, titled ten ways to start, 8 rounds. Bare process: A, files already in memory; B, the same with the page cache emptied first (drop_caches). Warm-ups: C1 and C50, 1 or 50 requests to itself, /ready says 503 until done; I50, 50 calls inside the process before the port opens. Containers: D the image; D2 the same image with the standard library compiled; E as D with the page cache emptied. Under traffic: F no warm-up; G 50 warm-ups and the /ready gate; both get 539 random arrivals a second for 3 s. The rule: told apart only if all 8 paired differences agree in sign and the two sets do not overlap. Caption: 18 guesses written first; 14 came out right.

Every run starts a brand-new copy of the service. There are ten ways to start it, which I call arms. In each of eight rounds, every arm runs once, in a new random order.

A is the plain case: a bare Python process, with its files already in the page cache. B is the same, but I empty the page cache just before the start. C1 and C50 send 1 or 50 requests to themselves after the port opens, and answer /ready with "503, not yet" until the warm-up is done.

I50 makes 50 lookups and model calls inside the process before it opens the port, with no warm-up. D runs the service in a container. D2 uses an image with the standard library compiled. Its compileall over /usr/local/lib/python3.13 also compiled 84 site-packages files that had none. They are most likely pip's own, which the service does not import; I did not check which. E is D with the page cache emptied.

and test the start under traffic. F has no warm-up, and arrivals begin when the port opens. G has 50 warm-ups and waits for to say 200.

The Machine

Like lessons 2 to 5, every timing comes from a small machine I rented from Amazon Web Services. None comes from my laptop, which is always busy with other work.

Five rows, each with a logo. EC2 c7g.medium, us-east-1: 1 vCPU (Neoverse-V1), 2 GiB, not burstable; client and service share the one core; on for 0.444 hours, $0.0161. Ubuntu 24.04, arm64: core at least 99.5% idle before every one of 110 runs; host steal at most 0.00%; timers stopped. Python 3.13.15, scikit-learn 1.9.1: numpy 2.5.3, pandas 3.0.6; the bare service in a venv; the client in the system's own Python, standard library only. FastAPI 0.142.2, uvicorn 0.54.0: lesson 2's handler and model file, with a clock reading at every step of the start. Docker 29.1.3, python:3.13.15-slim: the same pinned packages inside; standard library .pyc files 163 of 633 in the image, all in its compiled twin. Caption: every timing I quote comes from this box, never from my laptop.

The box is an EC2 c7g.medium, with one virtual CPU and 2 GiB of memory. My AWS account allows only one virtual CPU at a time, as lesson 5 explained. So the client that starts the service and sends requests runs on the same core as the service.

The client runs in a different Python on purpose. In the cold-cache arms, I empty the page cache just before the start. If the client imported numpy, it would read the very files the service is about to read, and the service would start warm. So the client uses only Ubuntu's own Python and its standard library. It shares no files with the service's Python, except the basic C library.

Before every one of the 110 runs, the core was at least 99.5% idle. The machine under my box took no measurable time from it.

The Whole Lab in One Picture

Before the recordings, here is the whole lab as one numbered sequence. The numbers on the arrows match the numbers on the recordings that follow.

A sequence diagram with five lifelines: my laptop, AWS, the box, client and service, and fourteen numbered arrows. 1, my laptop asks AWS whether any box is running. 2, key, group, SSH rule. 3, image, run-instances. 4, my laptop to the box: wait, then SSH. 5, Python, Docker, files. 6, build two images. 7, stop timers, look. 8, start, log out. 9, my laptop to AWS: read the console. 10, client to service: start, 100 requests. 11, service to client: answers, clock marks. 12, my laptop to the box: the demo. 13, the box to my laptop: raw files. 14, my laptop to AWS: terminate, delete. Below: arrows 10 and 11 repeat for each of 80 timed starts, 8 rounds of 10 ways, with nobody connected; step 9 reads the box's serial console through AWS and never touches the box. Caption: steps 1 to 9 and 12 to 14 are the recordings that follow, numbered the same; all on the box that measured the numbers.

Steps 1 to 8 happen once: check, rent the box, set it up, build the images and start the schedule. Step 9 is how I watched the schedule without touching the box. Arrows 10 and 11 are the timed work, repeated for 80 starts with nobody logged in. Steps 12 to 14 come after the schedule: the student demo, the copy of the results back to my laptop, and the deletion of everything.

Three things are not in this picture, and I want to be open about them. Before the schedule, I ran one untimed start of each arm over SSH, only to check that every arm worked. That was the first real test of the Dockerfiles. After the demo, I copied one more script to the box with scp. Then I ran one more block of timings over SSH, without a recording, to repeat a measurement that had gone wrong. The slide about images explains it. And the fetch in step 13 ran twice: the first rsync, right after the demo, is not recorded either; the recording is the second one.

Build the Lab Box Yourself

The rented box is part of the lab, so I show every step of it. You can rent the same kind of box and run the same schedule.

Which box you see. Every recording comes from the box that measured this lesson's numbers. No recording happened during a timed run, and nobody logged in while the schedule ran.

What you need first. An AWS account, the AWS command line tool (aws) set up with your keys, and ssh, scp, rsync and curl. All the steps live in one file, scripts/labs/serving/cold_files/cold_aws_steps.sh. You run it from the scripts/labs/serving folder, one step at a time: bash cold_files/cold_aws_steps.sh check, then access, and so on. Every code block below is copied word for word from that file.

The blocks use a few names the file sets at its top. Run these lines first, from the scripts/labs/serving folder, if you paste the blocks by hand. All of them are copied from the file, except HERE=$PWD, the folder B and the two short names S and C. The file works out from its own location, which a pasted line cannot do. It sets , and inside its step. and use the box's address, , so run those two lines only after step 4 has set it.

Steps 1 to 3: Check, Lock the Door, Rent the Box

Step 1: is any other course box running? My account allows one box of this kind at a time, and another lesson may be using it. So the first command counts the course's boxes that are not yet deleted. It must print 0.

A real terminal recording of step 1. The shell prints the command aws ec2 describe-instances with filters for the tag Project=ai-research-course and the states pending, running, stopping, stopped and shutting-down, asking for the number of instances; the answer is 0.

The recording shows the command and its answer, 0, so no other course box was alive. This is the command, copied from the step script:

  aws ec2 describe-instances --filters "Name=tag:Project,Values=ai-research-course" \
    "Name=instance-state-name,Values=pending,running,stopping,stopped,shutting-down" \
    --query "length(Reservations[].Instances[])"

Step 2: a key and a locked door. A key pair is how you prove to the box that you are allowed in. A security group is a door rule for the box. Mine opens only port 22, the SSH port, and only to my own address. In the rule, /32 after an address means that one address and no other. Both get tags, labels that say which project and lesson they belong to.

A real terminal recording of step 2, with the address, network and group ids replaced by placeholders. create-key-pair for ai-research-course-cold with the tags Project and Lesson writes the key to the key file without printing it; chmod 600 on the key file; curl reads my address; describe-vpcs finds the default network; create-security-group makes the group ai-research-course-cold with its tags; authorize-security-group-ingress prints a small table: from my address /32 (that one address only), port 22. Last line: the key file, readable only by its owner, 388 bytes.

The key material went straight into the file and was never printed, and the door rule names only my own address. These are the commands:

  aws ec2 create-key-pair --key-name ai-research-course-cold --key-type ed25519 \
    --tag-specifications "ResourceType=key-pair,Tags=[{Key=Project,Value=ai-research-course},{Key=Lesson,Value=cold}]" \
    --query KeyMaterial --output text > ~/lab-data/serving/keys/ai-research-course-cold.pem
  chmod 600 ~/lab-data/serving/keys/ai-research-course-cold.pem
  MYIP=$(curl -sf https://checkip.amazonaws.com)
  VPC=$(aws ec2 describe-vpcs --filters Name=isDefault,Values=true --query "Vpcs[0].VpcId" --output text)
  SG=$(aws ec2 create-security-group --group-name ai-research-course-cold --vpc-id "$VPC" \
    --description "lesson 6 cold start, ssh from one address" \
    --tag-specifications "ResourceType=security-group,Tags=[{Key=Project,Value=ai-research-course},{Key=Lesson,Value=cold}]" \
    --query GroupId --output text)
  aws ec2 authorize-security-group-ingress --group-id "$SG" --protocol tcp --port 22 --cidr "$MYIP/32" \
    --query "SecurityGroupRules[].{port:FromPort,from:CidrIpv4}" --output table

Steps 4 to 8: Wait, Set Up, Build the Images, Start

Step 4: wait until it runs. A new box takes a short while to start. The AWS tool waits for you. Then the script reads the box's public address into IP, and tries SSH every five seconds until the box answers.

A real terminal recording of step 4, with the instance id and address replaced by placeholders. aws ec2 wait instance-running; describe-instances reads the public address into IP; describe-instances prints running, c7g.medium, us-east-1a; ssh with ConnectTimeout=5 runs true; a second ssh prints ssh works: aarch64 Ubuntu 24.04.5 LTS.

This time SSH answered on the first try. These are the commands, with the retry loop:

  aws ec2 wait instance-running --instance-ids "$ID"
  IP=$(aws ec2 describe-instances --instance-ids "$ID" \
    --query "Reservations[0].Instances[0].PublicIpAddress" --output text)
  until ssh -i ~/lab-data/serving/keys/ai-research-course-cold.pem "${SSHO[@]}" -o ConnectTimeout=5 \
    ubuntu@"$IP" true 2>/dev/null; do sleep 5; done

Step 5: Python, and the lab files. The box gets Python 3.13.15 and the same pinned library versions as lessons 2 to 5, through uv, a fast installer for Python. It also gets Docker from Ubuntu's own packages. Then scp sends lesson 2's model, feature table and request list, this lesson's files, and what the demo needs. The last command builds the SQLite table, computes a reference score for every request row, and prints the files' sha256 fingerprints.

A real terminal recording of step 5, with the address replaced by a placeholder and my home folder shortened to a tilde. Over ssh: make the folders; install uv, Python 3.13.15 and a venv; install the pinned packages; apt-get installs docker.io and docker-buildx and prints Docker version 29.1.3. Then scp copies lesson 2's model and data, this lesson's cold_files with both Dockerfiles, the features chapter's files and data, and the demo. Last: build_store.py prints 26851 rows; cold_ref.py writes 26851 single-row scores, the first 0.10525502546736738; the sha256 of model.pkl starts bec4d11e, of store.parquet 62aadc30, of requests.parquet 407d62e6.

The three fingerprints match the ones lesson 2 recorded, so the box serves the same model. These are the commands that install Python, the packages and Docker:

Steps 9 and 12 to 14: Watch, Demo, Bring It Home, Delete

Step 9: watch without touching. The schedule writes one progress line to the box's serial console after every run. The serial console is a log that AWS keeps for the box, and you can read it through AWS without logging in.

A real terminal recording of step 9, with the instance id replaced by a placeholder. aws ec2 get-console-output with --latest, piped through grep COLD-PROGRESS and tail -12, prints progress lines with times: schedule start K=8 at 02:12:33; load before the schedule 0.10 0.38 0.23 after waiting 120s at 02:14:33; then r1-D2, r1-F, r1-G, r1-E, r1-D and r1-C50 done between 02:14:41 and 02:15:23.

The schedule waited 120 seconds for the box to calm down, then started round 1. Each start took about eight seconds, warm-up start included. This is the command:

  aws ec2 get-console-output --instance-id "$ID" --latest --output text | grep COLD-PROGRESS | tail -12

Step 12: the student demo, on the same box. After the schedule, I ran the demo you will meet later, so you can see what it prints on one core.

A real terminal recording of step 12, with the address replaced by a placeholder: python cold_demo.py on the box, its output also written to cold-demo-run.txt. It prints the machine, 1 cores, load 0.46, and test AP 0.5450. Part 1, a fresh Python process step by step, median of 5, ms: python ready 14.5, import numpy 61.0, import pandas 216.6, import sklearn 721.3, read model file 0.2, unpickle 0.8, all of it 1014.0; predict_proba on one row, call 1 3004.3 us, call 2 618.1, call 100 485.3. Part 2, -X importtime, top-level imports in ms: sklearn.ensemble 771.8, pandas 241.7, numpy 58.5, site 7.0, json 4.7. Part 3, the first real predict_proba call: no warm-up 3004.3 us, after one warm-up call 622.6 us.

This run happened after the last timed run of the schedule, so it could not disturb any timing. The same output is stored in results/cold-demo-run.txt. This is the command:

  ssh -i ~/lab-data/serving/keys/ai-research-course-cold.pem "${SSHO[@]}" ubuntu@"$IP" \
    'cd cold/serving/examples && ../../venv/bin/python cold_demo.py --save ~/cold/raw/cold-demo.json \
     | tee ~/cold/raw/cold-demo-run.txt'

Step 13: bring the results home first. The rule from lesson 3 is: the moment the runs finish, copy the raw files back, before anything else.

The Lab's Report, Running

This is a real recording of the report script, cold_report.py. It ran on my laptop, but it times nothing. It reads the raw files the box measured and does all the arithmetic again with its own code.

A terminal recording of cold_report.py in nine numbered sections. 1: the box, c7g.medium, 1 vCPU, measuring box 0.444 h and setup box 0.068 h at $0.0363 an hour = $0.0186; both boxes, keys, groups and disks gone; 28 stored files scanned with no ids, addresses or keys; 110 runs, the core at least 99.5% idle before each; cron ran during imp-warm-r2, r1-E and r2-A. 2: a table of the twelve start steps for A and B with B/A ratios, from python ready 18.2 against 139.5 ms to sklearn 824.3 against 2081.7 ms, and imports 0.969 of launch to port open. 3: launch to first answer for 8 ways, from A 1498.8 ms to E 5228.8 ms, and five comparisons, all told apart. 4: request 1, 2 and 100 per way with predict_proba inside, and the warm-up comparisons. 5: open loop, F and G per second, 0 of 6 told apart. 6: -X importtime warm 1597.8 ms and cold 4271.4 ms; image load 8.66 s, pull 2.28 s, and the schedule's own unclean runs 7.40 and 0.52 s; 32,379 of 32,379 scores equal; clock checks. 7: 14 guesses right, 4 wrong. 8 and 9: the demo and the playground match. Last line: all 469 checks agree with the stored lab.

The report reads the small copy of the raw files kept in the repo, about 1.8 MB, gzipped. It computes its own steps, medians, percentiles, paired differences and verdicts, and compares every one with the lab's results file. If any number disagreed, it would stop with an error. I checked that too: I changed one number in a copy of the results file, and the report printed "2 MISMATCHES" and stopped.

The system log had 452 lines during the 110 runs. Most came from and its helper containerd starting and stopping containers, from sudo emptying the page cache, and from systemd. Nine came from cron, which still ran because I stopped only the systemd timers. They fell in three runs, and each landed inside its arm's range. r2-A answered first at 1,484.9 ms, and A's range was 1,484.9 to 1,523.2. r1-E answered at 5,236.8 ms, inside E's 5,184.4 to 5,257.0. imp-warm-r2 took 1,597.8 ms, inside the warm runs' 1,585.8 to 1,620.7.

One Clock for the Client and the Service

To split one start into steps, I need to subtract a time the client read from a time the service read. That only works if both read the same clock.

A hand-drawn timeline from launch to port open with five marks: launch, Python ready, imports done, model loaded and port open; the client reads the first and the last, and the service reads the ones between, bare or in Docker. Below: the client reads the clock just before it starts the service, and again when a TCP connect first succeeds; the service reads the same clock at each step inside its own process, and a container shares the host's clock; in all 80 timed starts every service reading fell after the launch, and every reading up to the server's set-up fell before the port opened; in 2 runs the connect succeeded up to 1.4 ms before the service noted its socket was up, because that note is taken just after the server's start-up code returns. Caption: one clock is what lets me subtract a client reading from a service reading.

Every program here reads time.monotonic_ns(), a monotonic clock: one that only ever moves forward, so it is safe for measuring time between two readings. Python's documentation lists how it is built on each system, and on Linux it ends with "Otherwise, call clock_gettime(CLOCK_MONOTONIC)." That is one clock for the whole machine, and a container shares it unless you ask for a separate one.

I checked it on the box, not on trust. On my laptop, readings from two different Python versions did not agree, so the design asked for a check. In all 80 timed starts, every service reading came after the client's launch reading. Every reading up to the server's set-up came before the port opened.

One detail surprised me. My mark for "socket up" is read just after uvicorn's start-up code returns. The socket starts listening inside that code, so in 2 of 80 runs a connection succeeded slightly before my mark, by at most 1.4 ms. The clocks agreed; my mark was simply a little late. I changed the check to stop at the server's set-up, and I say so in the lab's notes.

Where One Start's 1.5 Seconds Went

Now arm A, the plain case, step by step. Each block is one step of the start. Its height is the median time over eight runs.

An isometric row of nine blocks, one per step of a start in arm A, each as tall as its median time in ms: start Python 18, stdlib 27, numpy 61, pandas 251, sklearn 824, fastapi 266, uvicorn 21, model + SQLite 2, server up 25. The sklearn block towers over the rest; the model block is almost flat. Below: the imports, from the standard library to uvicorn, took 1,449 ms; scikit-learn's alone 824 ms, the largest step in all 8 runs; reading the model file and unpickling it took 2.0 ms; 1,820 modules were loaded by the end of the imports. Caption: the model was the smallest block on the page.

Starting Python took 18 ms. That is from the client's launch to the first line of my service file.

Then came the imports, 1,449 ms in all. numpy took 61 ms, pandas 251, scikit-learn's ensemble module 824, FastAPI and pydantic 266, and uvicorn 21. By the end, 1,820 Python modules were loaded. I import pandas on its own line before scikit-learn only to show its time separately; scikit-learn would import it anyway.

Loading the model took 2.0 ms: 0.8 ms to read the file's bytes and 1.2 ms to unpickle them. The model is small, about 200 KB, and its class was already imported. Opening SQLite took 0.2 ms, and starting the server and its socket, the network end point that clients connect to, took about 25 ms.

So the model is not the slow part of starting this service. The libraries around it are. A bigger model would change the picture, and I come back to that in the limits.

Which Library Cost What

To check the import times another way, I asked Python itself. Python's -X importtime option prints the time of every import, with "cumulative time (including nested imports) and self time (excluding nested imports)".

Four rows, each with a real logo. sklearn.ensemble: 769 ms warm; 2,009 ms with the page cache emptied, 2.61 times as long. pandas: 235 ms warm; 1,013 ms cold, 4.32 times. fastapi: 241 ms warm; 456 ms cold, 1.90 times. numpy: 57 ms warm; 227 ms cold, 3.98 times. Below: Python's -X importtime option prints the time of every import, with the time of the imports inside it; the whole process took 1,598 ms warm and 4,271 ms cold; its number includes the libraries it imports in turn, such as scipy; numpy and pandas were already loaded by then. Caption: the largest top-level import was sklearn.ensemble in all 5 warm runs.

I ran it ten times: five with the files in memory and five with the page cache emptied, in a shuffled order. The order of imports matched the service: numpy, then pandas, then scikit-learn, then FastAPI and uvicorn.

scikit-learn's ensemble module was the largest in every run: 769 ms warm. It imports scipy and many of its own parts inside it, which all count in its number. pandas took 235 ms and FastAPI 241. These agree with my own clock marks to within about ten percent, so the two ways of measuring tell the same story.

With the page cache emptied, every library took longer, from 1.90 times for FastAPI to 4.32 times for pandas. The whole process took 4,271 ms instead of 1,598.

An Empty Page Cache Made Almost Every Step Slower

Arm B starts the same bare service, but just before the start I empty the page cache: sync; echo 3 > /proc/sys/vm/drop_caches. The kernel's documentation says writing to this file "will cause the kernel to drop clean caches". Its advice is to run sync first, so that more pages are clean and can be dropped.

A dot chart on a log scale of milliseconds against nine steps of the start: Python, stdlib, numpy, pandas, sklearn, fastapi, uvicorn, unpickle and SQLite. For each step one dot for A, warm cache, and one for B, cache emptied. B is above A at every step except unpickle, where the two dots sit on top of each other near 1 ms. Below: of these 9 steps, B was slower at 8, told apart; unpickle could not be told apart, since unpickling reads no file; Python itself started 7.69 times as late; the largest absolute cost was scikit-learn, +1,255 ms; emptying the cache with drop_caches took 72 ms, before the clock started. Caption: first answer, A 1,499 ms, B 4,215 ms.

The whole start took 4,215 ms instead of 1,499: 2.81 times as long, told apart. Every step that reads files from disk got slower. Python itself took 140 ms to start instead of 18, because its own program and standard library had to come from the disk. scikit-learn's import cost the most extra time: 1,255 ms more.

Unpickling did not change: 1.2 ms in both. It reads no file. The bytes were already read in the step before, which did get slower, from 0.8 ms to 2.1.

What does a cold cache mean in real life? It is a machine whose memory does not hold these files: just after a reboot, or after other work has pushed them out. It is not a brand-new disk made from a snapshot. AWS fetches such a disk's blocks from storage on first read, which can be much slower. I did not measure that case.

The first request was slower in B too: 9,432 us against 4,680 (not a declared comparison; B's lowest, 9,115 us, was above A's highest, 5,770). Its own request path had files left to read.

What the Container Added

Now arms D and D2. The service, the model and the packages are the same, but the service runs in a container: docker run -d --rm --network host. The option --network host lets the container use the box's own network, so the client connects the same way.

A bar chart of milliseconds for nine steps of the start, start, stdlib, numpy, pandas, sklearn, fastapi, uvicorn, SQLite and socket, with three bars per step: A bare, D container and D2 compiled. The start bar is near 18 for A and about 200 for D and D2. The sklearn bars are about 824, 1,040 and 889. Below: starting the container, until Python ran its first line, took 202 ms against 18 for a bare process; docker run itself returned at 199 ms; the imports took 1,797 ms in D and 1,589 in D2; the image held 163 compiled standard library files of 633; D2 held all of them, and its compileall also compiled 84 site-packages files that had none; Python is a different build in each, a different compiler, so D against A is not containers alone. Caption: first answer, A 1,499, D 2,043, D2 1,845 ms; D against A and D2 against D told apart.

The container's first answer came 539 ms later than the bare process's, told apart. The container image's Python is a different build, so this is the container and the build together, not separated. Two parts make up most of it.

First, starting the container took 202 ms before Python ran a line, against 18 ms for a bare process. has to set up the container's files, its limits and its process. The docker run command itself returned at 199 ms.

Second, the imports took 348 ms longer: 1,797 ms against 1,449. Here a detail of the image matters. A container started with --rm gets a fresh, thin writable layer every time, and the layer is thrown away when it stops. Docker's documentation says: "When you create a new container, you add a new writable layer on top of the underlying layers." So a .pyc file that Python writes during one start is gone at the next.

Most of the standard library had no .pyc files in the image. The box counted 163 compiled files out of 633 standard library modules. The official slim image ships none; pip compiled the ones it used itself during the build.

Getting the Image onto the Machine

A container cannot start until its image is on the machine. When a new machine joins, it must first get the image, from a file or from a registry, which is a server that stores images.

The schedule measured both: docker load of my image from a file, and docker pull of the base image from AWS's public copy of 's official images. Each one started from a machine with no images, and an emptied page cache.

The first try was not what the design asked for, and I found it only by looking at the numbers. Each pull took about 0.5 seconds.

That made me check, and containerd's store still held the base image's layers. containerd is the program under Docker that keeps images and runs containers; Docker 29's release notes say "containerd image store is now the default for fresh installs". The build in step 6 had left the base image's layers in that store. Docker's build cache, which keeps pieces of earlier builds to make later builds faster, was holding them. Removing the images removed their names, not their layers, which were still on disk. So "pull" found everything locally, and "load" found most of it too.

So I repeated both, on the same box, after the demo. This block was added after the results, and it is not in the recordings: I started it over SSH. Its script, cold_files/cold_image_again.sh, also empties the build cache. Blobs are the stored files that layers and image descriptions are kept in. Before each timed run the script checks that containerd's store holds 0 blobs, and all ten runs started from 0.

Loading my 209.5 MB image file took 8.66 seconds (8.64 to 8.92 over five runs). Pulling the base image, 43.3 MB as Docker reports it, took 2.28 seconds (2.25 to 2.40). The first, unclean try gave 7.40 and 0.52 seconds; I keep those numbers in the results, labelled, and I do not use them.

So on a new machine, getting the image can cost several times the start itself. A full pull of my own image from a registry far away would cost more. I did not measure that, because I have no registry for my image, and making one would be a new resource.

The First Request Paid About Three Milliseconds More

Back to the service once it has started. In arm A, the client sends 100 real requests, one at a time, as soon as the port opens. How does the first compare with the rest?

Three hand-drawn bars of time from send to answer, median of 8 runs in arm A: request 1 4,680 us, request 2 1,849 us, request 100 1,190 us. The lower, darker part of each bar is predict_proba inside the handler: 2,946 us on the first request and 725 us on the second. Below: request 1 was slower than request 2 in every run, told apart; the median of the eight paired differences was 2,981 us, a median of differences, so not 4,680 minus 1,849; request 1 ranged 4,186 to 5,770 us and request 2 1,180 to 2,620; request 2 was slower than request 100 in all eight runs, but the two sets overlapped, so they could not be told apart. Caption: most of the extra was the model's first call, not the network.

The first request took 4,680 microseconds, and the second 1,849 (medians of eight runs). Request 1 was slower in all eight runs. The median of the eight paired differences was 2,981 us (a median of differences, so not 4,680 minus 1,849). Request 1 ranged 4,186 to 5,770 us and request 2 ranged 1,180 to 2,620. The median of the eight per-run ratios was 2.84. Inside the handler, the model's own call took 2,946 us the first time and 725 us the second. So most of the extra was the model's first predict_proba call. The rest was the first trip through the web framework's code.

Why is the model's first call slow? I did not measure inside it. One possible reason is that the first call prepares things the later calls reuse. Examples are code that the processor reads for the first time, and helper objects that scikit-learn builds once. Lesson 3 saw the same first-call effect in its maximum .

Request 2 was slower than request 100 in all eight runs, 1,849 against 1,190 us. But the two sets of eight overlapped, so by the rule they could not be told apart. My guess had been that the first request would cost at least three times the second. It cost 2.84 times, so that guess was wrong, though close.

One Warm-Up Removed Most of the First Request's Extra

If the first call costs extra, let someone else make it. A warm-up sends the service a request, or a model call, before real users arrive.

Three panels of the first real request's time, median of 8 runs. No warm-up: 4,680 us; A, the port opening is the signal; the first real request is the service's first. One warm-up: 1,655 us; C1, one request to itself, then /ready says 200; the warm-up itself took 4,912 us. 50 calls inside: 2,056 us; I50, 50 lookups and predicts before the port opens; no HTTP warm-up. Below: C50, 50 HTTP warm-ups, 1,195 us; C1 and C50 below A, told apart; C50 against C1 could not be told apart, median -422 us; I50 against C50 could not be told apart either. Caption: someone always pays for the first call; with a warm-up, it is the service itself.

With one warm-up request (C1), the first real request took 1,655 us instead of 4,680, told apart. The warm-up request itself took 4,912 us, about what the first request took in A. So the extra cost did not disappear. The service paid it, before any user was let in.

Fifty warm-up requests (C50) gave 1,195 us. Against one warm-up, the difference could not be told apart. So one request did most of the work here.

Fifty calls inside the process (I50) gave 2,056 us. These warmed the model and SQLite, but not the web framework's own path. My guess was that I50 would land between A and C50, told apart from both. It was below A, told apart, but against C50 it could not be told apart. So that guess was wrong. The path's own first-time cost was too small to show against the noise here.

The Readiness Gate

A warm-up only helps if no user arrives before it ends. Something must hold traffic back. That is the readiness check.

A hand-drawn timeline: port open, then a long box of /ready answering 503, 503 and so on while 50 warm-up requests run, then 200, then traffic. Below: in C50 the port opened at 1,491 ms and /ready first answered 200 at 1,567 ms; the client asked /ready every 5 ms, 11 to 12 times per run; until then, a load balancer that only checks the TCP port would already have sent users in; the first real request then went out at 1,569 ms, 69.5 ms later than A's, told apart. Caption: a port that is open is not a service that is ready.

In arms C1 and C50, my service answers GET /ready with status 503, "not available", until its warm-up is done, and with 200 after. The client plays the : it asks every 5 ms and sends real requests only after a 200.

, a common system for running containers, describes the same idea. Its documentation says: "Readiness probes determine when a container is ready to accept traffic." It also says this is useful when an application must do "time-consuming initial tasks, such as establishing network connections, loading files, and warming caches."

The important part is what the check asks. Some load balancers check only that a connection succeeds; for AWS's Network Load Balancer, its documentation says "The default is the TCP protocol." In C50 the port opened 76 ms before the service was warm. With one warm-up (C1) the gap was about 16 ms (ready 1,506.3 minus port 1,490.5, a difference of medians). A check that only opens a connection would have sent users in during those milliseconds, straight into the slow first call.

What the Warm-Up Cost

A warm-up is not free. It makes the start a little longer, and it puts the cost of the first call on the service instead of on a user.

A two-column ledger titled spent before traffic and saved on the first request. C50: the first real request went out 69.5 ms later than in A; it saved 3,202 us: its first real request took 0.26 of A's time. I50: the first real request went out 50.1 ms later than in A; it saved 2,594 us: its first real request took 0.44 of A's time. A cold start already took 1,499 ms before any warm-up; so a warm-up adds a little to a start that is already long, and it moves the first call's cost off a user.

Fifty warm-ups delayed the first real request by 69.5 ms, compared with A, told apart. Fifty in-process calls delayed it by 50.1 ms, also told apart. Against a start of 1,499 ms, both are small.

What did they buy? About 3 milliseconds on one request. That sounds like very little, and on its own it is. But it is the request that a user waits for after a deploy, a restart or a scale-up. And if a service starts many times a day, for example when it scales up and down with traffic, every start has a first user.

I would still use one warm-up request behind a readiness check. It is cheap, it covers the whole request path, and here it removed most of the extra.

Under Steady Traffic, the Warm-Up Could Not Be Told Apart

So far one request came at a time. Real users arrive at random, many at once. Arms F and G start the service and then send open-loop traffic: 539 random arrivals a second for 3 seconds. That is half of lesson 5's measured capacity S, 1,077.1 a second. counts from the moment each request was meant to arrive.

A dot chart of p99 latency in ms, from 0 to 25, for the 1st, 2nd and 3rd second of traffic; each dot is one run of 8; F, no warm-up, and G, 50 warm-ups and /ready, side by side. In the first second both spread widely, from about 5 to 23 ms; in the second and third seconds both sit near 4 to 9 ms. Below: median p99 in the first second, F 6.51 ms, G 8.19 ms; in the third, F 5.37, G 5.01; compared round by round, 0 of 6 comparisons, p50, p99 and max in the first and the third second, were told apart; each run held 1,552 to 1,688 requests in its 3 seconds, so one slow first call is one of hundreds in the first second alone. Caption: at 539 random arrivals a second, I could not tell a warmed service from a cold one: the run-to-run spread was larger than any difference between them.

None of the six comparisons was told apart. In the first second, F's median p99 was 6.51 ms and G's 8.19 ms. In the third second they were 5.37 and 5.01. One run in each arm reached 21 to 23 ms; the next highest were 15.4 (F) and 14.3 (G).

I had guessed that G's first-second maximum would be lower than F's, and that F's first second would be at least three times worse than its third. Both guesses were wrong. F's first-second p99 was only 1.40 times its third.

Why? One possible reason is size. At 539 arrivals a second, about 540 requests land in the first second. The first call's extra 3 ms touches only the few requests queued behind it. The random bunching of arrivals, which made the queue grow and shrink, moved the p99 by more than that. This does not mean warm-ups never matter under traffic. A bigger model, a slower first call or more traffic at the start could change it. Here, at this rate, I could not see it.

My Eighteen Guesses Before the Run, Checked

I wrote eighteen guesses into the lab before it ran. Each is quoted here word for word from the docstring. Fourteen were right and four were wrong.

  1. "H1 A: launch -> python_ready between 15 and 80 ms." Right: 18.2 ms.

  2. "H2 A: the sklearn import step (after numpy and pandas) is the largest single step of the start." Right, in all 8 runs.

  3. "H3 A: the import steps (stdlib to uvicorn) are at least 70% of launch -> port open." Right: 96.9%.

  4. "H4 A: unpickle under 20 ms." Right: 1.22 ms.

  5. "H5 A: launch -> first answer between 1.0 and 4.0 s." Right: 1.499 s.

  6. "H6 A: request 1 at least 3 x request 2 (median of the runs), and request 100 within 1.2 x request 2." Wrong: request 1 was 2.84 times request 2. Request 100 was 0.73 times request 2.

  7. "H7 B: launch -> first answer at least 2 x A's, told apart." Right: 2.81 times.

  8. "H8 C1: the first real request within 1.5 x A's request 2 (one warm-up removes most of the first-request cost)." Right: 1.04 times.

  9. "H9 C50 vs C1: the first real request cannot be told apart." Right.

  10. "H10 I50: the first real request below A's, told apart, and above C50's, told apart (the path is still cold)." Wrong: below A's, told apart, but against C50's it could not be told apart.

Try It Yourself

The full lab needs a rented box and . The demo, cold_demo.py, runs on your own computer and needs neither. It times a fresh Python process, step by step, the same way the lab does.

A page in four labelled zones, headed cold_demo.py, designed before it ran. Train: the features chapter's model; it must score test AP 0.5450; saved with pickle. Five fresh processes: each times starting Python, importing numpy, pandas and sklearn, reading and unpickling the model, then 100 one-row calls. -X importtime: the five biggest top-level imports, once. A warm-up: five more processes with one call first; the first real call with and without it. Caption: on the box, 1 core, start to model loaded 1,014 ms; call 1 3,004 us, after one warm-up 623 us.

I wrote the demo's design into its docstring after the lab's design and before the demo first ran.

A real screenshot of VS Code with cold_demo.py open at the top of the file, showing its docstring: what it needs, how to run it, what it writes, and the design written before it first ran.

Before you run this lab. The demo uses lesson 2's Python environment, with the same library versions as the lab's box. If you made it for lessons 2 to 5, use it again. If not, make it once, inside the scripts/labs/serving/examples folder:

python3 -m venv venv-sv
source venv-sv/bin/activate        # on Windows: venv-sv\Scripts\activate
pip install -r lat_requirements.txt

I used Python 3.13. Then run python ../../features/fetch_data.py once, which downloads the shop data. Now run python cold_demo.py. It needs no Docker, no web server and no cloud account. It writes a model file into a temporary folder and deletes the folder at the end.

The first line it prints says the times come from your machine. Your disk, your processor and whatever else is running all move them. The demo cannot empty your page cache without administrator rights, so every process after the first starts with its files already in memory.

Pick a Way to Start

This box holds the real results from the lab. It needs nothing but Python, so it runs in your browser. It has no model and no clock: it only looks up what the box measured.

Press Run. It shows arm A, the plain case, and then the arm you pick, so you can always compare. Try START = "B" for an emptied page cache, "D2" for the compiled container, or "C1" for one warm-up. The differences it prints are one median minus another, so they can differ a little from the paired medians on the slides.

The report script writes this box from the lab's raw files, runs it for all eight ways, and checks that each run prints the lab's numbers.

The Lab's Code, Piece by Piece

The lab has one file that runs on my laptop, a few small files that run on the box, and the step script you saw above.

cold_start.py, on my laptop, holds the design in its docstring, written before any run, with dated notes added after. Its commands are check (lesson 2's files match by sha256), status (reads the box's serial console), collect, cost and factcheck. collect calls cold_stats.py, which gzips the box's raw files into results/cold-raw/, passes every log line through cold_redact.py, and does all the arithmetic.

cold_files/cold_service.py is lesson 2's service, rewritten so that every step of starting it reads the clock. It adds GET /ready, which says 503 until a warm-up is done, and GET /startup, which hands the clock readings to the client after the run.

cold_files/cold_client.py starts the service, waits for the port or for , sends the requests and stops the service. It uses only the standard library, for the reason on the machine slide. Its open-loop mode is lesson 4's design, rewritten without numpy.

How to Make a Cold Start Hurt Less

Here is the order I would follow, using only what this lab did.

A flowchart. Time one start with a clock mark per step leads to a diamond: imports the biggest part? Yes leads to import less, or start fewer times; no leads to look at the step that is biggest. Both lead to warm up before /ready says 200, then to keep files and layers near the machine. Below: here, imports were 96.9% of the time to the port in A; one HTTP warm-up took the first real request from 4,680 to 1,655 us; an emptied page cache made the start 2.81 times as long. Caption: measure the start before you shorten it.

The chart starts by measuring, because the right fix depends on which step is biggest. Here are the same steps in words.

  1. Time one start, step by step. Read one clock at each step, as cold_service.py does, or run python -X importtime. Here the imports were 96.9% of the time until the port opened.

  2. If imports are the biggest part, import less, or start less often. Import only what the service uses. Start the service once and keep it running, rather than once per request or per job. Here, one start cost 1.5 seconds, and one request about 1 millisecond.

  3. Put compiled files in the image. Run python -m compileall on the standard library when you build it. Here that saved 200 ms on every container start.

  4. Warm up, then say ready. Send one real request through the whole path, and answer the readiness check with 200 only after it. Here it took the first real request from 4,680 to 1,655 us.

  5. Make the ask the right question. A check that a port is open is not enough. Here the port opened 76 ms before the service was warm with 50 warm-ups (C50), and about 16 ms before with one (C1).

When Cold Starts Matter, and When They Do Not

Cold starts matter when a service starts often. If it scales up and down with traffic, or runs as a short job for each batch, every start pays the import time again. Here that was 1.5 seconds on a warm machine and 4.2 seconds on a cold one.

They matter for the first users after a deploy. Every new copy has a first request. Here it cost about 3 ms more than the next, unless a warm-up paid it first.

They matter most on a new machine. Then the files are not in memory, and the image may not be there at all. Here an emptied page cache made the start 2.81 times as long, and loading the image took 8.66 seconds before the start could even begin.

They matter less for a service that starts once and runs for weeks. One slow start a week is not worth much effort. Under steady traffic here, I could not even tell a warmed start from a cold one.

Do not judge the model file first. It is easy to blame the model for a slow start. Here the model took 2.0 ms of a 1,499 ms start. Measure before you shrink the model.

What This Lab Cannot Tell You

Two columns titled shows and cannot show. Shows: one small model and its libraries on one 1-vCPU box; the client shares the core. A page cache emptied with drop_caches; images loaded from a file or pulled from a nearby registry. The bare and the container Python are different builds of 3.13.15. Cannot show: a brand-new disk made from a snapshot, whose blocks AWS fetches on first read. A machine that must first download its own image from far away, or a large model file. How much of D against A is the container and how much is the build.

One small model. This model loads in 2 ms. A model of several gigabytes would make loading the model the biggest step, and the answer would change. I did not measure one.

One core, and the client on it. Starting a service is mostly one thread of work, so I expect the shape to hold on bigger machines. But I measured only this one.

A cold cache, not a cold disk. Emptying the page cache simulates a machine that has not read these files recently. A new disk made from a snapshot can be much slower on first read. Pulling a large image from far away can take much longer too.

Two builds of Python. The bare and container Pythons differ (Clang against GCC), so D against A mixes the container and the build. D2 against D does not. But D2's layer also compiled 84 site-packages files that had none (most likely pip's own, which the service does not import; I did not check which).

An emptied page cache is not a reboot. drop_caches keeps pages that running programs have mapped. About 154 MB stayed cached after each drop, including the programs dockerd and containerd. So arm E was cold for the image's files, but itself was warm, unlike just after a reboot.

Labelled additions and repeats. The clock check was changed after the results, as the clock slide explains. The image timings were repeated after the results with a cleaner start, as the image slide explains. Both changes are in the lab's notes, and the first numbers are kept.

What to Do on Monday

A hand-drawn grid of six cards, titled five habits for cold starts. 1, time the start: a clock reading per step, imports, model, server. 2, gate on ready: /ready says 200 only after the warm-up. 3, warm up once: one real request through the whole path. 4, ship compiled files: .pyc for the standard library in the image. 5, keep starts rare: a start costs seconds; do not restart per request. The reason: here a start took 1,499 ms warm and 4,215 ms cold; the model 2.0 ms. Caption: the first request is paid for by someone; make it the service, not a user.

If you take one thing to work on Monday, time one start of your model service. Run python -X importtime -c "import your_service" and sort by the cumulative column. You will probably find, as I did, that the libraries cost far more than the model.

Then look at how your service says it is ready. If the only checks that a port is open, add a /ready route that answers 200 only after one real warm-up request has gone through.

The one idea to keep: a service is not ready when its port opens. It is ready when someone has already paid for the first request, and that someone should be the service.

Knowledge Check

Knowledge Check

4 questions - Score 80% to pass

Q1

In arm A, a bare process with its files in memory, which part took most of the time before the port opened?

Q2

Why did the image with a compiled standard library (D2) start faster than the plain image (D)?

Q3

What did one warm-up request behind a /ready check (C1) do to the first real request?

Q4

Under 539 random arrivals a second, what did the lab find for F (no warm-up) against G (50 warm-ups and /ready)?

F
G
/ready

To keep the files warm for a warm arm, the lab first starts and stops the same kind of service once, without timing it. Then it waits until the core is at least 95% idle.

The rule, written before the runs, is lesson 3's rule. Two arms are "told apart from run-to-run noise" only if all eight paired differences have the same sign, and the two sets of eight values do not overlap. I claim that one arm beat another only when I compared those two arms directly.

HERE
B
S
C
copy
S
C
IP
export AWS_DEFAULT_REGION=us-east-1 AWS_PAGER=""
NAME=ai-research-course-cold
HERE=$PWD
STATE=${COLD_STATE:-$HOME/lab-data/serving/cold}
KEY=$HOME/lab-data/serving/keys/$NAME.pem
K=${K:-8}
mkdir -p "$STATE" "$HOME/lab-data/serving/keys"
SSHO=(-o LogLevel=ERROR -o StrictHostKeyChecking=accept-new -o "UserKnownHostsFile=$STATE/known_hosts")
[ -f "$STATE/id" ] && ID=$(cat "$STATE/id")
[ -f "$STATE/ip" ] && IP=$(cat "$STATE/ip")
[ -f "$STATE/sg" ] && SG=$(cat "$STATE/sg")
  B=/home/ubuntu/cold
  S=(ssh -i "$KEY" "${SSHO[@]}" ubuntu@"$IP")
  C=(scp -q -i "$KEY" "${SSHO[@]}")

What it costs. A c7g.medium costs $0.0363 an hour on demand in us-east-1, from AWS's price list, and AWS bills it by the second. My measuring box was on for 0.444 hours, from launch to terminated, which is $0.0161 for compute. An earlier box, on 2026-10-04, ran the set-up steps but took no timing and was deleted the same day: 0.068 hours, $0.0025. So the whole lesson cost $0.0186. The disks were deleted with the boxes. These numbers are in results/cold-cost.json and results/cold-cost-setup-box.json. The schedule itself took 14 minutes.

What the recordings hide. Every line passed through a small filter, cold_redact.py, a copy of lesson 5's. It hides my account number, every IP address, every name AWS makes up for a resource, host names, my user name and anything from a key file. Where you see <ip> or i-<id>, your terminal shows the real value. Lines that start with + are the shell printing each command just before it runs it.

Keep the key file outside any git folder. Mine lives in ~/lab-data/serving/keys/, and the last step deletes it.

Step 3: rent the box. First the script asks AWS's parameter store (SSM) for the id of the newest Ubuntu 24.04 image for arm processors. So you never copy an id that has gone out of date. Then run-instances asks for one c7g.medium with a 16 GiB disk that is deleted with the box, with the tags on both.

A real terminal recording of step 3, with the image, group and instance ids replaced by placeholders. aws ssm get-parameter reads /aws/service/canonical/ubuntu/server/24.04/stable/current/arm64/hvm/ebs-gp3/ami-id into AMI; aws ec2 run-instances with that image, type c7g.medium, the key ai-research-course-cold, a 16 GiB gp3 disk deleted on termination and the tags Project, Lesson and Name, asking only for the instance id, which goes into ID; describe-instances prints the id, c7g.medium, pending and the image id.

The box started in the "pending" state, which is normal for the first few seconds. These are the two commands, the image lookup and the launch:

  AMI=$(aws ssm get-parameter --query Parameter.Value --output text \
    --name /aws/service/canonical/ubuntu/server/24.04/stable/current/arm64/hvm/ebs-gp3/ami-id)
  ID=$(aws ec2 run-instances --image-id "$AMI" --instance-type c7g.medium --count 1 \
    --key-name ai-research-course-cold --security-group-ids "$SG" \
    --block-device-mappings '[{"DeviceName":"/dev/sda1","Ebs":{"VolumeSize":16,"VolumeType":"gp3","DeleteOnTermination":true}}]' \
    --tag-specifications \
    "ResourceType=instance,Tags=[{Key=Project,Value=ai-research-course},{Key=Lesson,Value=cold},{Key=Name,Value=ai-research-course-cold}]" \
    "ResourceType=volume,Tags=[{Key=Project,Value=ai-research-course},{Key=Lesson,Value=cold}]" \
    --query "Instances[0].InstanceId" --output text)

The script also writes ID and SG into small files in ~/lab-data/serving/cold/, so the later steps can read them back.

  "${S[@]}" "curl -LsSf https://astral.sh/uv/install.sh | sh > /dev/null 2>&1 && \
    ~/.local/bin/uv python install 3.13.15 > /dev/null 2>&1 && \
    ~/.local/bin/uv venv --allow-existing --python 3.13.15 $B/venv > /dev/null 2>&1"
  "${C[@]}" "$HERE/examples/lat_requirements.txt" ubuntu@"$IP":$B/
  "${S[@]}" "~/.local/bin/uv pip install --python $B/venv/bin/python -q -r $B/lat_requirements.txt"
  "${S[@]}" "sudo apt-get update -q > /dev/null && \
    sudo DEBIAN_FRONTEND=noninteractive apt-get install -y -q docker.io docker-buildx > /dev/null && \
    sudo usermod -aG docker ubuntu && docker --version"

S and C are short names for the ssh and scp commands with the key, from the block above. The full list of files to copy is in the copy step.

Step 6: build the two container images. The first image, cold-svc:lab, follows packaging lesson 5's slim pattern. It starts from the official python:3.13.15-slim image, pinned by its digest, a fingerprint of the image's exact contents. It adds the same pinned packages, the model, the SQLite file and the service.

An image is built in layers, each a set of file changes stacked on the one below. The second image, cold-svc:pyc, adds one layer: Python's standard library compiled to .pyc files with compileall. That compileall over /usr/local/lib/python3.13 also compiled 84 site-packages files that had none. They are most likely pip's own, which the service does not import; I did not check which.

A real terminal recording of step 6, with the address replaced by a placeholder. Over ssh: copy the requirements, model, SQLite file, rows and service into a build folder; docker build cold-svc:lab from Dockerfile; docker build cold-svc:pyc from Dockerfile.pyc; docker images prints cold-svc:pyc 962MB and cold-svc:lab 948MB.

The compiled image is 14 MB larger, as Docker counts it. This is the command:

  ssh -i ~/lab-data/serving/keys/ai-research-course-cold.pem "${SSHO[@]}" ubuntu@"$IP" \
    'cd cold && cp lat_requirements.txt model.pkl store.sqlite cold_rows.json cold_service.py ctx/ && \
     docker build -q -t cold-svc:lab -f Dockerfile ctx > /dev/null && \
     docker build -q -t cold-svc:pyc -f Dockerfile.pyc ctx > /dev/null && \
     docker images --format "{{.Repository}}:{{.Tag}}  {{.Size}}"'

Step 7: stop the timers and look at the box. Ubuntu runs small jobs on a timer. One of those starting during a run would land in the results, so the script stops every timer first. Then it prints what the box is, saves a fuller record, and lists every installed package with its version.

A real terminal recording of step 7, with the address replaced by a placeholder. Over ssh: stop every active timer and the unattended-upgrades service, after which systemctl prints 0 timers listed; lscpu prints architecture aarch64, 1 CPU, vendor ARM, model name Neoverse-V1, 1 thread per core, 1 core per socket; nproc prints 1; uptime prints up 5 min with load average 0.79, 0.58, 0.26. Then the box info is saved, and Python 3.13.15 and the full pip freeze are printed, including fastapi 0.142.2, numpy 2.5.3, pandas 3.0.6, scikit-learn 1.9.1, uvicorn 0.54.0 and uvloop 0.23.0.

The load average, Linux's running average of how many programs want the core, was still high from step 6. So the schedule's first job is to wait until it falls to 0.10. These are the commands that stop the timers and look:

  ssh -i ~/lab-data/serving/keys/ai-research-course-cold.pem "${SSHO[@]}" ubuntu@"$IP" \
    'for t in $(systemctl list-units --type=timer --state=active --no-legend --plain | cut -d" " -f1); do
       sudo systemctl stop "$t"; done; sudo systemctl stop unattended-upgrades.service
     systemctl list-timers --no-pager | tail -2'
  ssh -i ~/lab-data/serving/keys/ai-research-course-cold.pem "${SSHO[@]}" ubuntu@"$IP" 'lscpu | head -12; nproc; uptime'

Step 8: start the schedule, then leave. The schedule starts in the background with setsid nohup, so it keeps running after SSH disconnects. Then I log out, and nobody logs in until it ends.

A real terminal recording of step 8, with the address replaced by a placeholder. Over ssh: cd cold, then setsid nohup bash cold_schedule.sh 8 with its output sent to raw/schedule.log, in the background; two seconds later the log shows 02:12:33 schedule start K=8.

The log's first line appeared, and the SSH session closed. This is the command:

  ssh -i ~/lab-data/serving/keys/ai-research-course-cold.pem "${SSHO[@]}" ubuntu@"$IP" \
    "cd cold && (setsid nohup bash cold_schedule.sh $K > raw/schedule.log 2>&1 < /dev/null &); \
     sleep 2; cat raw/schedule.log"

A real terminal recording of step 13, with the address replaced by a placeholder. Over ssh, the box records its info after the schedule and prints the size of its raw folder, 5.4M; then rsync copies the box's cold/raw folder to my laptop; ls counts 258 files; du prints 5.4M for the copy.

I ran this step twice. The first fetch came right after the demo. Then I repeated the image timings (see the slide on images), and this recording is the second fetch, which brought those files too. Rsync only adds and updates files, so running it twice is safe. This is the copy command:

  rsync -a -e "ssh -i $KEY ${SSHO[*]}" ubuntu@"$IP":cold/raw/ ~/lab-data/serving/cold/raw/

Step 14: delete the box, the key and the door rule, and check. A box you forget keeps costing money. The script deletes the box, waits until AWS says "terminated", then deletes the key pair and the security group. Then it asks AWS again.

A real terminal recording of step 14, with the instance and group ids replaced by placeholders. terminate-instances prints shutting-down; wait instance-terminated; delete-key-pair prints true; delete-security-group prints true. Then the checks: describe-instances prints terminated; the count of key pairs named ai-research-course-cold is 0; the count of security groups of that name is 0; the count of volumes tagged Lesson=cold is 0; the count of course instances not terminated is 0. Last, the key file is removed.

Every count printed 0 and the box printed terminated. I asked AWS once more afterwards, from my laptop, and all four counts were still 0. These are the teardown commands:

  aws ec2 terminate-instances --instance-ids "$ID" --query "TerminatingInstances[0].CurrentState.Name" --output text
  aws ec2 wait instance-terminated --instance-ids "$ID"
  aws ec2 delete-key-pair --key-name ai-research-course-cold --query Return
  until aws ec2 delete-security-group --group-id "$SG" --query Return 2>/dev/null; do sleep 10; done
  aws ec2 describe-instances --instance-ids "$ID" --query "Reservations[0].Instances[0].State.Name" --output text
  aws ec2 describe-key-pairs --filters "Name=key-name,Values=ai-research-course-cold" --query "length(KeyPairs)"
  aws ec2 describe-security-groups --filters "Name=group-name,Values=ai-research-course-cold" \
    --query "length(SecurityGroups)"
  aws ec2 describe-volumes --filters "Name=tag:Lesson,Values=cold" --query "length(Volumes)"
  aws ec2 describe-instances --filters "Name=tag:Project,Values=ai-research-course" \
    "Name=instance-state-name,Values=pending,running,stopping,stopped,shutting-down" \
    --query "length(Reservations[].Instances[])"
  rm -f ~/lab-data/serving/keys/ai-research-course-cold.pem

So every start compiled again whichever standard library modules it imported that had no .pyc file, and threw those copies away.

The bare Python on the box held 206 compiled standard library files after its runs, which is roughly how many modules this service imports. Image D2 compiled all 633 ahead of time with python -m compileall. Its compileall over /usr/local/lib/python3.13 also compiled 84 site-packages files that had none. They are most likely pip's own, which the service does not import; I did not check which. Its imports took 1,589 ms, and its first answer came 200 ms earlier than D's, told apart. Python's tutorial says it plainly: "the only thing that’s faster about .pyc files is the speed with which they are loaded."

One honest limit. The bare Python and the container's Python are both 3.13.15, but they are different builds, made with different compilers. So some of D against A may come from the build, not from the container. D2 against D is a much cleaner comparison: the same image plus one layer. That layer holds the compiled standard library and the 84 extra site-packages files, so the 200 ms belongs to the layer as a whole.

AWS
  • "H11 D: launch -> first answer above A's, told apart, by at least 0.3 s (median difference)." Right: 539.5 ms.

  • "H12 D2: launch -> first answer below D's, told apart." Right: 199.9 ms earlier.

  • "H13 E: launch -> first answer above D's, told apart." Right: 3,185 ms later.

  • "H14 G: first-second max below F's, told apart." Wrong: it could not be told apart.

  • "H15 F: first-second p99 at least 3 x third-second p99 (median of the runs)." Wrong: 1.40 times.

  • "H16 every returned score equals the box's single-row score to the bit, in every arm." Right: 32,379 of 32,379. This is a check that every arm served the right model on the right rows, not a finding. The box's reference matched lesson 2's stored scores on 26,818 of 26,851 rows, with a largest gap of 1.1e-16.

  • "H17 load of the image from the file (cold cache) between 5 and 60 s." Right: 8.66 s, judged on the clean repeat added after the results.

  • "H18 -X importtime, warm cache: the largest top-level cumulative import is sklearn.ensemble." Right, in all 5 warm runs.

  • r"""What does the first request pay? Time a model's cold start, piece by piece, on YOUR machine.
    
    Lesson 6 of 'Serving and Inference Basics'. It uses lesson 2's Python environment, made once inside this folder
    (Python 3.13 is what the lab used):
        python3 -m venv venv-sv
        source venv-sv/bin/activate              # Windows: venv-sv\Scripts\activate
        pip install -r lat_requirements.txt      # numpy, pandas, pyarrow, scikit-learn, threadpoolctl, ...
    It also needs the shop data from the features chapter: run python ../../features/fetch_data.py once first. Then:
        python cold_demo.py                  # train, then print parts 2 to 4 below, numbered 1 to 3
        python cold_demo.py --save out.json  # and save the numbers
    If a package is missing, it prints one line saying why and stops. It needs no Docker, no web server and no
    administrator rights. It writes a model file and a few rows into a temporary folder, deletes that folder at the
    end, and writes nothing else except out.json if you ask.
    Every time it prints comes from YOUR machine, as it is right now: other programs, your disk and whether these files
    were read recently all move them. Read them for what they show about your machine.
    
    Design, written 2026-10-04 after the lab's design (cold_start.py) and before this file first ran.
      1. Train the features chapter's model (it must score test AP 0.5450), as lesson 5's wrk_demo.py does, and save it
         with pickle, with the first 100 test rows, into a temporary folder.
      2. Five fresh Python processes, one after another (this same python, OMP_NUM_THREADS=1). Each reads the clock on
         its first line, then imports numpy, pandas and sklearn.ensemble, reads the model file, unpickles it, and calls
         predict_proba on one row 100 times, each on the next test row. The clock is time.monotonic_ns(), read by this
         process just before it starts each child, and by the child after each step. Printed: the median of the five
         for each step, and for predict calls number 1, 2 and 100.
      3. python -X importtime of the same imports, once: the five biggest top-level imports by their total time.
      4. A warm-up: five more fresh processes, the same steps, but ONE predict_proba call on a row that is not a test
         row first (the warm-up), then the 100 calls. Printed: call number 1 with and without the warm-up.
      The page cache cannot be emptied without administrator rights, so every child here starts with its files
      already read recently (the first child may not): what the lesson calls a warm cache.
    
    Author: Roni Das
    Created: 2026-10-04
    """
    import importlib.util
    import json
    import os
    import pickle
    import statistics
    import subprocess
    import sys
    import tempfile
    import time
    from pathlib import Path
    
    NEEDED = ("numpy", "pandas", "pyarrow", "sklearn", "threadpoolctl")
    missing = [m for m in NEEDED if importlib.util.find_spec(m) is None]
    if missing:
        sys.exit(f"cold_demo.py needs {', '.join(missing)}: make the environment in its docstring "
                 f"(pip install -r lat_requirements.txt), then run it with that environment's python.")
    
    RUNS = 5
    CHILD = r'''
    import time
    T = [("python ready", time.monotonic_ns())]
    import sys
    def mark(k):
        T.append((k, time.monotonic_ns()))
    import json, pickle
    from pathlib import Path
    from time import perf_counter_ns
    import numpy as np
    mark("import numpy")
    import pandas
    mark("import pandas")
    import sklearn.ensemble
    mark("import sklearn")
    d = Path(sys.argv[1])
    blob = (d / "model.pkl").read_bytes()
    mark("read model file")
    model = pickle.loads(blob)
    mark("unpickle")
    X = np.load(d / "rows.npy")
    calls = []
    if sys.argv[2] == "warm":
        model.predict_proba(np.load(d / "warm.npy"))
    for i in range(100):
        a = perf_counter_ns()
        model.predict_proba(X[i:i + 1])
        calls.append(perf_counter_ns() - a)
    print(json.dumps({"marks": T, "calls": calls}))
    '''
    
    
    def build():
        """The features chapter's model and its test rows (lesson 5's wrk_demo.py build step)."""
        import numpy as np
        from sklearn.metrics import average_precision_score
        here = Path(__file__).resolve().parent
        sys.path.insert(0, str(here.parents[1] / "features"))
        import task
        from what_a_feature_is import HAND_COLS, hgb, joined
    
        ev = task.load_events()
        lab_tr, _, lab_te = task.splits(ev)
        tr = joined(ev, lab_tr, task.TRAIN_CUTOFFS)
        te = joined(ev, lab_te, task.TEST_CUTOFFS)
        model = hgb(0).fit(tr[HAND_COLS].to_numpy(float), tr["label"].to_numpy())
        X = np.ascontiguousarray(te[HAND_COLS].to_numpy(np.float64))
        p = model.predict_proba(X)[:, 1]
        y, cut = te["label"].to_numpy(), te["cutoff"].to_numpy()
        ap = float(np.mean([average_precision_score(y[cut == c], p[cut == c]) for c in np.unique(cut)]))
        assert round(ap, 4) == 0.5450, ap
        warm = np.ascontiguousarray(tr[HAND_COLS].to_numpy(np.float64)[:1])
        return model, X, warm, ap
    
    
    def child(folder: Path, mode: str) -> dict:
        env = dict(os.environ, OMP_NUM_THREADS="1")
        t = time.monotonic_ns()
        r = subprocess.run([sys.executable, "-c", CHILD, str(folder), mode], capture_output=True, text=True, env=env)
        if r.returncode:
            sys.exit(r.stderr[-800:])
        d = json.loads(r.stdout)
        marks = [("start", t)] + [tuple(m) for m in d["marks"]]
        steps = {marks[i][0]: (marks[i][1] - marks[i - 1][1]) / 1e6 for i in range(1, len(marks))}
        return {"steps_ms": steps, "calls_us": [c / 1000 for c in d["calls"]]}
    
    
    def importtime(folder: Path) -> list:
        code = "import json, pickle; import numpy; import pandas; import sklearn.ensemble"
        r = subprocess.run([sys.executable, "-X", "importtime", "-c", code], capture_output=True, text=True,
                           env=dict(os.environ, OMP_NUM_THREADS="1"))
        top = []
        for line in r.stderr.splitlines()[1:]:
            if not line.startswith("import time:"):
                continue
            _, cum, name = line[len("import time:"):].split("|")
            if not name[1:].startswith(" "):          # one space only: a top-level import (Python's start-up or the -c line)
                top.append((int(cum) / 1000, name.strip()))
        return sorted(top, reverse=True)[:5]
    
    
    def main() -> None:
        try:
            load = f"{os.getloadavg()[0]:.2f}"
        except (AttributeError, OSError):
            load = "not available on this system"
        print(f"Every time below was measured on THIS machine ({os.cpu_count()} cores, 1-minute load {load}).")
        import numpy as np
        model, X, warm, ap = build()
        print(f"model trained: test AP {ap:.4f}; {len(X):,} test rows")
        with tempfile.TemporaryDirectory() as tmp:
            folder = Path(tmp)
            (folder / "model.pkl").write_bytes(pickle.dumps(model, protocol=5))
            np.save(folder / "rows.npy", X[:100])
            np.save(folder / "warm.npy", warm)
            cold = [child(folder, "plain") for _ in range(RUNS)]
            top = importtime(folder)
            hot = [child(folder, "warm") for _ in range(RUNS)]
    
        med = lambda xs: statistics.median(xs)  # noqa: E731
        steps = list(cold[0]["steps_ms"])
        print(f"\n1. a fresh Python process, step by step: median of {RUNS} processes, milliseconds")
        for k in steps:
            print(f"   {k:16s} {med([c['steps_ms'][k] for c in cold]):9.1f}")
        total = med([sum(c["steps_ms"].values()) for c in cold])
        print(f"   {'all of it':16s} {total:9.1f}   (start to model loaded; median of the {RUNS} totals)")
        calls = {n: med([c["calls_us"][n - 1] for c in cold]) for n in (1, 2, 100)}
        print("   predict_proba on one row, microseconds (us): "
              + ", ".join(f"call {n} {v:.1f}" for n, v in calls.items()))
    
        print("\n2. python -X importtime: the five biggest top-level imports, total ms (one run)")
        for ms, name in top:
            print(f"   {name:16s} {ms:9.1f}")
    
        w1 = med([h["calls_us"][0] for h in hot])
        print("\n3. the first real predict_proba call, median of 5 processes, us")
        print(f"   no warm-up                  {calls[1]:9.1f}")
        print(f"   after one warm-up call      {w1:9.1f}")
        if "--save" in sys.argv:
            out = {"load": load, "cores": os.cpu_count(), "test_ap": ap, "runs": RUNS,
                   "steps_ms": {k: med([c["steps_ms"][k] for c in cold]) for k in steps}, "total_ms": total,
                   "calls_us": {str(n): v for n, v in calls.items()}, "importtime_top": top, "warm_call1_us": w1}
            Path(sys.argv[sys.argv.index("--save") + 1]).write_text(json.dumps(out, indent=1))
    
    
    if __name__ == "__main__":
        main()
    

    You saw the demo running on the lab's box in step 12, after the schedule. That exact run is stored in results/cold-demo-run.txt, and the report checks that every number it printed matches the numbers it saved.

    On the box it told the same story as the lab. Starting Python and loading the model took 1,014 ms, and scikit-learn's import was 721 ms of it. Reading and unpickling the model took about 1 ms. The first predict_proba call took 3,004 us, the second 618 us, and after one warm-up call the first real call took 623 us.

    Its numbers are a little lower than the lab's start, because the demo does not import FastAPI or uvicorn and does not start a web server.

    And here is the same demo in VS Code on my laptop, run with python cold_demo.py in the examples folder with venv-sv active.

    A real screenshot of VS Code's terminal on my laptop, in the venv-sv environment, after running python cold_demo.py. The first line says every time below was measured on this machine: 10 cores, 1-minute load 14.25. Then: model trained, test AP 0.5450, 26,851 test rows. Part 1, a fresh Python process step by step, median of 5 processes, in ms: python ready 17.8, import numpy 27.0, import pandas 150.9, import sklearn 532.5, read model file 0.1, unpickle 0.3, all of it 729.2; predict_proba on one row, call 1 1801.4 us, call 2 197.5, call 100 156.5. Part 2, python -X importtime, the five biggest top-level imports in ms: sklearn.ensemble 535.1, pandas 153.5, numpy 24.3, site 3.9, pickle 1.1. Part 3, the first real predict_proba call, median of 5 processes: no warm-up 1801.4 us, after one warm-up call 199.5 us.

    This run is from my laptop, not the box. It is a different machine, with 10 cores, and it was heavily loaded when I ran it: the 1-minute load was 14.25. So these numbers are not to be compared with the box's numbers in any way. Look only at the shape, which is the same as on the box. scikit-learn's import was the biggest step, 532.5 of the 729.2 ms from start to model loaded. Reading and unpickling the model took 0.4 ms. The first predict_proba call took 1,801.4 us, about nine times the second call's 197.5 us. After one warm-up call, the first real call took 199.5 us.

    /ready

    cold_files/cold_run_once.sh warms or empties the page cache, waits for an idle core, runs the client, and records the load and the system log. cold_files/cold_schedule.sh runs everything in order. cold_files/cold_image_again.sh is the repeat of the image timings, added after the results. Dockerfile and Dockerfile.pyc build the two images.

    cold_report.py does not import any of the above. It reads the small copy of the raw files with its own code and recomputes every number. It also checks the guesses, the demo's stored run, the fact-check, the teardown, the redaction, and it writes the playground.

  • Keep files close to the machine. Files in memory started the service 2.81 times faster than files on disk. An image already on the machine starts at once; one that must be loaded took 8.66 seconds here.