Let me start at a desk in an office, early in the morning.
The desk opens at nine. The first person in the queue walks up at one minute past. The clerk has to switch on the lights and the computer, wait for it to start, open the right program and find the right drawer. Only then can she help. The second person waits only for the first one to finish.

Nobody did anything wrong. The desk was simply not ready yet. A new copy of a prediction service has the same problem. Before it can answer anyone, it must start a program, load its tools, load the model and open its door. And the very first question it answers may still be slower than the ones after it.
In this lesson I measure that start, piece by piece, on a rented machine. I ask where the time goes, and which parts a real user would feel. Then I try four ways to hide those parts from users. I keep the files in memory, send a warm-up request, put a readiness check in front, and run the service in a container.
This is lesson 6 of the chapter on serving a model. Lesson 2, latency anatomy, took one request apart. Lesson 3, tail latency, found that the slowest request of a run was always the first one, but it measured only after the service had started. Lesson 5, workers and threads, showed that one worker with one thread was the best choice on one core. So in this lesson the service always runs with one worker and one thread.
The model, the service code and the rented machine are the same as in lessons 2 to 5, so I do not explain them again. In short: the model is the features chapter's gradient boosted tree model, which scores one customer of a real online shop and reached a test AP of 0.5450. The service is a small FastAPI web service run by uvicorn. It reads the customer's six features from SQLite.
The container image follows packaging lesson 5, container images, which showed how to build a slim image and what was inside it. I do not repeat that here. This lesson asks a different question about the same kind of image: how long does it take to start?
Please read this slide slowly if any word is new. Every slide after it uses these words.

A cold start is everything a new copy of a service does before it can answer, plus anything extra its first request pays. A process is a running program with its own memory. Starting a new process here means starting Python again from nothing.
To import a library is to make Python read the library's files and run its set-up code. This happens once per process. A .pyc file is a module that Python has already turned into its own internal code, called bytecode. When the .pyc file exists and is up to date, the next import skips that step. To unpickle the model means to turn the bytes of the saved model file back into a model object in memory, as packaging lesson 1 showed.
The page cache is a part of memory where Linux keeps copies of files it has read recently. When a file is in it, reading the file is fast. I call that a warm cache. When it is not, Linux must read from the disk: a cold cache.
A container is a process started by from an image, which is the packed set of files it needs: Python, the libraries and the model. A warm-up is a request or a call that the service makes to itself before real users arrive. A readiness check is a question that a asks the service again and again: are you ready for traffic yet? Here it is a request to GET /ready. The port opening is the moment a network connection to the service first succeeds.
is the time from a request to its answer. The p99 is the 99th percentile: the time that 99 of every 100 requests stay under.
Here is the main result first. Each lane is one way of starting the service, from the moment it was started to the moment the first answer came back. Each lane is the median of eight starts.

A fresh service answered for the first time after 1,499 milliseconds. That was a bare Python process on the box, with all its files already in memory.
Almost all of that was Python importing libraries. The imports took 1,449 ms, which is 96.9% of the time until the port opened. scikit-learn alone took 824 ms. Reading the model file and unpickling it took 2.0 ms.
Files that were not in memory made the start 2.81 times as long: 4,215 ms. The container image, whose Python is a different build, added 539 ms (container and build not separated). A container whose image held compiled copies of Python's standard library gave 200 ms of that back.
The first request itself was slower than the second: 4,680 microseconds against 1,849. One warm-up request, sent by the service to itself before it said "ready", brought the first real request down to 1,655 microseconds.
But under 539 random arrivals a second, I could not tell a warmed service from a cold one: the run-to-run spread was larger than any difference between them.
I wrote the lab's design into the docstring of scripts/labs/serving/cold_start.py on 2026-10-04, before any timed run. It names every way of starting, every comparison, the rule for "told apart from noise", and eighteen guesses.

Every run starts a brand-new copy of the service. There are ten ways to start it, which I call arms. In each of eight rounds, every arm runs once, in a new random order.
A is the plain case: a bare Python process, with its files already in the page cache. B is the same, but I empty the page cache just before the start. C1 and C50 send 1 or 50 requests to themselves after the port opens, and answer /ready with "503, not yet" until the warm-up is done.
I50 makes 50 lookups and model calls inside the process before it opens the port, with no warm-up. D runs the service in a container. D2 uses an image with the standard library compiled. Its compileall over /usr/local/lib/python3.13 also compiled 84 site-packages files that had none. They are most likely pip's own, which the service does not import; I did not check which. E is D with the page cache emptied.
and test the start under traffic. F has no warm-up, and arrivals begin when the port opens. G has 50 warm-ups and waits for to say 200.
Like lessons 2 to 5, every timing comes from a small machine I rented from Amazon Web Services. None comes from my laptop, which is always busy with other work.

The box is an EC2 c7g.medium, with one virtual CPU and 2 GiB of memory. My AWS account allows only one virtual CPU at a time, as lesson 5 explained. So the client that starts the service and sends requests runs on the same core as the service.
The client runs in a different Python on purpose. In the cold-cache arms, I empty the page cache just before the start. If the client imported numpy, it would read the very files the service is about to read, and the service would start warm. So the client uses only Ubuntu's own Python and its standard library. It shares no files with the service's Python, except the basic C library.
Before every one of the 110 runs, the core was at least 99.5% idle. The machine under my box took no measurable time from it.
Before the recordings, here is the whole lab as one numbered sequence. The numbers on the arrows match the numbers on the recordings that follow.

Steps 1 to 8 happen once: check, rent the box, set it up, build the images and start the schedule. Step 9 is how I watched the schedule without touching the box. Arrows 10 and 11 are the timed work, repeated for 80 starts with nobody logged in. Steps 12 to 14 come after the schedule: the student demo, the copy of the results back to my laptop, and the deletion of everything.
Three things are not in this picture, and I want to be open about them. Before the schedule, I ran one untimed start of each arm over SSH, only to check that every arm worked. That was the first real test of the Dockerfiles. After the demo, I copied one more script to the box with scp. Then I ran one more block of timings over SSH, without a recording, to repeat a measurement that had gone wrong. The slide about images explains it. And the fetch in step 13 ran twice: the first rsync, right after the demo, is not recorded either; the recording is the second one.
The rented box is part of the lab, so I show every step of it. You can rent the same kind of box and run the same schedule.
Which box you see. Every recording comes from the box that measured this lesson's numbers. No recording happened during a timed run, and nobody logged in while the schedule ran.
What you need first. An AWS account, the AWS command line tool (aws) set up with your keys, and ssh, scp, rsync and curl. All the steps live in one file, scripts/labs/serving/cold_files/cold_aws_steps.sh. You run it from the scripts/labs/serving folder, one step at a time: bash cold_files/cold_aws_steps.sh check, then access, and so on. Every code block below is copied word for word from that file.
The blocks use a few names the file sets at its top. Run these lines first, from the scripts/labs/serving folder, if you paste the blocks by hand. All of them are copied from the file, except HERE=$PWD, the folder B and the two short names S and C. The file works out from its own location, which a pasted line cannot do. It sets , and inside its step. and use the box's address, , so run those two lines only after step 4 has set it.
Step 1: is any other course box running? My account allows one box of this kind at a time, and another lesson may be using it. So the first command counts the course's boxes that are not yet deleted. It must print 0.

The recording shows the command and its answer, 0, so no other course box was alive. This is the command, copied from the step script:
aws ec2 describe-instances --filters "Name=tag:Project,Values=ai-research-course" \
"Name=instance-state-name,Values=pending,running,stopping,stopped,shutting-down" \
--query "length(Reservations[].Instances[])"
Step 2: a key and a locked door. A key pair is how you prove to the box that you are allowed in. A security group is a door rule for the box. Mine opens only port 22, the SSH port, and only to my own address. In the rule, /32 after an address means that one address and no other. Both get tags, labels that say which project and lesson they belong to.

The key material went straight into the file and was never printed, and the door rule names only my own address. These are the commands:
aws ec2 create-key-pair --key-name ai-research-course-cold --key-type ed25519 \
--tag-specifications "ResourceType=key-pair,Tags=[{Key=Project,Value=ai-research-course},{Key=Lesson,Value=cold}]" \
--query KeyMaterial --output text > ~/lab-data/serving/keys/ai-research-course-cold.pem
chmod 600 ~/lab-data/serving/keys/ai-research-course-cold.pem
MYIP=$(curl -sf https://checkip.amazonaws.com)
VPC=$(aws ec2 describe-vpcs --filters Name=isDefault,Values=true --query "Vpcs[0].VpcId" --output text)
SG=$(aws ec2 create-security-group --group-name ai-research-course-cold --vpc-id "$VPC" \
--description "lesson 6 cold start, ssh from one address" \
--tag-specifications "ResourceType=security-group,Tags=[{Key=Project,Value=ai-research-course},{Key=Lesson,Value=cold}]" \
--query GroupId --output text)
aws ec2 authorize-security-group-ingress --group-id "$SG" --protocol tcp --port 22 --cidr "$MYIP/32" \
--query "SecurityGroupRules[].{port:FromPort,from:CidrIpv4}" --output table
Step 4: wait until it runs. A new box takes a short while to start. The AWS tool waits for you. Then the script reads the box's public address into IP, and tries SSH every five seconds until the box answers.

This time SSH answered on the first try. These are the commands, with the retry loop:
aws ec2 wait instance-running --instance-ids "$ID"
IP=$(aws ec2 describe-instances --instance-ids "$ID" \
--query "Reservations[0].Instances[0].PublicIpAddress" --output text)
until ssh -i ~/lab-data/serving/keys/ai-research-course-cold.pem "${SSHO[@]}" -o ConnectTimeout=5 \
ubuntu@"$IP" true 2>/dev/null; do sleep 5; done
Step 5: Python, and the lab files. The box gets Python 3.13.15 and the same pinned library versions as lessons 2 to 5, through uv, a fast installer for Python. It also gets Docker from Ubuntu's own packages. Then scp sends lesson 2's model, feature table and request list, this lesson's files, and what the demo needs. The last command builds the SQLite table, computes a reference score for every request row, and prints the files' sha256 fingerprints.

The three fingerprints match the ones lesson 2 recorded, so the box serves the same model. These are the commands that install Python, the packages and Docker:
Step 9: watch without touching. The schedule writes one progress line to the box's serial console after every run. The serial console is a log that AWS keeps for the box, and you can read it through AWS without logging in.

The schedule waited 120 seconds for the box to calm down, then started round 1. Each start took about eight seconds, warm-up start included. This is the command:
aws ec2 get-console-output --instance-id "$ID" --latest --output text | grep COLD-PROGRESS | tail -12
Step 12: the student demo, on the same box. After the schedule, I ran the demo you will meet later, so you can see what it prints on one core.

This run happened after the last timed run of the schedule, so it could not disturb any timing. The same output is stored in results/cold-demo-run.txt. This is the command:
ssh -i ~/lab-data/serving/keys/ai-research-course-cold.pem "${SSHO[@]}" ubuntu@"$IP" \
'cd cold/serving/examples && ../../venv/bin/python cold_demo.py --save ~/cold/raw/cold-demo.json \
| tee ~/cold/raw/cold-demo-run.txt'
Step 13: bring the results home first. The rule from lesson 3 is: the moment the runs finish, copy the raw files back, before anything else.
This is a real recording of the report script, cold_report.py. It ran on my laptop, but it times nothing. It reads the raw files the box measured and does all the arithmetic again with its own code.

The report reads the small copy of the raw files kept in the repo, about 1.8 MB, gzipped. It computes its own steps, medians, percentiles, paired differences and verdicts, and compares every one with the lab's results file. If any number disagreed, it would stop with an error. I checked that too: I changed one number in a copy of the results file, and the report printed "2 MISMATCHES" and stopped.
The system log had 452 lines during the 110 runs. Most came from and its helper containerd starting and stopping containers, from sudo emptying the page cache, and from systemd. Nine came from cron, which still ran because I stopped only the systemd timers. They fell in three runs, and each landed inside its arm's range. r2-A answered first at 1,484.9 ms, and A's range was 1,484.9 to 1,523.2. r1-E answered at 5,236.8 ms, inside E's 5,184.4 to 5,257.0. imp-warm-r2 took 1,597.8 ms, inside the warm runs' 1,585.8 to 1,620.7.
To split one start into steps, I need to subtract a time the client read from a time the service read. That only works if both read the same clock.

Every program here reads time.monotonic_ns(), a monotonic clock: one that only ever moves forward, so it is safe for measuring time between two readings. Python's documentation lists how it is built on each system, and on Linux it ends with "Otherwise, call clock_gettime(CLOCK_MONOTONIC)." That is one clock for the whole machine, and a container shares it unless you ask for a separate one.
I checked it on the box, not on trust. On my laptop, readings from two different Python versions did not agree, so the design asked for a check. In all 80 timed starts, every service reading came after the client's launch reading. Every reading up to the server's set-up came before the port opened.
One detail surprised me. My mark for "socket up" is read just after uvicorn's start-up code returns. The socket starts listening inside that code, so in 2 of 80 runs a connection succeeded slightly before my mark, by at most 1.4 ms. The clocks agreed; my mark was simply a little late. I changed the check to stop at the server's set-up, and I say so in the lab's notes.
Now arm A, the plain case, step by step. Each block is one step of the start. Its height is the median time over eight runs.

Starting Python took 18 ms. That is from the client's launch to the first line of my service file.
Then came the imports, 1,449 ms in all. numpy took 61 ms, pandas 251, scikit-learn's ensemble module 824, FastAPI and pydantic 266, and uvicorn 21. By the end, 1,820 Python modules were loaded. I import pandas on its own line before scikit-learn only to show its time separately; scikit-learn would import it anyway.
Loading the model took 2.0 ms: 0.8 ms to read the file's bytes and 1.2 ms to unpickle them. The model is small, about 200 KB, and its class was already imported. Opening SQLite took 0.2 ms, and starting the server and its socket, the network end point that clients connect to, took about 25 ms.
So the model is not the slow part of starting this service. The libraries around it are. A bigger model would change the picture, and I come back to that in the limits.
To check the import times another way, I asked Python itself. Python's -X importtime option prints the time of every import, with "cumulative time (including nested imports) and self time (excluding nested imports)".

I ran it ten times: five with the files in memory and five with the page cache emptied, in a shuffled order. The order of imports matched the service: numpy, then pandas, then scikit-learn, then FastAPI and uvicorn.
scikit-learn's ensemble module was the largest in every run: 769 ms warm. It imports scipy and many of its own parts inside it, which all count in its number. pandas took 235 ms and FastAPI 241. These agree with my own clock marks to within about ten percent, so the two ways of measuring tell the same story.
With the page cache emptied, every library took longer, from 1.90 times for FastAPI to 4.32 times for pandas. The whole process took 4,271 ms instead of 1,598.
Arm B starts the same bare service, but just before the start I empty the page cache: sync; echo 3 > /proc/sys/vm/drop_caches. The kernel's documentation says writing to this file "will cause the kernel to drop clean caches". Its advice is to run sync first, so that more pages are clean and can be dropped.

The whole start took 4,215 ms instead of 1,499: 2.81 times as long, told apart. Every step that reads files from disk got slower. Python itself took 140 ms to start instead of 18, because its own program and standard library had to come from the disk. scikit-learn's import cost the most extra time: 1,255 ms more.
Unpickling did not change: 1.2 ms in both. It reads no file. The bytes were already read in the step before, which did get slower, from 0.8 ms to 2.1.
What does a cold cache mean in real life? It is a machine whose memory does not hold these files: just after a reboot, or after other work has pushed them out. It is not a brand-new disk made from a snapshot. AWS fetches such a disk's blocks from storage on first read, which can be much slower. I did not measure that case.
The first request was slower in B too: 9,432 us against 4,680 (not a declared comparison; B's lowest, 9,115 us, was above A's highest, 5,770). Its own request path had files left to read.
Now arms D and D2. The service, the model and the packages are the same, but the service runs in a container: docker run -d --rm --network host. The option --network host lets the container use the box's own network, so the client connects the same way.

The container's first answer came 539 ms later than the bare process's, told apart. The container image's Python is a different build, so this is the container and the build together, not separated. Two parts make up most of it.
First, starting the container took 202 ms before Python ran a line, against 18 ms for a bare process. has to set up the container's files, its limits and its process. The docker run command itself returned at 199 ms.
Second, the imports took 348 ms longer: 1,797 ms against 1,449. Here a detail of the image matters. A container started with --rm gets a fresh, thin writable layer every time, and the layer is thrown away when it stops. Docker's documentation says: "When you create a new container, you add a new writable layer on top of the underlying layers." So a .pyc file that Python writes during one start is gone at the next.
Most of the standard library had no .pyc files in the image. The box counted 163 compiled files out of 633 standard library modules. The official slim image ships none; pip compiled the ones it used itself during the build.
A container cannot start until its image is on the machine. When a new machine joins, it must first get the image, from a file or from a registry, which is a server that stores images.
The schedule measured both: docker load of my image from a file, and docker pull of the base image from AWS's public copy of 's official images. Each one started from a machine with no images, and an emptied page cache.
The first try was not what the design asked for, and I found it only by looking at the numbers. Each pull took about 0.5 seconds.
That made me check, and containerd's store still held the base image's layers. containerd is the program under Docker that keeps images and runs containers; Docker 29's release notes say "containerd image store is now the default for fresh installs". The build in step 6 had left the base image's layers in that store. Docker's build cache, which keeps pieces of earlier builds to make later builds faster, was holding them. Removing the images removed their names, not their layers, which were still on disk. So "pull" found everything locally, and "load" found most of it too.
So I repeated both, on the same box, after the demo. This block was added after the results, and it is not in the recordings: I started it over SSH. Its script, cold_files/cold_image_again.sh, also empties the build cache. Blobs are the stored files that layers and image descriptions are kept in. Before each timed run the script checks that containerd's store holds 0 blobs, and all ten runs started from 0.
Loading my 209.5 MB image file took 8.66 seconds (8.64 to 8.92 over five runs). Pulling the base image, 43.3 MB as Docker reports it, took 2.28 seconds (2.25 to 2.40). The first, unclean try gave 7.40 and 0.52 seconds; I keep those numbers in the results, labelled, and I do not use them.
So on a new machine, getting the image can cost several times the start itself. A full pull of my own image from a registry far away would cost more. I did not measure that, because I have no registry for my image, and making one would be a new resource.
Back to the service once it has started. In arm A, the client sends 100 real requests, one at a time, as soon as the port opens. How does the first compare with the rest?

The first request took 4,680 microseconds, and the second 1,849 (medians of eight runs). Request 1 was slower in all eight runs. The median of the eight paired differences was 2,981 us (a median of differences, so not 4,680 minus 1,849). Request 1 ranged 4,186 to 5,770 us and request 2 ranged 1,180 to 2,620. The median of the eight per-run ratios was 2.84. Inside the handler, the model's own call took 2,946 us the first time and 725 us the second. So most of the extra was the model's first predict_proba call. The rest was the first trip through the web framework's code.
Why is the model's first call slow? I did not measure inside it. One possible reason is that the first call prepares things the later calls reuse. Examples are code that the processor reads for the first time, and helper objects that scikit-learn builds once. Lesson 3 saw the same first-call effect in its maximum .
Request 2 was slower than request 100 in all eight runs, 1,849 against 1,190 us. But the two sets of eight overlapped, so by the rule they could not be told apart. My guess had been that the first request would cost at least three times the second. It cost 2.84 times, so that guess was wrong, though close.
If the first call costs extra, let someone else make it. A warm-up sends the service a request, or a model call, before real users arrive.

With one warm-up request (C1), the first real request took 1,655 us instead of 4,680, told apart. The warm-up request itself took 4,912 us, about what the first request took in A. So the extra cost did not disappear. The service paid it, before any user was let in.
Fifty warm-up requests (C50) gave 1,195 us. Against one warm-up, the difference could not be told apart. So one request did most of the work here.
Fifty calls inside the process (I50) gave 2,056 us. These warmed the model and SQLite, but not the web framework's own path. My guess was that I50 would land between A and C50, told apart from both. It was below A, told apart, but against C50 it could not be told apart. So that guess was wrong. The path's own first-time cost was too small to show against the noise here.
A warm-up only helps if no user arrives before it ends. Something must hold traffic back. That is the readiness check.

In arms C1 and C50, my service answers GET /ready with status 503, "not available", until its warm-up is done, and with 200 after. The client plays the : it asks every 5 ms and sends real requests only after a 200.
, a common system for running containers, describes the same idea. Its documentation says: "Readiness probes determine when a container is ready to accept traffic." It also says this is useful when an application must do "time-consuming initial tasks, such as establishing network connections, loading files, and warming caches."
The important part is what the check asks. Some load balancers check only that a connection succeeds; for AWS's Network Load Balancer, its documentation says "The default is the TCP protocol." In C50 the port opened 76 ms before the service was warm. With one warm-up (C1) the gap was about 16 ms (ready 1,506.3 minus port 1,490.5, a difference of medians). A check that only opens a connection would have sent users in during those milliseconds, straight into the slow first call.
A warm-up is not free. It makes the start a little longer, and it puts the cost of the first call on the service instead of on a user.

Fifty warm-ups delayed the first real request by 69.5 ms, compared with A, told apart. Fifty in-process calls delayed it by 50.1 ms, also told apart. Against a start of 1,499 ms, both are small.
What did they buy? About 3 milliseconds on one request. That sounds like very little, and on its own it is. But it is the request that a user waits for after a deploy, a restart or a scale-up. And if a service starts many times a day, for example when it scales up and down with traffic, every start has a first user.
I would still use one warm-up request behind a readiness check. It is cheap, it covers the whole request path, and here it removed most of the extra.
So far one request came at a time. Real users arrive at random, many at once. Arms F and G start the service and then send open-loop traffic: 539 random arrivals a second for 3 seconds. That is half of lesson 5's measured capacity S, 1,077.1 a second. counts from the moment each request was meant to arrive.

None of the six comparisons was told apart. In the first second, F's median p99 was 6.51 ms and G's 8.19 ms. In the third second they were 5.37 and 5.01. One run in each arm reached 21 to 23 ms; the next highest were 15.4 (F) and 14.3 (G).
I had guessed that G's first-second maximum would be lower than F's, and that F's first second would be at least three times worse than its third. Both guesses were wrong. F's first-second p99 was only 1.40 times its third.
Why? One possible reason is size. At 539 arrivals a second, about 540 requests land in the first second. The first call's extra 3 ms touches only the few requests queued behind it. The random bunching of arrivals, which made the queue grow and shrink, moved the p99 by more than that. This does not mean warm-ups never matter under traffic. A bigger model, a slower first call or more traffic at the start could change it. Here, at this rate, I could not see it.
I wrote eighteen guesses into the lab before it ran. Each is quoted here word for word from the docstring. Fourteen were right and four were wrong.
"H1 A: launch -> python_ready between 15 and 80 ms." Right: 18.2 ms.
"H2 A: the sklearn import step (after numpy and pandas) is the largest single step of the start." Right, in all 8 runs.
"H3 A: the import steps (stdlib to uvicorn) are at least 70% of launch -> port open." Right: 96.9%.
"H4 A: unpickle under 20 ms." Right: 1.22 ms.
"H5 A: launch -> first answer between 1.0 and 4.0 s." Right: 1.499 s.
"H6 A: request 1 at least 3 x request 2 (median of the runs), and request 100 within 1.2 x request 2." Wrong: request 1 was 2.84 times request 2. Request 100 was 0.73 times request 2.
"H7 B: launch -> first answer at least 2 x A's, told apart." Right: 2.81 times.
"H8 C1: the first real request within 1.5 x A's request 2 (one warm-up removes most of the first-request cost)." Right: 1.04 times.
"H9 C50 vs C1: the first real request cannot be told apart." Right.
"H10 I50: the first real request below A's, told apart, and above C50's, told apart (the path is still cold)." Wrong: below A's, told apart, but against C50's it could not be told apart.
The full lab needs a rented box and . The demo, cold_demo.py, runs on your own computer and needs neither. It times a fresh Python process, step by step, the same way the lab does.

I wrote the demo's design into its docstring after the lab's design and before the demo first ran.

Before you run this lab. The demo uses lesson 2's Python environment, with the same library versions as the lab's box. If you made it for lessons 2 to 5, use it again. If not, make it once, inside the scripts/labs/serving/examples folder:
python3 -m venv venv-sv
source venv-sv/bin/activate # on Windows: venv-sv\Scripts\activate
pip install -r lat_requirements.txt
I used Python 3.13. Then run python ../../features/fetch_data.py once, which downloads the shop data. Now run python cold_demo.py. It needs no Docker, no web server and no cloud account. It writes a model file into a temporary folder and deletes the folder at the end.
The first line it prints says the times come from your machine. Your disk, your processor and whatever else is running all move them. The demo cannot empty your page cache without administrator rights, so every process after the first starts with its files already in memory.
This box holds the real results from the lab. It needs nothing but Python, so it runs in your browser. It has no model and no clock: it only looks up what the box measured.
Press Run. It shows arm A, the plain case, and then the arm you pick, so you can always compare. Try START = "B" for an emptied page cache, "D2" for the compiled container, or "C1" for one warm-up. The differences it prints are one median minus another, so they can differ a little from the paired medians on the slides.
The report script writes this box from the lab's raw files, runs it for all eight ways, and checks that each run prints the lab's numbers.
The lab has one file that runs on my laptop, a few small files that run on the box, and the step script you saw above.
cold_start.py, on my laptop, holds the design in its docstring, written before any run, with dated notes added after. Its commands are check (lesson 2's files match by sha256), status (reads the box's serial console), collect, cost and factcheck. collect calls cold_stats.py, which gzips the box's raw files into results/cold-raw/, passes every log line through cold_redact.py, and does all the arithmetic.
cold_files/cold_service.py is lesson 2's service, rewritten so that every step of starting it reads the clock. It adds GET /ready, which says 503 until a warm-up is done, and GET /startup, which hands the clock readings to the client after the run.
cold_files/cold_client.py starts the service, waits for the port or for , sends the requests and stops the service. It uses only the standard library, for the reason on the machine slide. Its open-loop mode is lesson 4's design, rewritten without numpy.
Here is the order I would follow, using only what this lab did.

The chart starts by measuring, because the right fix depends on which step is biggest. Here are the same steps in words.
Time one start, step by step. Read one clock at each step, as cold_service.py does, or run python -X importtime. Here the imports were 96.9% of the time until the port opened.
If imports are the biggest part, import less, or start less often. Import only what the service uses. Start the service once and keep it running, rather than once per request or per job. Here, one start cost 1.5 seconds, and one request about 1 millisecond.
Put compiled files in the image. Run python -m compileall on the standard library when you build it. Here that saved 200 ms on every container start.
Warm up, then say ready. Send one real request through the whole path, and answer the readiness check with 200 only after it. Here it took the first real request from 4,680 to 1,655 us.
Make the ask the right question. A check that a port is open is not enough. Here the port opened 76 ms before the service was warm with 50 warm-ups (C50), and about 16 ms before with one (C1).
Cold starts matter when a service starts often. If it scales up and down with traffic, or runs as a short job for each batch, every start pays the import time again. Here that was 1.5 seconds on a warm machine and 4.2 seconds on a cold one.
They matter for the first users after a deploy. Every new copy has a first request. Here it cost about 3 ms more than the next, unless a warm-up paid it first.
They matter most on a new machine. Then the files are not in memory, and the image may not be there at all. Here an emptied page cache made the start 2.81 times as long, and loading the image took 8.66 seconds before the start could even begin.
They matter less for a service that starts once and runs for weeks. One slow start a week is not worth much effort. Under steady traffic here, I could not even tell a warmed start from a cold one.
Do not judge the model file first. It is easy to blame the model for a slow start. Here the model took 2.0 ms of a 1,499 ms start. Measure before you shrink the model.

One small model. This model loads in 2 ms. A model of several gigabytes would make loading the model the biggest step, and the answer would change. I did not measure one.
One core, and the client on it. Starting a service is mostly one thread of work, so I expect the shape to hold on bigger machines. But I measured only this one.
A cold cache, not a cold disk. Emptying the page cache simulates a machine that has not read these files recently. A new disk made from a snapshot can be much slower on first read. Pulling a large image from far away can take much longer too.
Two builds of Python. The bare and container Pythons differ (Clang against GCC), so D against A mixes the container and the build. D2 against D does not. But D2's layer also compiled 84 site-packages files that had none (most likely pip's own, which the service does not import; I did not check which).
An emptied page cache is not a reboot. drop_caches keeps pages that running programs have mapped. About 154 MB stayed cached after each drop, including the programs dockerd and containerd. So arm E was cold for the image's files, but itself was warm, unlike just after a reboot.
Labelled additions and repeats. The clock check was changed after the results, as the clock slide explains. The image timings were repeated after the results with a cleaner start, as the image slide explains. Both changes are in the lab's notes, and the first numbers are kept.

If you take one thing to work on Monday, time one start of your model service. Run python -X importtime -c "import your_service" and sort by the cumulative column. You will probably find, as I did, that the libraries cost far more than the model.
Then look at how your service says it is ready. If the only checks that a port is open, add a /ready route that answers 200 only after one real warm-up request has gone through.
The one idea to keep: a service is not ready when its port opens. It is ready when someone has already paid for the first request, and that someone should be the service.
4 questions - Score 80% to pass
In arm A, a bare process with its files in memory, which part took most of the time before the port opened?
Why did the image with a compiled standard library (D2) start faster than the plain image (D)?
What did one warm-up request behind a /ready check (C1) do to the first real request?
Under 539 random arrivals a second, what did the lab find for F (no warm-up) against G (50 warm-ups and /ready)?
/readyTo keep the files warm for a warm arm, the lab first starts and stops the same kind of service once, without timing it. Then it waits until the core is at least 95% idle.
The rule, written before the runs, is lesson 3's rule. Two arms are "told apart from run-to-run noise" only if all eight paired differences have the same sign, and the two sets of eight values do not overlap. I claim that one arm beat another only when I compared those two arms directly.
HEREBSCcopySCIPexport AWS_DEFAULT_REGION=us-east-1 AWS_PAGER=""
NAME=ai-research-course-cold
HERE=$PWD
STATE=${COLD_STATE:-$HOME/lab-data/serving/cold}
KEY=$HOME/lab-data/serving/keys/$NAME.pem
K=${K:-8}
mkdir -p "$STATE" "$HOME/lab-data/serving/keys"
SSHO=(-o LogLevel=ERROR -o StrictHostKeyChecking=accept-new -o "UserKnownHostsFile=$STATE/known_hosts")
[ -f "$STATE/id" ] && ID=$(cat "$STATE/id")
[ -f "$STATE/ip" ] && IP=$(cat "$STATE/ip")
[ -f "$STATE/sg" ] && SG=$(cat "$STATE/sg")
B=/home/ubuntu/cold
S=(ssh -i "$KEY" "${SSHO[@]}" ubuntu@"$IP")
C=(scp -q -i "$KEY" "${SSHO[@]}")
What it costs. A c7g.medium costs $0.0363 an hour on demand in us-east-1, from AWS's price list, and AWS bills it by the second. My measuring box was on for 0.444 hours, from launch to terminated, which is $0.0161 for compute. An earlier box, on 2026-10-04, ran the set-up steps but took no timing and was deleted the same day: 0.068 hours, $0.0025. So the whole lesson cost $0.0186. The disks were deleted with the boxes. These numbers are in results/cold-cost.json and results/cold-cost-setup-box.json. The schedule itself took 14 minutes.
What the recordings hide. Every line passed through a small filter, cold_redact.py, a copy of lesson 5's. It hides my account number, every IP address, every name AWS makes up for a resource, host names, my user name and anything from a key file. Where you see <ip> or i-<id>, your terminal shows the real value. Lines that start with + are the shell printing each command just before it runs it.
Keep the key file outside any git folder. Mine lives in ~/lab-data/serving/keys/, and the last step deletes it.
Step 3: rent the box. First the script asks AWS's parameter store (SSM) for the id of the newest Ubuntu 24.04 image for arm processors. So you never copy an id that has gone out of date. Then run-instances asks for one c7g.medium with a 16 GiB disk that is deleted with the box, with the tags on both.

The box started in the "pending" state, which is normal for the first few seconds. These are the two commands, the image lookup and the launch:
AMI=$(aws ssm get-parameter --query Parameter.Value --output text \
--name /aws/service/canonical/ubuntu/server/24.04/stable/current/arm64/hvm/ebs-gp3/ami-id)
ID=$(aws ec2 run-instances --image-id "$AMI" --instance-type c7g.medium --count 1 \
--key-name ai-research-course-cold --security-group-ids "$SG" \
--block-device-mappings '[{"DeviceName":"/dev/sda1","Ebs":{"VolumeSize":16,"VolumeType":"gp3","DeleteOnTermination":true}}]' \
--tag-specifications \
"ResourceType=instance,Tags=[{Key=Project,Value=ai-research-course},{Key=Lesson,Value=cold},{Key=Name,Value=ai-research-course-cold}]" \
"ResourceType=volume,Tags=[{Key=Project,Value=ai-research-course},{Key=Lesson,Value=cold}]" \
--query "Instances[0].InstanceId" --output text)
The script also writes ID and SG into small files in ~/lab-data/serving/cold/, so the later steps can read them back.
"${S[@]}" "curl -LsSf https://astral.sh/uv/install.sh | sh > /dev/null 2>&1 && \
~/.local/bin/uv python install 3.13.15 > /dev/null 2>&1 && \
~/.local/bin/uv venv --allow-existing --python 3.13.15 $B/venv > /dev/null 2>&1"
"${C[@]}" "$HERE/examples/lat_requirements.txt" ubuntu@"$IP":$B/
"${S[@]}" "~/.local/bin/uv pip install --python $B/venv/bin/python -q -r $B/lat_requirements.txt"
"${S[@]}" "sudo apt-get update -q > /dev/null && \
sudo DEBIAN_FRONTEND=noninteractive apt-get install -y -q docker.io docker-buildx > /dev/null && \
sudo usermod -aG docker ubuntu && docker --version"
S and C are short names for the ssh and scp commands with the key, from the block above. The full list of files to copy is in the copy step.
Step 6: build the two container images. The first image, cold-svc:lab, follows packaging lesson 5's slim pattern. It starts from the official python:3.13.15-slim image, pinned by its digest, a fingerprint of the image's exact contents. It adds the same pinned packages, the model, the SQLite file and the service.
An image is built in layers, each a set of file changes stacked on the one below. The second image, cold-svc:pyc, adds one layer: Python's standard library compiled to .pyc files with compileall. That compileall over /usr/local/lib/python3.13 also compiled 84 site-packages files that had none. They are most likely pip's own, which the service does not import; I did not check which.

The compiled image is 14 MB larger, as Docker counts it. This is the command:
ssh -i ~/lab-data/serving/keys/ai-research-course-cold.pem "${SSHO[@]}" ubuntu@"$IP" \
'cd cold && cp lat_requirements.txt model.pkl store.sqlite cold_rows.json cold_service.py ctx/ && \
docker build -q -t cold-svc:lab -f Dockerfile ctx > /dev/null && \
docker build -q -t cold-svc:pyc -f Dockerfile.pyc ctx > /dev/null && \
docker images --format "{{.Repository}}:{{.Tag}} {{.Size}}"'
Step 7: stop the timers and look at the box. Ubuntu runs small jobs on a timer. One of those starting during a run would land in the results, so the script stops every timer first. Then it prints what the box is, saves a fuller record, and lists every installed package with its version.

The load average, Linux's running average of how many programs want the core, was still high from step 6. So the schedule's first job is to wait until it falls to 0.10. These are the commands that stop the timers and look:
ssh -i ~/lab-data/serving/keys/ai-research-course-cold.pem "${SSHO[@]}" ubuntu@"$IP" \
'for t in $(systemctl list-units --type=timer --state=active --no-legend --plain | cut -d" " -f1); do
sudo systemctl stop "$t"; done; sudo systemctl stop unattended-upgrades.service
systemctl list-timers --no-pager | tail -2'
ssh -i ~/lab-data/serving/keys/ai-research-course-cold.pem "${SSHO[@]}" ubuntu@"$IP" 'lscpu | head -12; nproc; uptime'
Step 8: start the schedule, then leave. The schedule starts in the background with setsid nohup, so it keeps running after SSH disconnects. Then I log out, and nobody logs in until it ends.

The log's first line appeared, and the SSH session closed. This is the command:
ssh -i ~/lab-data/serving/keys/ai-research-course-cold.pem "${SSHO[@]}" ubuntu@"$IP" \
"cd cold && (setsid nohup bash cold_schedule.sh $K > raw/schedule.log 2>&1 < /dev/null &); \
sleep 2; cat raw/schedule.log"

I ran this step twice. The first fetch came right after the demo. Then I repeated the image timings (see the slide on images), and this recording is the second fetch, which brought those files too. Rsync only adds and updates files, so running it twice is safe. This is the copy command:
rsync -a -e "ssh -i $KEY ${SSHO[*]}" ubuntu@"$IP":cold/raw/ ~/lab-data/serving/cold/raw/
Step 14: delete the box, the key and the door rule, and check. A box you forget keeps costing money. The script deletes the box, waits until AWS says "terminated", then deletes the key pair and the security group. Then it asks AWS again.

Every count printed 0 and the box printed terminated. I asked AWS once more afterwards, from my laptop, and all four counts were still 0. These are the teardown commands:
aws ec2 terminate-instances --instance-ids "$ID" --query "TerminatingInstances[0].CurrentState.Name" --output text
aws ec2 wait instance-terminated --instance-ids "$ID"
aws ec2 delete-key-pair --key-name ai-research-course-cold --query Return
until aws ec2 delete-security-group --group-id "$SG" --query Return 2>/dev/null; do sleep 10; done
aws ec2 describe-instances --instance-ids "$ID" --query "Reservations[0].Instances[0].State.Name" --output text
aws ec2 describe-key-pairs --filters "Name=key-name,Values=ai-research-course-cold" --query "length(KeyPairs)"
aws ec2 describe-security-groups --filters "Name=group-name,Values=ai-research-course-cold" \
--query "length(SecurityGroups)"
aws ec2 describe-volumes --filters "Name=tag:Lesson,Values=cold" --query "length(Volumes)"
aws ec2 describe-instances --filters "Name=tag:Project,Values=ai-research-course" \
"Name=instance-state-name,Values=pending,running,stopping,stopped,shutting-down" \
--query "length(Reservations[].Instances[])"
rm -f ~/lab-data/serving/keys/ai-research-course-cold.pem
So every start compiled again whichever standard library modules it imported that had no .pyc file, and threw those copies away.
The bare Python on the box held 206 compiled standard library files after its runs, which is roughly how many modules this service imports. Image D2 compiled all 633 ahead of time with python -m compileall. Its compileall over /usr/local/lib/python3.13 also compiled 84 site-packages files that had none. They are most likely pip's own, which the service does not import; I did not check which. Its imports took 1,589 ms, and its first answer came 200 ms earlier than D's, told apart. Python's tutorial says it plainly: "the only thing that’s faster about .pyc files is the speed with which they are loaded."
One honest limit. The bare Python and the container's Python are both 3.13.15, but they are different builds, made with different compilers. So some of D against A may come from the build, not from the container. D2 against D is a much cleaner comparison: the same image plus one layer. That layer holds the compiled standard library and the 84 extra site-packages files, so the 200 ms belongs to the layer as a whole.
"H11 D: launch -> first answer above A's, told apart, by at least 0.3 s (median difference)." Right: 539.5 ms.
"H12 D2: launch -> first answer below D's, told apart." Right: 199.9 ms earlier.
"H13 E: launch -> first answer above D's, told apart." Right: 3,185 ms later.
"H14 G: first-second max below F's, told apart." Wrong: it could not be told apart.
"H15 F: first-second p99 at least 3 x third-second p99 (median of the runs)." Wrong: 1.40 times.
"H16 every returned score equals the box's single-row score to the bit, in every arm." Right: 32,379 of 32,379. This is a check that every arm served the right model on the right rows, not a finding. The box's reference matched lesson 2's stored scores on 26,818 of 26,851 rows, with a largest gap of 1.1e-16.
"H17 load of the image from the file (cold cache) between 5 and 60 s." Right: 8.66 s, judged on the clean repeat added after the results.
"H18 -X importtime, warm cache: the largest top-level cumulative import is sklearn.ensemble." Right, in all 5 warm runs.
r"""What does the first request pay? Time a model's cold start, piece by piece, on YOUR machine.
Lesson 6 of 'Serving and Inference Basics'. It uses lesson 2's Python environment, made once inside this folder
(Python 3.13 is what the lab used):
python3 -m venv venv-sv
source venv-sv/bin/activate # Windows: venv-sv\Scripts\activate
pip install -r lat_requirements.txt # numpy, pandas, pyarrow, scikit-learn, threadpoolctl, ...
It also needs the shop data from the features chapter: run python ../../features/fetch_data.py once first. Then:
python cold_demo.py # train, then print parts 2 to 4 below, numbered 1 to 3
python cold_demo.py --save out.json # and save the numbers
If a package is missing, it prints one line saying why and stops. It needs no Docker, no web server and no
administrator rights. It writes a model file and a few rows into a temporary folder, deletes that folder at the
end, and writes nothing else except out.json if you ask.
Every time it prints comes from YOUR machine, as it is right now: other programs, your disk and whether these files
were read recently all move them. Read them for what they show about your machine.
Design, written 2026-10-04 after the lab's design (cold_start.py) and before this file first ran.
1. Train the features chapter's model (it must score test AP 0.5450), as lesson 5's wrk_demo.py does, and save it
with pickle, with the first 100 test rows, into a temporary folder.
2. Five fresh Python processes, one after another (this same python, OMP_NUM_THREADS=1). Each reads the clock on
its first line, then imports numpy, pandas and sklearn.ensemble, reads the model file, unpickles it, and calls
predict_proba on one row 100 times, each on the next test row. The clock is time.monotonic_ns(), read by this
process just before it starts each child, and by the child after each step. Printed: the median of the five
for each step, and for predict calls number 1, 2 and 100.
3. python -X importtime of the same imports, once: the five biggest top-level imports by their total time.
4. A warm-up: five more fresh processes, the same steps, but ONE predict_proba call on a row that is not a test
row first (the warm-up), then the 100 calls. Printed: call number 1 with and without the warm-up.
The page cache cannot be emptied without administrator rights, so every child here starts with its files
already read recently (the first child may not): what the lesson calls a warm cache.
Author: Roni Das
Created: 2026-10-04
"""
import importlib.util
import json
import os
import pickle
import statistics
import subprocess
import sys
import tempfile
import time
from pathlib import Path
NEEDED = ("numpy", "pandas", "pyarrow", "sklearn", "threadpoolctl")
missing = [m for m in NEEDED if importlib.util.find_spec(m) is None]
if missing:
sys.exit(f"cold_demo.py needs {', '.join(missing)}: make the environment in its docstring "
f"(pip install -r lat_requirements.txt), then run it with that environment's python.")
RUNS = 5
CHILD = r'''
import time
T = [("python ready", time.monotonic_ns())]
import sys
def mark(k):
T.append((k, time.monotonic_ns()))
import json, pickle
from pathlib import Path
from time import perf_counter_ns
import numpy as np
mark("import numpy")
import pandas
mark("import pandas")
import sklearn.ensemble
mark("import sklearn")
d = Path(sys.argv[1])
blob = (d / "model.pkl").read_bytes()
mark("read model file")
model = pickle.loads(blob)
mark("unpickle")
X = np.load(d / "rows.npy")
calls = []
if sys.argv[2] == "warm":
model.predict_proba(np.load(d / "warm.npy"))
for i in range(100):
a = perf_counter_ns()
model.predict_proba(X[i:i + 1])
calls.append(perf_counter_ns() - a)
print(json.dumps({"marks": T, "calls": calls}))
'''
def build():
"""The features chapter's model and its test rows (lesson 5's wrk_demo.py build step)."""
import numpy as np
from sklearn.metrics import average_precision_score
here = Path(__file__).resolve().parent
sys.path.insert(0, str(here.parents[1] / "features"))
import task
from what_a_feature_is import HAND_COLS, hgb, joined
ev = task.load_events()
lab_tr, _, lab_te = task.splits(ev)
tr = joined(ev, lab_tr, task.TRAIN_CUTOFFS)
te = joined(ev, lab_te, task.TEST_CUTOFFS)
model = hgb(0).fit(tr[HAND_COLS].to_numpy(float), tr["label"].to_numpy())
X = np.ascontiguousarray(te[HAND_COLS].to_numpy(np.float64))
p = model.predict_proba(X)[:, 1]
y, cut = te["label"].to_numpy(), te["cutoff"].to_numpy()
ap = float(np.mean([average_precision_score(y[cut == c], p[cut == c]) for c in np.unique(cut)]))
assert round(ap, 4) == 0.5450, ap
warm = np.ascontiguousarray(tr[HAND_COLS].to_numpy(np.float64)[:1])
return model, X, warm, ap
def child(folder: Path, mode: str) -> dict:
env = dict(os.environ, OMP_NUM_THREADS="1")
t = time.monotonic_ns()
r = subprocess.run([sys.executable, "-c", CHILD, str(folder), mode], capture_output=True, text=True, env=env)
if r.returncode:
sys.exit(r.stderr[-800:])
d = json.loads(r.stdout)
marks = [("start", t)] + [tuple(m) for m in d["marks"]]
steps = {marks[i][0]: (marks[i][1] - marks[i - 1][1]) / 1e6 for i in range(1, len(marks))}
return {"steps_ms": steps, "calls_us": [c / 1000 for c in d["calls"]]}
def importtime(folder: Path) -> list:
code = "import json, pickle; import numpy; import pandas; import sklearn.ensemble"
r = subprocess.run([sys.executable, "-X", "importtime", "-c", code], capture_output=True, text=True,
env=dict(os.environ, OMP_NUM_THREADS="1"))
top = []
for line in r.stderr.splitlines()[1:]:
if not line.startswith("import time:"):
continue
_, cum, name = line[len("import time:"):].split("|")
if not name[1:].startswith(" "): # one space only: a top-level import (Python's start-up or the -c line)
top.append((int(cum) / 1000, name.strip()))
return sorted(top, reverse=True)[:5]
def main() -> None:
try:
load = f"{os.getloadavg()[0]:.2f}"
except (AttributeError, OSError):
load = "not available on this system"
print(f"Every time below was measured on THIS machine ({os.cpu_count()} cores, 1-minute load {load}).")
import numpy as np
model, X, warm, ap = build()
print(f"model trained: test AP {ap:.4f}; {len(X):,} test rows")
with tempfile.TemporaryDirectory() as tmp:
folder = Path(tmp)
(folder / "model.pkl").write_bytes(pickle.dumps(model, protocol=5))
np.save(folder / "rows.npy", X[:100])
np.save(folder / "warm.npy", warm)
cold = [child(folder, "plain") for _ in range(RUNS)]
top = importtime(folder)
hot = [child(folder, "warm") for _ in range(RUNS)]
med = lambda xs: statistics.median(xs) # noqa: E731
steps = list(cold[0]["steps_ms"])
print(f"\n1. a fresh Python process, step by step: median of {RUNS} processes, milliseconds")
for k in steps:
print(f" {k:16s} {med([c['steps_ms'][k] for c in cold]):9.1f}")
total = med([sum(c["steps_ms"].values()) for c in cold])
print(f" {'all of it':16s} {total:9.1f} (start to model loaded; median of the {RUNS} totals)")
calls = {n: med([c["calls_us"][n - 1] for c in cold]) for n in (1, 2, 100)}
print(" predict_proba on one row, microseconds (us): "
+ ", ".join(f"call {n} {v:.1f}" for n, v in calls.items()))
print("\n2. python -X importtime: the five biggest top-level imports, total ms (one run)")
for ms, name in top:
print(f" {name:16s} {ms:9.1f}")
w1 = med([h["calls_us"][0] for h in hot])
print("\n3. the first real predict_proba call, median of 5 processes, us")
print(f" no warm-up {calls[1]:9.1f}")
print(f" after one warm-up call {w1:9.1f}")
if "--save" in sys.argv:
out = {"load": load, "cores": os.cpu_count(), "test_ap": ap, "runs": RUNS,
"steps_ms": {k: med([c["steps_ms"][k] for c in cold]) for k in steps}, "total_ms": total,
"calls_us": {str(n): v for n, v in calls.items()}, "importtime_top": top, "warm_call1_us": w1}
Path(sys.argv[sys.argv.index("--save") + 1]).write_text(json.dumps(out, indent=1))
if __name__ == "__main__":
main()
You saw the demo running on the lab's box in step 12, after the schedule. That exact run is stored in results/cold-demo-run.txt, and the report checks that every number it printed matches the numbers it saved.
On the box it told the same story as the lab. Starting Python and loading the model took 1,014 ms, and scikit-learn's import was 721 ms of it. Reading and unpickling the model took about 1 ms. The first predict_proba call took 3,004 us, the second 618 us, and after one warm-up call the first real call took 623 us.
Its numbers are a little lower than the lab's start, because the demo does not import FastAPI or uvicorn and does not start a web server.
And here is the same demo in VS Code on my laptop, run with python cold_demo.py in the examples folder with venv-sv active.

This run is from my laptop, not the box. It is a different machine, with 10 cores, and it was heavily loaded when I ran it: the 1-minute load was 14.25. So these numbers are not to be compared with the box's numbers in any way. Look only at the shape, which is the same as on the box. scikit-learn's import was the biggest step, 532.5 of the 729.2 ms from start to model loaded. Reading and unpickling the model took 0.4 ms. The first predict_proba call took 1,801.4 us, about nine times the second call's 197.5 us. After one warm-up call, the first real call took 199.5 us.
/readycold_files/cold_run_once.sh warms or empties the page cache, waits for an idle core, runs the client, and records the load and the system log. cold_files/cold_schedule.sh runs everything in order. cold_files/cold_image_again.sh is the repeat of the image timings, added after the results. Dockerfile and Dockerfile.pyc build the two images.
cold_report.py does not import any of the above. It reads the small copy of the raw files with its own code and recomputes every number. It also checks the guesses, the demo's stored run, the fact-check, the teardown, the redaction, and it writes the playground.
Keep files close to the machine. Files in memory started the service 2.81 times faster than files on disk. An image already on the machine starts at once; one that must be loaded took 8.66 seconds here.