Serving And Inference

Workers and Threads: Processes, Threads, or Both on One Core?

0 of 29 complete

0%

Contents

Back|Serving And InferenceWorkers and Threads: Processes, Threads, or Both on One Core?
1/29
86 min left
  1. Home
  2. AI Engineering: Data, RAG and Agents
  3. Serving and Inference Basics
  4. Workers and Threads: Processes, Threads, or Both on One Core?
Prerequisites
Batching Requests: Does It Help a Small CPU Model?requiredTail Latency: What Sits in the Slowest 1% of RequestsrequiredLatency Anatomy: Where the Time in One Prediction Request Goesrequired
Related Topics
Model Signatures: The Right Numbers in the Wrong Shape, and What a Schema Check CatchesPackaging, Registry and Versioning
1 of 29
Previous lesson
Batching Requests: Does It Help a Small CPU Model?

System Design

  • Foundation
  • Intermediate
  • Advanced
  • Capstone

AI Engineering

  • Foundation
  • Data, RAG and Agents
  • Evaluation, LLM Ops and Security

systemdesign.academy

  • Home
  • Glossary
  • Interview prep
  • Reviews
  • About
  • Privacy
  • Terms

Three People and One Laptop

Let me start in a small office.

Three colleagues share one table and one laptop. Each of them has their own notebook with their own notes. Only one person can type on the laptop at a time, so they pass it round.

A flat illustration of three colleagues at one table, two women and a man, talking and writing in notebooks around a single open laptop, with charts on paper and a plant. Below the picture: three colleagues and one laptop; a fourth person makes the table busier, not the laptop faster; asking three people to type one short line together only adds the time it takes to pass the laptop round.

Now the manager wants more work done at the same table, with the same one laptop, and has two ideas.

Idea one: hire a fourth person. The fourth person brings a fourth notebook and takes a turn at the laptop. But the laptop does not get faster. Each person now waits a little longer for their turn, and passing the laptop round takes time too. If the work is all typing, four people finish no more than three did. They may finish a little less.

Idea two: ask three people to type one short line together. Someone has to call the other two over and wait for them to sit down. Then each gets a third of the line, and everyone waits for the last one to finish. For one short line, that costs much more than typing it alone. For a long report on a table with three laptops, it could be a good idea. On one laptop it never is.

A prediction service has the same two ideas. It can run more worker processes, like more people. Or it can let the model use more threads inside one call, like several people typing one line. In this lesson I measure both, and both together, on a rented machine with exactly one core: the one laptop of the story.

Where This Lesson Starts

This is lesson 5 of the chapter on serving a model. Lesson 2, latency anatomy, took one request apart. Lesson 3, tail latency, looked at the slowest requests. Lesson 4, batching requests, let requests share one model call. All three ran the service as one process, with one thread for the model.

This lesson changes those two numbers. I use the same model, the same service code and the same kind of rented machine as lessons 2 to 4, so I do not explain them again. In short: the model is the features chapter's gradient boosted tree model. It scores one customer of a real online shop, and it scored a test AP of 0.5450. The service is a small FastAPI web service run by uvicorn, and it reads the customer's six features from SQLite. If any of that is new, please read lesson 2 first.

One thing about the machine changes the question. My AWS account is allowed only one virtual CPU of this kind at a time. I checked whether a cheaper "burstable" machine would give me two. It would not: every t4g size has 2 vCPUs, and the same limit covers that family too. So the box has one core. The question in this lesson becomes: on one core, do more workers or more threads help, or do they only get in each other's way?

What a machine with more cores would add, I say in words only. I give no numbers for a machine I did not measure.

The Words You Need First

Please read this slide slowly if any word is new. Every slide after it uses these words.

A hand-drawn grid of twelve cards, each a word with its meaning: core, process, thread, worker (W), OpenMP thread (T), GIL, context switch, oversubscription, RSS, Pss, closed loop and open loop. Below: request, latency, p50, p99, keeps up and microsecond (us) mean what they meant in lessons 2 to 4; 1 ms = 1,000 us.

A core is one part of the processor that runs one stream of instructions at a time. It is the laptop in the story. My box has exactly one. A machine with eight cores can run eight things at the same moment.

A process is a running program with its own memory. Two processes cannot read each other's memory, like two people with their own notebooks. A thread is a line of work inside a process. Threads in one process share its memory, like two people reading the same notebook.

A worker is one copy of the web service, running as its own process. uvicorn, the server from lesson 2, starts W of them when you pass --workers W. Each worker loads the model for itself. Inside a worker, one event loop does the work: a single thread that takes the next ready job and runs it to the end before taking another. A job is something like reading a request or answering it.

An OpenMP thread is a helper thread that scikit-learn starts inside one call to predict_proba, to share the call's rows between helpers. OpenMP is a standard way for C code to do this. I set how many with the OMP_NUM_THREADS setting, and I call that number T. A setting of this lab is written W x T. So 4x1 means four workers with one thread each, and 1x4 means one worker whose model calls use four threads.

The GIL, the global interpreter lock, is a lock inside Python. Only the thread that holds it may run Python code. Code written in C can let go of it while it works, so that other threads can run Python in the meantime.

A is the operating system stopping one thread on the core and starting another. means more threads want to run than there are cores to run them, so they wait and switch. On a one-core box, every setting except 1x1 is oversubscribed.

The Headline: On One Core, Nothing Beat One Worker With One Thread

Here is the main result first. Each column is one setting. Each dot is one round: the most answers per second that setting gave, with 16 connections each sending a new request the moment its answer arrived.

A dot chart with one column per setting, 1x1, 1x2, 1x4, 2x1, 2x2, 2x4, 4x1, 4x2, 4x4, 1x8 and 4x8, five dots per column for the five rounds, against answers per second from 0 to 1,200, with a dashed line labelled 1x1 median. The 1x1, 2x1 and 4x1 dots sit near the line, around 1,050 to 1,070; 1x2 near 700; 1x4 near 490; 2x2 near 460; 4x2 three dots near 380 and two near 464; 2x4 near 290; 1x8 near 306; 4x4 near 250 with one dot near 240; 4x8 four dots near 190 and one near 164. Below: median of 5 rounds, 1x1 1,071, 2x1 1,049, 4x1 1,049, 1x4 486, 4x4 253, lowest 4x8 192; the 4x settings started 4 workers, but only 2 or 3 got connections in most rounds; compared with 1x1 round by round, 10 of 10 gave fewer answers, told apart, and 0 gave more. Caption: one core does one thing at a time; more processes and threads only took turns on it.

One worker with one thread gave the most: 1,071 answers a second. Every one of the other ten settings gave fewer, and the difference held in all five rounds.

More workers cost a little. Two workers gave 1,049 a second, and four started workers also 1,049, about 0.98 of 1x1. (With four, only 2 or 3 got connections in most rounds; a later slide shows why.) That is a small loss, but it was there in every round. And four workers used 4.14 times the memory.

More OpenMP threads cost a lot. One worker with 2 threads gave 701 a second, with 4 threads 486, and with 8 threads 306. Four started workers with 8 threads each gave 192, only 0.18 of 1x1; 2 of the 4 got connections in four of the five rounds.

Under random arrivals, the threads made the service fall behind. At 646 requests a second, one worker with one thread had a p99 of 7.00 ms. One worker with 4 threads had a queue that never emptied, and a p99 of 2.7 seconds.

So on one core, the best choice was the simplest one: one process and one thread. The rest of the lesson shows where these numbers come from, and why.

How the Lab Was Built

I wrote the lab's design into the docstring of scripts/labs/serving/workers_and_threads.py on 2026-10-04, before any run of this lesson. It names every measurement, every setting, the rule for "told apart from noise", and seventeen guesses.

A page in five labelled zones, titled three parts, one core, 250 runs. Part P, predict alone: predict_proba on 1, 32, 1,024 and all 26,851 rows with 1, 2, 4 and 8 OpenMP threads; 5 runs, each a fresh process; CPU time against wall time. Part G, the GIL: a side thread counts in Python while the main thread runs predict_proba, sorted() or a Python loop; 5 runs. Part B, the service: 11 settings, 1, 2 or 4 workers x 1, 2 or 4 threads, plus 1x8 and 4x8; each the most answers per second (closed loop), and random arrivals at 0.3, 0.6, 0.85 x S (323, 646, 916 a second), 8 s each; 5 rounds, shuffled. Memory: RSS and Pss of every process, read from /proc at the end of each closed-loop run. The rule: a setting differs from 1x1 only if all 5 paired rounds agree in sign and the two sets of values do not overlap. Caption: 17 guesses written first; 14 came out right; the question: on one core, does anything beat one worker with one thread?

Part P times predict_proba alone, with no web server, with 1, 2, 4 and 8 OpenMP threads. It uses 1 row, 32 rows, 1,024 rows and all 26,851 test rows. Each thread count runs as a fresh Python process, five times, in a shuffled order. It also records how much CPU time the process used during the timed calls, so I can see whether the core was busy or waiting.

Part G checks the GIL. One thread counts in a Python loop. Meanwhile the main thread calls predict_proba on all rows, or sorts a list, or runs its own Python loop. If the counting thread keeps going while the main thread works, the main thread's work let the lock go.

Part B is the service. First I measured S, the most answers per second of 1x1 with a closed-loop client, as in lesson 4. Then each of the 11 settings got a closed-loop run and three open-loop runs, at 0.3, 0.6 and 0.85 times S. That is 44 runs a round, five rounds, in a new random order each round. Within a round, every setting at one rate got exactly the same arrival times and customers. So I could compare each setting with 1x1 on the very same traffic.

The Machine, and the Client That Shares It

Like lessons 2 to 4, every timing comes from a small machine I rented from Amazon Web Services. None comes from my laptop, which is always busy with other work.

Five rows, each with a logo. EC2 c7g.medium, us-east-1: 1 vCPU (Neoverse-V1), 2 GiB, not burstable; the account allows 1 vCPU at a time; on for 1.579 hours, $0.0573. Ubuntu 24.04, arm64: core at least 99.0% idle before every one of 250 runs; host steal at most 0.00%; timers stopped. Python 3.13.15, scikit-learn 1.9.1: numpy 2.5.3, threadpoolctl 3.7.0; OpenMP threads set with OMP_NUM_THREADS; switch interval 0.005 s. FastAPI 0.142.2, uvicorn 0.54.0: --workers 1, 2 or 4, uvloop 0.23.0, httptools 0.8.0; lesson 2's handler, unchanged. SQLite, lesson 2's feature table: 26,851 rows; the same model file and request list as lessons 2 to 4, same sha256. Caption: every timing I quote comes from this box, never from my laptop.

The box is an EC2 c7g.medium. AWS's own API lists it with 1 virtual CPU, 1 core, 2,048 MiB of memory and no bursting. Its AWS page says C7g instances use Graviton3 processors.

A few names in the figure. uvloop is a faster event loop that uvicorn can use in place of Python's own. httptools is a fast reader of messages. Lessons 2 to 4 used both. Steal is time the machine under a rented box gives to other customers' boxes instead of mine; Linux counts it, and here it was 0.00%.

The client and the service share that one core, as in lesson 4. So every answer costs some core time in the client too. In the closed-loop runs of 1x1, the client used about 0.08 of the core. Its work per answer is the same in every setting, so the comparisons between settings are fair. But the absolute numbers are lower than a service would reach with its client on another machine.

Before every run, the service was started and every worker had loaded the model. Then the core had to be at least 95% idle over two seconds before the client began. In all 250 runs it was at least 99.0% idle. The one-minute load average is recorded too, but it decays too slowly to wait for between runs, so it reads high right after a busy run.

The Whole Lab in One Picture

Before the recordings, here is the whole lab as one numbered sequence. The numbers on the arrows match the numbers on the recordings that follow.

A sequence diagram with five lifelines: my laptop, AWS, the box, client and W workers, and thirteen numbered arrows. 1, my laptop asks AWS whether any box is running. 2, key, group, SSH rule. 3, image, run-instances. 4, my laptop to the box: wait, then SSH. 5, Python, files. 6, stop timers, look. 7, start, log out. 8, my laptop to AWS: read the console. 9, client to W workers: requests. 10, W workers to client: answers, pid, time. 11, my laptop to the box: the demo. 12, the box to my laptop: raw files. 13, my laptop to AWS: terminate, delete. Below: arrows 9 and 10 repeat 1,077,019 times in the counted part of the service runs (warm-ups excluded), with nobody connected; the client, the W worker processes and their threads all share the box's one core; step 8 reads the box's serial console through AWS and never touches the box. Caption: steps 1 to 8 and 11 to 13 are the recordings that follow, numbered the same; all on the box that measured the numbers.

Steps 1 to 7 happen once: check, rent the box, set it up, and start the schedule. Step 8 is how I watched the schedule without touching the box. Arrows 9 and 10 are the timed work: 1,077,019 answers in the counted part of the runs, warm-ups excluded, with nobody logged in. Each answer carries two things lesson 2's answer did not: the id of the worker process that answered, and how long its own predict_proba call took. Steps 11 to 13 come after the schedule: the student demo, the copy of the results back to my laptop, and the deletion of everything.

Build the Lab Box Yourself

The rented box is part of the lab, so I show every step of it. You can rent the same kind of box and run the same schedule. The flow figure numbers the steps, and each recording below carries the same number.

Which box you see. Every recording comes from the box that measured this lesson's numbers. I planned the order so that no recording happened during a timed run, and nobody logged in while the schedule ran. Steps 1 to 7 come before the schedule. Step 8 reads the box from outside without logging in. Steps 11 to 13 come after the schedule finished.

What you need first. An AWS account, the AWS command line tool (aws) set up with your keys, and ssh, scp, rsync and curl. All the steps live in one file, scripts/labs/serving/wrk_files/wrk_aws_steps.sh. You run it from the scripts/labs/serving folder, one step at a time: bash wrk_files/wrk_aws_steps.sh check, then access, and so on. Every code block below is copied word for word from that file.

The blocks use a few names the file sets at its top; run these lines first, from the scripts/labs/serving folder, if you paste the blocks by hand. All of them are copied from the file, except HERE=$PWD: the file works out that folder from its own location, which a pasted line cannot do.

export AWS_DEFAULT_REGION=us-east-1 AWS_PAGER=""
NAME=ai-research-course-wrk
HERE=$PWD
STATE=${WRK_STATE:-$HOME/lab-data/serving/wrk}
KEY=$HOME/lab-data/serving/keys/$NAME.pem
K=${K:-5}; SAT_S=${SAT_S:-8}; OPEN_S=${OPEN_S:-8}
mkdir -p "$STATE" "$HOME/lab-data/serving/keys"
SSHO=(-o LogLevel=ERROR -o StrictHostKeyChecking=accept-new -o "UserKnownHostsFile=$STATE/known_hosts")
[ -f "$STATE/id" ] && ID=$(cat "$STATE/id")
[ -f "$STATE/ip" ] && IP=$(cat "$STATE/ip")
[ -f "$STATE/sg" ] && SG=$(cat "$STATE/sg")
  B=/home/ubuntu/wrk

Steps 1 to 3: Check, Lock the Door, Rent the Box

Step 1: is any other course box running? My account allows one box of this kind at a time, and another lesson may be using it. So the first command counts the course's boxes that are not yet deleted. It must print 0.

A real terminal recording of step 1. The shell prints the command aws ec2 describe-instances with filters for the tag Project=ai-research-course and the states pending, running, stopping, stopped and shutting-down, asking for the number of instances; the answer is 0.

The recording shows the command and its answer, 0, so no other course box was alive. This is the command, copied from the step script:

  aws ec2 describe-instances --filters "Name=tag:Project,Values=ai-research-course" \
    "Name=instance-state-name,Values=pending,running,stopping,stopped,shutting-down" \
    --query "length(Reservations[].Instances[])"

Step 2: a key and a locked door. A key pair is how you prove to the box that you are allowed in. AWS keeps one half, and you keep the other half in a file that only you can read. A security group is a door rule for the box. Mine opens only port 22, the SSH port, and only to my own address. Both get tags, labels that say which project and lesson they belong to.

A real terminal recording of step 2, with the address, network and group ids replaced by placeholders. create-key-pair for ai-research-course-wrk with the tags Project and Lesson writes the key to the key file without printing it; chmod 600 on the key file; curl reads my address; describe-vpcs finds the default network; create-security-group makes the group ai-research-course-wrk with its tags; authorize-security-group-ingress prints a small table: from my address /32, port 22. Last line: the key file, readable only by its owner, 388 bytes.

The key material went straight into the file and was never printed, and the door rule names only my own address. These are the commands:

  aws ec2 create-key-pair --key-name ai-research-course-wrk --key-type ed25519 \
    --tag-specifications "ResourceType=key-pair,Tags=[{Key=Project,Value=ai-research-course},{Key=Lesson,Value=wrk}]" \
    --query KeyMaterial --output text > ~/lab-data/serving/keys/ai-research-course-wrk.pem
  chmod 600 ~/lab-data/serving/keys/ai-research-course-wrk.pem
  MYIP=$(curl -sf https://checkip.amazonaws.com)
  VPC=$(aws ec2 describe-vpcs --filters Name=isDefault,Values=true --query "Vpcs[0].VpcId" --output text)
  SG=$(aws ec2 create-security-group --group-name ai-research-course-wrk --vpc-id "$VPC" \
    --description "lesson 5 workers and threads, ssh from one address" \
    --tag-specifications "ResourceType=security-group,Tags=[{Key=Project,Value=ai-research-course},{Key=Lesson,Value=wrk}]" \
    --query GroupId --output text)
  aws ec2 authorize-security-group-ingress --group-id "$SG" --protocol tcp --port 22 --cidr "$MYIP/32" \
    --query "SecurityGroupRules[].{port:FromPort,from:CidrIpv4}" --output table

Steps 4 to 7: Wait, Set Up, Look, Start

Step 4: wait until it runs. A new box takes a short while to start. The AWS tool can wait for you. Then the script reads the box's public address into the variable IP, and tries SSH every five seconds until the box answers.

A real terminal recording of step 4, with the instance id and address replaced by placeholders. aws ec2 wait instance-running; describe-instances reads the public address into IP; describe-instances prints running, c7g.medium, us-east-1b. Then ssh with ConnectTimeout=5 runs true and fails twice with a sleep 5 between, and a third ssh answers: ssh works, aarch64, Ubuntu 24.04.5 LTS.

The first two SSH tries failed because the box was still starting, and the third one answered. These are the commands, with the retry loop:

  aws ec2 wait instance-running --instance-ids "$ID"
  IP=$(aws ec2 describe-instances --instance-ids "$ID" \
    --query "Reservations[0].Instances[0].PublicIpAddress" --output text)
  until ssh -i ~/lab-data/serving/keys/ai-research-course-wrk.pem "${SSHO[@]}" -o ConnectTimeout=5 \
    ubuntu@"$IP" true 2>/dev/null; do sleep 5; done

Step 5: Python and the lab files. The box gets the same Python, 3.13.15, and the same pinned library versions as lessons 2 to 4, through uv, a fast installer for Python. Then scp, a copy command that works through SSH, sends the files over. They are lesson 2's model file, feature table, request list and service, this lesson's files, and the files the demo needs. The last command builds the SQLite table on the box and prints the files' sha256 fingerprints, so you can check they match the ones on your computer.

A real terminal recording of step 5, with the address replaced by a placeholder and my home folder shortened to a tilde. Over ssh: make the folders; install uv, Python 3.13.15 and a venv; install the pinned packages from lat_requirements.txt. Then scp copies lesson 2's model, data and service files, this lesson's wrk_files, the features chapter's files and data, and the demo. Last: build_store.py prints 26851 rows; the sha256 of model.pkl starts bec4d11e, of store.parquet 62aadc30, of requests.parquet 407d62e6; then the folder's list, from box_info.sh to wrk_service.py.

Steps 8 and 11 to 13: Watch, Demo, Bring It Home, Delete

Step 8: watch without touching. While it runs, the schedule writes one progress line to the box's serial console after every run. The serial console is a log that AWS keeps for the box, and you can read it through AWS without logging in. So I could follow the runs and never disturb them.

A real terminal recording of step 8, with the instance id replaced by a placeholder. aws ec2 get-console-output with --latest, piped through grep WRK-PROGRESS and tail -12, prints progress lines with times: gil-r1 to gil-r5 done between 09:21:56 and 09:22:52; sat0-r1 to sat0-r5 done; S=1077.125 and the rates 323, 646 and 916 at 09:24:01; open-r1-w2t2-m2 done at 09:24:24.

These lines came from AWS's copy of the console while the runs went on, so the box was never touched. This is the command to read them:

  aws ec2 get-console-output --instance-id "$ID" --latest --output text | grep WRK-PROGRESS | tail -12

Step 11: the student demo, on the same box. After the schedule, I ran the demo you will meet later, so you can see what it prints on one core. Its output went into a file at the same moment it went to the screen.

A real terminal recording of step 11, with the address replaced by a placeholder: python wrk_demo.py on the box, its output also written to wrk-demo-run.txt. It prints the machine, 1 cores, load 0.69, and test AP 0.5450. Part 1: predict_proba on 1 row 509.2 us with 1 thread, 1016.6 with 2, 1653.4 with 4; on all rows 47985.7, 49035.5 and 49726.8 us. Part 2: the side thread's share, 0.79 during predict_proba and 0.48 during sorted(). Part 3: rows per second with 1, 2 and 4 Python threads 1985, 1873, 1843; with 1, 2 and 4 worker processes 1977, 1735, 1411. Below the recording: its sorted() line, 0.48, came from a demo bug fixed after the box was gone (see Try It Yourself); the lab's own check measured 0.173.

This run happened after the last timed run of the schedule had finished, so it could not disturb any timing. Its sorted() line, 0.48, came from a demo bug fixed after the box was gone (see Try It Yourself); the lab's own check measured 0.173. This is the command:

  ssh -i ~/lab-data/serving/keys/ai-research-course-wrk.pem "${SSHO[@]}" ubuntu@"$IP" \
    'cd wrk/serving/examples && ../../venv/bin/python wrk_demo.py --save ~/wrk/raw/wrk-demo.json \
     | tee ~/wrk/raw/wrk-demo-run.txt'

Step 12: bring the results home first. Lesson 3 lost a whole box of results because it was deleted before anyone copied them. So the rule is: the moment the runs finish, copy the raw files back, before anything else.

The Lab's Report, Running

This is a real recording of the report script, wrk_report.py. It ran on my laptop, but it times nothing. It reads the raw files the box measured and does all the arithmetic again with its own code.

A terminal recording of wrk_report.py in eight numbered sections. 1: the box, c7g.medium, 1 vCPU, 1.58 h at $0.0363 an hour = $0.0573, everything deleted, stored files scanned with no ids, addresses or keys; 250 runs, the core at least 99.0% idle before each. 2: predict_proba per call by thread count for 1, 32, 1,024 and 26,851 rows, from 510.5 us with 1 thread on 1 row to 2842.4 us with 8 threads, every CPU-to-wall ratio 1.00, every score equal. 3: the side thread's share: 0.797 during predict_proba, 0.173 during sorted(), 0.498 during a Python loop. 4: a table of answers per second for 11 settings, from 1071.4 for 1x1 to 191.8 for 4x8, every one told apart from 1x1, with memory per setting. 5: a table of p50 and p99 in ms at three rates, with stars where a setting fell behind, and the scores, 826,617 and 282,392 answers all equal. 6: 14 guesses right, 3 wrong. 7 and 8: the demo and the playground match. Last line: all 1242 checks agree with the stored lab.

The report reads the small copy of the raw files kept in the repo. It computes its own percentiles, its own answers per second, its own paired differences and its own verdicts. Then it compares every one with the lab's results file. If any number disagreed, it would stop with an error. I checked that too: I changed one number in a copy of the results file, and the report printed "2 MISMATCHES" and stopped.

Section 1 is the quiet check. Before every one of the 250 runs, the core was at least 99.0% idle over two seconds. The machine under my box took no measurable time from it. No SSH session was open at the start or end of any run.

The system log had 80 lines during runs. 54 came from the AWS management agent failing to reach AWS, because my box had no permission to talk to it. 24 came from cron, which still ran because I stopped only the systemd timers. They were a statistics job (debian-sa1) and the hourly job, in 8 timed runs (open-r3-w4t8-m2, open-r4-w4t4-m1, open-r4-w4t8-m1, open-r5-w4t4-m1, open-r5-w4t8-m0, predict-r1-t4, sat-r1-w4t8 and sat-r3-w1t4). One line came from the kernel and one from systemd. No worker process died and restarted in any run.

Part P: More Threads Made a One-Row Call Slower

First, the model alone, with no web server. How long does one call to predict_proba take with 1, 2, 4 and 8 OpenMP threads, on one core?

A line chart of time per call as a multiple of the time with 1 thread, against the number of OpenMP threads, 1, 2, 4 and 8, on one core. Four lines: 1 row and 32 rows climb steeply, to about 5.6 and 5.2 times at 8 threads; 1,024 rows climbs gently to about 1.9 times; 26,851 rows stays almost flat near 1. A dashed line at 1 is labelled same as 1 thread. Below: 1 row, 1 thread 510.5 us, 2 threads 1001.1 us, 4 threads 1627.6 us, 8 threads 2842.4 us; all 26,851 rows, 1: 52.6 ms, 2: 53.2 ms, 4: 54.1 ms, 8: 56.0 ms. Caption: on one core, helper threads cannot work at the same moment; they only add the cost of waking them and taking turns.

One row took 510.5 microseconds with 1 thread, 1,001.1 with 2, 1,627.6 with 4 and 2,842.4 with 8. So 4 threads made one row 3.19 times slower, and 8 threads 5.57 times slower.

All 26,851 rows took about the same with any thread count: 52.6 ms with 1 thread and 54.1 ms with 4, only 1.03 times as long.

Why does a thread count hurt so much for one row and so little for many? Think back to the office.

To type one short line together, someone has to call the helpers over, give each a share, and wait for the last one to finish. scikit-learn's tree code does the same: at each call it wakes its helper threads, gives each a share of the rows, and waits for all of them. With one row there is almost nothing to share, so the call is mostly the waking and waiting. On one core, the helpers cannot even work at the same moment. Each one has to wait for its turn on the core, so every helper adds a context switch.

With 26,851 rows the waking and waiting costs about the same as before, but now it is a tiny part of a 52-millisecond call. The rows still get scored one after another on the one core, so there is no speed-up either.

Part P also measured the process's CPU time against the wall time of each call. At every thread count and every batch size the ratio was 1.00. So the core was never idle during a call: all the extra time went to the process's own threads waking, waiting their turn and switching. One possible part of it, which I did not test, is spin-waiting. An OpenMP helper that has finished may keep the core busy checking for new work for a short while before it sleeps. I could not see from this how it split between the threads. The service runs below can.

Part G: predict_proba Lets Python's Lock Go, and Waits to Get It Back

Before the service, one more question. Python has the GIL: only one thread can run Python code at a time. Does predict_proba hold the GIL while it works, or let it go? And what happens to it when another thread is busy?

The Python documentation defines the GIL as "The mechanism used by the CPython interpreter to assure that only one thread executes Python bytecode at a time". It also says that "some extension modules, either standard or third-party, are designed so as to release the GIL when doing computationally intensive tasks". scikit-learn's tree code is written in Cython, a language close to C. In version 1.9.1 its loop over the rows of one tree reads prange(..., nogil=True, ...), which means "run without the GIL". The loop over the trees, though, is Python: the source reads for predictors_of_ith_iteration in self._predictors:, one call per tree.

To see what happens, I ran a side thread that counts in a Python loop. Then I measured how fast it counted while the main thread did three kinds of work. I report it as a share of how fast it counted while the main thread slept.

Three panels, each with a number. predict_proba: 0.80; all 26,851 rows; one call 270.2 ms; 5 runs 0.80 to 0.80. sorted(): 0.17; 200,000 floats; one call 37.0 ms; 5 runs 0.17 to 0.18. Python loop: 0.50; 200,000 steps; one call 14.1 ms; 5 runs 0.50 to 0.50. Caption: predict_proba alone took 52.6 ms; with the side thread, 270.2 ms; 1 - 52.6 / 270.2 = 0.805, the share the side thread got.

During predict_proba on all rows, the side thread kept 0.797 of its speed, and one call took 270.2 ms instead of 52.6 ms. So the model call did let the lock go: otherwise the side thread could not have run most of the time.

During sorted(), which holds the lock, it kept only 0.173. Sorting a list of Python numbers is C code that keeps the lock while it sorts, because it compares Python objects. The side thread got to run only between one sort and the next.

Both threads run Python code, and Python makes the running thread hand the lock over every , which is 0.005 seconds by default. So they shared the core about half and half.

Part B: The Most Each Setting Could Answer

Now the service. Before the settings, I measured S, the most answers per second of 1x1. Five closed-loop runs gave 1,072.1 to 1,078.3 a second, and their median, S, was 1,077.1. The open-loop rates were 0.3, 0.6 and 0.85 times S: 323, 646 and 916 requests a second.

Then every setting got one closed-loop run in each of the five rounds. Here they are on one grid.

A grid with rows 1 worker, 2 workers and 4 workers, and columns T=1, T=2, T=4 and T=8. Each cell holds the answers per second as a share of 1x1 and the median itself. 1 worker: 1.00 (1,071 a second, framed), 0.65 (701), 0.45 (486), 0.29 (306). 2 workers: 0.98 (1,049), 0.43 (460), 0.27 (289), not run. 4 workers: 0.98 (1,049), 0.36 (383), 0.24 (253), 0.18 (192). Below: a darker cell gave fewer; moving right adds OpenMP threads per worker; moving down adds worker processes; the 4-worker row started 4 workers, but only 2 or 3 got connections in most rounds. Caption: lowest, 4 workers x 8 threads, 0.18 of 1x1.

Moving down the first column, adding workers, cost about 2%. 2x1 gave 0.979 of 1x1 and 4x1 gave 0.981. In every round both were below 1x1, so the loss is told apart, but it is small. Each extra worker is one more process that the core switches to and from. Its own copy of the model and libraries also competes for the processor's small, fast memory caches. That is one possible reason; I did not measure the switches or the caches.

Moving right along any row, adding threads, cost a lot. One worker went from 1,071 to 701, 486 and 306 answers a second as T went from 1 to 8.

Both together were the worst. 2x2 gave 460, less than 1x2's 701, and 4x4 gave 253. 4x4 started four workers of four threads each, but in 4 of the 5 rounds only three got connections, so twelve threads were busy, not sixteen.

The other eight settings varied by at most 16.1 answers a second across the five rounds. Three did not. 4x2 gave 380 to 383 in three rounds and 464 in two: the two in which only two of its four workers got connections. 4x8 gave 191 to 196 in four rounds and 164 in the one where three workers got connections. 4x4 gave 252 to 256 with three workers, and 240 in the one round where all four did.

Where the One Core's Time Went

Why did threads hurt so much in the service? During every closed-loop run, the client read from /proc, Linux's view of every running thread, how much CPU time each of the service's threads used. I split it into the main threads, which read requests and run the model, and all the other threads, which here are OpenMP's helpers.

Four hand-drawn bars against settings 1x1, 1x2, 1x4 and 1x8: 0.00, 0.13, 0.32 and 0.47. Below: each bar is the CPU time of the worker's threads other than its main thread, divided by the length of the counted part; 1.00 would be the whole core; the main thread, the one that reads requests and runs the model, got 1x1 0.92, 1x2 0.82, 1x4 0.65, 1x8 0.50. Caption: every share of the core a helper used was a share the request handler did not get.

With 8 threads, the helpers used 0.47 of the core, and the main thread got only 0.50. With 1 thread, the main thread got 0.92 of the core and the client most of the rest.

The helpers were not doing useful work in parallel, because there was nothing in parallel to do. A request scores one row, and one row cannot be split usefully. So the helpers' time is the cost of the waking, waiting and switching from Part P, paid on every request. Spin-waiting, helpers busy-checking for work before they sleep, may be part of their CPU time; I did not test that.

scikit-learn's documentation on parallelism says it plainly: "It is generally recommended to avoid using significantly more processes or threads than the number of CPUs on a machine." And of too many threads, it "leads to oversubscription of threads for physical CPU resources and thus to scheduling overhead".

Started with four workers, of which 2 to 3 got connections. Five hand-drawn bars, each one second of the core split three ways, for 1x1, 4x1, 4x2, 4x4 and 4x8. The 1x1 and 4x1 bars are almost all main threads with a small client part. From 4x2 to 4x8 the helper part grows and the main part shrinks. Below: the workers' main threads, the workers' other threads (OpenMP's helpers), and the client; 1x1: 0.92 / 0.00 / 0.08; 4x1: 0.93 / 0.00 / 0.07; 4x2: 0.81 / 0.15 / 0.03; 4x4: 0.59 / 0.39 / 0.02; 4x8: 0.47 / 0.51 / 0.02; the rest of the second went to the system and to switching; each 4x setting started 4 workers, and 2 to 3 of them got connections. Caption: time spent by helper threads was time the event loop did not get.

With four started workers, the story is the same. (In these runs 2 to 3 of the four got connections.) Four workers with one thread each kept 0.93 of the core in their main threads, as much as one worker. Four workers with 8 threads each gave 0.51 of the core to helpers.

Under Random Arrivals: Who Kept Up

The closed loop shows the most each setting could do. Real users arrive at random, so here are the open-loop runs.

A line chart of p99 latency on a log scale from 1 ms to 100 s, against three arrival rates, 0.3x, 0.6x and 0.85x of S, for six settings: 1x1, 2x1, 4x1, 1x2, 1x4 and 4x4. The 1x1, 2x1 and 4x1 lines stay low, between about 3 ms and 52 ms. The 1x2 line rises from about 7 ms to about 3 s; the 1x4 and 4x4 lines rise to seconds already at 0.6x and reach about 7 s and 22 s at 0.85x. Below: fell behind in at least one round: 1x2 at 0.85x; 1x4 at 0.6 and 0.85x; 4x4 at 0.3, 0.6 and 0.85x; a queue that never empties grows with the run's length, so those p99s are not steady numbers; p99 at 0.6x S: 1x1 7.00 ms, 2x1 8.23 ms, 4x1 9.31 ms, 1x2 75.3 ms, 1x4 2719 ms, 4x4 13396 ms. Caption: same arrivals for every setting in a round: only workers and threads differ.

1x1 had the lowest p99 at every rate: 3.39 ms at 0.3 times S, 7.00 ms at 0.6 and 29.60 ms at 0.85. I compared it directly with each other setting, round by round. At 0.6 and 0.85 times S, its p99 was lower than all ten others, told apart. At 0.3 times S it was lower than nine; against 2x1 (3.66 ms) the difference cannot be told apart from noise.

More workers cost a little at the tail. At 0.85 times S, 2x1 had a p99 of 42.78 ms and 4x1 51.96 ms, against 29.60 ms for 1x1. Their medians were almost the same as 1x1's.

Extra threads made the service fall behind. 1x4's most was 486 answers a second, so at 646 arrivals a second its queue grew for the whole run, and its p99 reached 2.7 seconds. Every setting with 4 or 8 threads fell behind at 0.6 times S or earlier. Their p99s above that point only mean "the queue never emptied"; a longer run would have made them larger.

One of my guesses was wrong here. I expected every setting to keep up at 0.3 times S, 323 requests a second. Four did not in every round: 2x4, 4x4, 1x8 and 4x8. Their closed-loop most was already below 323, or close to it.

4x4 kept up in 3 of 5 rounds, and the rounds split by how many workers answered. It kept up in the three rounds where 3 workers answered, and fell behind in the two where all 4 did. More busy workers meant more helper threads wanting the core. Its closed-loop most, 253 a second, was measured with a different client and 16 connections, so the two numbers do not line up exactly.

Which Worker Answered

That last point needs a closer look. Each answer carried the id of the worker process that sent it, so I could count how the work was shared.

Three groups of five strips, one strip per round, each split by worker, with how many workers answered at the right. 2 workers x 1 thread: each round split about half and half, 2 of 2. 4 workers x 1 thread: round 1 2 of 4, rounds 2 to 5 3 of 4. 4 workers x 4 threads: 3 of 4 in four rounds and 4 of 4 in one. Below: each strip is one round, split by worker; at its right, how many workers answered at all; 16 connections were opened at the start and each stayed with the worker that accepted it, so a worker that accepted more connections answered more. Caption: on one core it hardly matters who answers: they all wait for the same core.

With four workers and 16 connections, one or two workers often answered nothing at all. uvicorn's workers share one listening socket. A connection goes to whichever worker accepts it first, and then it stays there for the whole run. For 4x1, the 16 connections opened at the start went to two or three of the four workers in every round.

In the open-loop runs, with 512 connections, 4x1 still had only two to four workers answering, depending on the run. So a setting called "4 workers" often served with three.

This table was added after the results, from the answers' process ids. It shows how many of the four started workers answered in each closed-loop round, and that round's answers per second.

settinground 1round 2round 3round 4round 5
4x12: 1,0513: 1,0443: 1,049

Memory: Each Worker Carries Its Own Copy

Workers are not free in memory. At the end of every closed-loop run, the client read each service process's RSS and Pss from /proc.

An isometric drawing of three rows of blocks, one block per worker process, each as tall as its RSS, the median of 5 rounds. Row 1 worker x 1 thread: one block; total RSS 226 MB, total Pss 194 MB. Row 2 workers x 1 thread: two blocks; total RSS 492 MB, total Pss 359 MB. Row 4 workers x 1 thread: four blocks of about the same height; total RSS 934 MB, total Pss 619 MB. Below: one block per worker process, as tall as its RSS (226 MB for one worker, median); uvicorn starts each worker as a fresh Python (spawn), so each loads its own model and libraries; Pss is lower than RSS because pages of the same library files are shared with every other process that loads them, the client included. Caption: the 4-worker service used 4.14 times the RSS of the 1-worker one, its supervisor process included.

Each worker used 218 to 228 MB of RSS, whatever its thread count. One worker was 226 MB in all. Four workers were 934 MB, 4.14 times as much. That is more than four copies. With more than one worker, uvicorn also runs a supervisor process (about 27 MB), which starts the workers, checks they are alive and restarts any that die. Python adds a small helper process (about 17 MB).

Why does every worker hold a full copy? uvicorn's documentation says: "Uvicorn includes a --workers option that allows you to run multiple worker processes." And: "Unlike gunicorn, uvicorn does not use pre-fork, but uses" spawn. Spawn starts each worker as a brand-new Python, which imports numpy, scikit-learn and FastAPI and loads the model from the file, all over again. Pre-fork, which gunicorn uses, loads the model once and then copies the whole process with fork. The copies share the parent's memory pages until one of them writes to a page. This is called copy-on-write.

Pss shows what is really shared. Four workers' Pss was 619 MB, against 934 MB of RSS. The difference is pages of the same library files on disk, which the operating system loads once and lets every process map. Each worker's own model, arrays and Python objects are not shared. Lesson 10 of this chapter measures memory when many models live on one box.

What a Box With More Cores Would Change

Everything above is one small tree model on one core. I could not rent a box with more cores, so I give no numbers for one. But the measurements here show which parts of the answer depend on the core count, and why.

A two-column ledger titled this box, measured, and more cores, reasoned. Left: here, 1 core, 4 workers gave 0.98 of 1x1's answers per second; they took turns. Here: 1 row with 4 threads took 3.19 times as long as with 1 thread. Here: beside a busy Python thread, one predict_proba call took 270.2 ms, not 52.6: each tree waited to get the GIL back. Here: each worker held 226 MB (RSS); 4 workers 934 MB in all. Right: with N cores, up to N workers can each have a core; answers per second can grow, until something shared runs out. With N cores, threads can split a big batch; a one-row call has nothing to split, and waking them still costs. With N cores, the tree work can overlap on other cores, but every retake of the lock still waits behind a busy Python thread, so a busy thread beside the model slows every call, on any number of cores. Memory still grows with each worker, cores or not; measure it before adding workers.

Workers are what more cores would change most. On one core, four workers took turns and gave 0.98 of one worker's answers. With four cores, each worker could have its own core. Much of each answer here is Python code: reading the request, checking it, and the Python part of the model call. Lesson 2 timed each step. Only one thread at a time can run Python code in a process. So more processes, not more threads, is the usual way to use more cores for a service like this one. How much it would gain, I cannot say without measuring.

One-row threads would still not help much. A one-row call has nothing to share between helpers. More cores remove the waiting for a turn, but not the waking and the gathering.

A busy Python thread beside the model still costs on any machine. With more cores, the tree work can overlap on other cores. But every retake of the lock still waits behind a busy Python thread, so a busy thread beside the model slows every call, on any number of cores. The box used the standard Python build. Python 3.13 also offers an optional free-threaded build without the GIL, which I did not measure. Its documentation says "Starting with the 3.13 release, CPython has experimental support for a build of Python called free threading where the global interpreter lock (GIL) is disabled."

The danger of multiplying is the same on any machine. scikit-learn uses as many threads as there are cores by default. Its own documentation gives the example of eight jobs each running a model with eight threads on eight cores, "a total of 8 * 8 = 64 threads". Four uvicorn workers on a four-core box, each with four OpenMP threads by default, is the same mistake: sixteen threads for four cores. That is exactly the 4x4 setting here, on a smaller scale.

My Seventeen Guesses Before the Run, Checked

I wrote seventeen guesses into the lab before it ran. Each is quoted here word for word from the docstring. Fourteen were right and three were wrong.

  1. "P1. predict on 1 row at T = 1: median per call between 450 and 650 us (lesson 4: 501.7 us)." Right: 510.5 us.

  2. "P2. predict on 1 row at T = 4: median at least 1.5 x T = 1's." Right: 3.19 times.

  3. "P3. predict on all 26,851 rows at T = 4: median within 15% of T = 1's (one core: no speed-up, small loss)." Right: 1.030 times.

  4. "P4. every score at every T equals T = 1's single-row score to the bit." Right.

  5. "G1. side thread during predict_proba(26,851 rows): between 0.35 and 0.65 of its alone rate." Wrong: 0.797. I expected the two threads to share the core about half and half. Instead the model call spent most of its time waiting to take the lock back, one switch interval per tree, as the GIL slide explains. The side thread ran meanwhile.

  6. "G2. side thread during sorted(): below 0.15 of its alone rate." Wrong, though close: 0.173. Each sort took only 37 ms, so the side thread's short turns between sorts added up to a bit more than I guessed.

  7. "G3. side thread during a pure-Python loop: between 0.35 and 0.65 of its alone rate." Right: 0.498.

  8. "S1. S (1x1, closed loop) between 1,000 and 1,200 answers per second." Right: 1,077.1.

  9. "S2. 2x1 and 4x1 closed-loop answers per second below 1x1's, told apart, each by less than 15%." Right: 0.979 and 0.981.

Try It Yourself

The full lab needs a rented box and about an hour and a half. The demo, wrk_demo.py, runs on your own computer. It times OpenMP threads, checks the GIL, and shares one-row calls between Python threads and between worker processes.

A page in four labelled zones, headed wrk_demo.py, designed before it ran. Train: the features chapter's model; it must score test AP 0.5450. OpenMP threads: predict_proba on 1 row and on all rows with 1, 2 and 4 threads. The GIL: a side thread counts while the main thread runs predict_proba or sorted(). Threads against processes: 3,000 one-row calls shared over 1, 2 and 4 Python threads, then 1, 2 and 4 worker processes: rows per second. Caption: on the box (1 core) 4 processes gave 1,411 rows a second and 1 gave 1,977.

I wrote the demo's design into its docstring after the lab's design and before the demo first ran.

A real screenshot of VS Code with wrk_demo.py open at the top of the file, showing its docstring: what it needs, how to run it, and the design written before it first ran.

Before you run this lab. The demo uses lesson 2's Python environment, with the same library versions as the lab's box. If you made it for lessons 2 to 4, use it again. If not, make it once, inside the scripts/labs/serving/examples folder:

python3 -m venv venv-sv
source venv-sv/bin/activate        # on Windows: venv-sv\Scripts\activate
pip install -r lat_requirements.txt

I used Python 3.13. Then run python ../../features/fetch_data.py once, which downloads the shop data. Now run python wrk_demo.py. It needs no GPU, no web server and no cloud account. If a package is missing, it prints one line saying why and stops.

The first line it prints says the times come from your machine, with its number of cores and its load at that moment. Your machine probably has more cores than my box, so its numbers answer a different question: they are not a check on the lab. Do more threads slow down your one-row calls? Does the side thread keep going during ? Do more worker processes give more rows a second on your cores, and more Python threads not?

Pick Workers, Threads and a Rate

This box holds the real results from the lab. It needs nothing but Python, so it runs in your browser. It has no model and no clock: it only looks up what the box measured.

Press Run. Then try W = 1 with T = 4 at RATE = 0.6, where one worker fell behind. Try W = 4, T = 8. Every run shows 1 worker x 1 thread first, at the same rate, so you always compare with it.

The report script writes this box from the lab's raw files, runs it with six settings, and checks that each run prints the lab's numbers.

The Lab's Code, Piece by Piece

The lab has one file that runs on my laptop, a few small files that run on the box, and the step script you saw above.

workers_and_threads.py, on my laptop, holds the design in its docstring, written before any run. Its commands are check (lesson 2's files match by sha256), status (reads the box's serial console), collect, cost and factcheck. collect calls wrk_stats.py, which copies the box's small raw files into results/wrk-raw/, passes every log line through wrk_redact.py, and does all the arithmetic.

wrk_files/wrk_service.py is the service. It imports lesson 2's service file unchanged, so the model, the SQLite table and the request checks are lesson 2's, and the handler is lesson 2's too. It adds two fields to each answer: the process id, and the nanoseconds of its own predict_proba call. When a worker has loaded the model, it writes a small "ready" file, so a run starts only when every worker is ready.

wrk_files/wrk_client.py is lesson 4's client, copied with credit. It keeps the same closed loop, the same open loop with random arrivals and 512 kept-alive connections, and counted from each request's arrival. Changed for several workers: it warms every connection. Around the counted part, it reads for the CPU time of every service thread, and afterwards for each process's RSS and Pss. It also checks every score against the box's own single-row score.

How to Choose Workers and Threads

Here is the order I would follow, using only what this lab did.

A flowchart. Count the cores the service really gets leads to a diamond: one core? Yes leads to 1 worker, OMP_NUM_THREADS=1. No leads to workers x threads at most the cores; start with threads = 1, then to measure each setting; compare it directly with the simplest. Both paths lead to check memory: one model copy per worker. Below: here, on one core, 1x1 gave 1,071 answers a second; 10 of 10 other settings gave fewer, told apart, and none gave more; 4 workers used 934 MB of memory in all, against 226 MB for one. Caption: never more threads in total than cores, and never on trust: measure on the box you serve from.

The chart starts with the one fact everything else depends on: how many cores the service really gets. Here are the same steps in words, each with the number from this lab behind it.

  1. Count the real cores. Run nproc on the machine, and check any CPU limit your container sets. My box printed 1.

  2. Set OMP_NUM_THREADS on purpose. scikit-learn takes it as given, even above the core count. Left unset, it uses the number of cores, in every worker. Here, 4 threads made a one-row call 3.19 times slower.

  3. Keep workers times threads at most the cores. Here, every setting that broke this rule gave fewer answers than 1x1, and those with 4 or more threads fell behind at random arrivals that 1x1 handled easily.

  4. Measure the most a setting can answer, and its under random arrivals, and compare directly. Here, 2x1 and 4x1 looked almost the same as 1x1 at a glance, but compared round by round they gave about 2% fewer answers.

  5. Count the memory of every worker. Here each worker used about 220 MB, so four workers used 934 MB of a 2 GiB box.

When More Workers or Threads Help, and When They Do Not

More worker processes help when there are free cores for them. Each worker can run Python on its own core, and the GIL is per process, so it does not stop them. On one core they only took turns: 0.98 of one worker's answers.

More OpenMP threads help when one call has a lot of rows to share and there are free cores. A batch job that scores millions of rows on a big machine is the right place. A service that scores one row per request is the wrong place: here, more threads only slowed every call down.

Never let the two multiply past the cores. Workers times threads is the number of threads that want the core at once. On this box, the worst setting, four workers with eight threads, gave 0.18 of the answers of the simplest one.

Do not put a busy Python thread beside the model. predict_proba lets the GIL go for each tree's work. But beside a busy Python thread, each of its 54 trees waited about one switch interval to take it back. One call took 270.2 ms instead of 52.6. On more cores the tree work can overlap, but those waits for the lock remain.

Do not add workers without counting memory. Each one here was a full copy, about 220 MB.

Do not copy my handler as it is into a busy service. Like lessons 2 to 4, it is an async def that calls SQLite and predict_proba, which block the event loop while they run. With one thread per worker, that is what I measured on purpose.

What This Lab Cannot Tell You

Two columns titled shows and cannot show. Shows: one small tree model on one 1-vCPU box; the client shares the core. 11 settings, 3 steady rates, 8 seconds each, 5 rounds. uvicorn's own workers (spawn), one listening socket. Cannot show: a box with more cores, because the account allows one vCPU at a time. Long runs, rising traffic, or a big model whose calls take far longer. gunicorn's pre-fork workers, which can share memory pages (lesson 10).

One core, and the client on it. The account's limit made this a one-core lesson. Every answer about more cores in this lesson is reasoning, not measurement. The client also used part of the core, about 0.08 of it in 1x1's closed-loop runs.

One small model. A one-row call here is about half a millisecond, most of it fixed cost. A model whose calls take far longer would change how much threads cost compared with the work.

Steady rates for 8 seconds. Real traffic rises and falls. Above a setting's limit, its latencies grow with the run's length, so they are not stable numbers.

Uneven workers. With four workers, some accepted no connections, so "4 workers" often served with three. I report it as measured; I did not try to balance them.

Labelled additions. Two figures use readings the design named, but I chose to draw them after seeing the results. They are where the core's time went per setting (CPU per thread), and the count of workers that answered (the answer's process id). The demo's sort list was fixed after the box run, as the demo slide says. The step 5 recording was the second run of that step on the same box.

What to Do on Monday

A hand-drawn grid of six cards, titled five habits for workers and threads. 1, count the real cores: nproc on the box, and any container CPU limit. 2, workers x threads: at most the cores; set OMP_NUM_THREADS on purpose. 3, measure, do not trust: closed loop for the most, open loop for latency. 4, compare directly: each setting against the simplest one, round by round. 5, count the memory: RSS and Pss of every worker, not just one. The reason: here, on one core, 1x1 gave the most answers; 1 row with 4 threads took 3.19x as long. Caption: more processes and threads help only when there are cores for them to run on.

If you take one thing to work on Monday, look at how your model service is started. Find the number of workers, find OMP_NUM_THREADS, and run nproc where it runs. If workers times threads is more than the cores, you are paying for switching, not for work.

Then set OMP_NUM_THREADS=1 for a service that scores a few rows per request, and set the workers to at most the cores. Measure both the most it answers and its under random arrivals, against the setting you had before.

The one idea to keep: processes and threads are ways to use more cores. Where there are no more cores, they only take turns, and each turn has a price.

Knowledge Check

Knowledge Check

4 questions - Score 80% to pass

Q1

On the one-core box, how did four started workers with one thread each (2 or 3 got connections) compare with one worker and one thread?

Q2

Why did one-row predict_proba calls get slower with more OpenMP threads on one core?

Q3

What did the GIL check show about predict_proba on all 26,851 rows?

Q4

A service scores one row per request on a box with 4 cores. Which start-up follows this lesson's advice?

context switch
Oversubscription

RSS is how much memory a process is using right now. It counts shared pages in full for every process that uses them. Pss counts each shared page only once in total, split between the processes that share it. A closed loop client sends a new request only when an answer comes back, so it measures the most a service can do. An open loop client sends requests at random moments on its own clock, as real users do, so it shows waiting. Mine draws the moments as Poisson arrivals: random, with a steady average rate, so sometimes several come close together.

The rule, written before the runs, is lesson 3's rule. A setting is "told apart from run-to-run noise" only if two things hold. All five paired differences have the same sign. AND its five values and the other setting's five values do not overlap. When I say one setting beat another, I compared those two directly.

What it costs. A c7g.medium costs $0.0363 an hour on demand in us-east-1, from AWS's price list, and AWS bills it by the second. My box was on for 1.579 hours, from launch to terminated, which is $0.0573 for compute. Its 16 GiB disk was deleted with it. These numbers are in results/wrk-cost.json. The schedule itself took about 1 hour and 27 minutes.

What the recordings hide. Every line passed through a small filter, wrk_redact.py. It builds on lesson 3's tail_redact.py. It hides my account number and every IP address. It also hides every name AWS makes up for a resource, such as the box's id, the security group's id and the image id. And it hides host names, my user name and anything from a key file. Where you see <ip> or i-<id>, your terminal shows the real value. Lines that start with + are the shell printing each command just before it runs it. I then scanned every stored file and recording again for those patterns, and every count was 0.

Keep the key file outside any git folder. Mine lives in ~/lab-data/serving/keys/, and the last step deletes it.

Step 3: rent the box. First the script asks AWS's parameter store (SSM) for the id of the newest Ubuntu 24.04 image for arm processors in this region. So you never copy an id that has gone out of date. Then run-instances asks for one c7g.medium with a 16 GiB disk that is deleted with the box, with the tags on both. The box's id goes into the variable ID.

A real terminal recording of step 3, with the image, group and instance ids replaced by placeholders. aws ssm get-parameter reads /aws/service/canonical/ubuntu/server/24.04/stable/current/arm64/hvm/ebs-gp3/ami-id into AMI; aws ec2 run-instances with that image, type c7g.medium, the key ai-research-course-wrk, a 16 GiB gp3 disk deleted on termination and the tags Project, Lesson and Name, asking only for the instance id, which goes into ID; describe-instances prints the id, c7g.medium, pending and the image id.

The box started in the "pending" state, which is normal for the first few seconds. These are the two commands, the image lookup and the launch:

  AMI=$(aws ssm get-parameter --query Parameter.Value --output text \
    --name /aws/service/canonical/ubuntu/server/24.04/stable/current/arm64/hvm/ebs-gp3/ami-id)
  ID=$(aws ec2 run-instances --image-id "$AMI" --instance-type c7g.medium --count 1 \
    --key-name ai-research-course-wrk --security-group-ids "$SG" \
    --block-device-mappings '[{"DeviceName":"/dev/sda1","Ebs":{"VolumeSize":16,"VolumeType":"gp3","DeleteOnTermination":true}}]' \
    --tag-specifications \
    "ResourceType=instance,Tags=[{Key=Project,Value=ai-research-course},{Key=Lesson,Value=wrk},{Key=Name,Value=ai-research-course-wrk}]" \
    "ResourceType=volume,Tags=[{Key=Project,Value=ai-research-course},{Key=Lesson,Value=wrk}]" \
    --query "Instances[0].InstanceId" --output text)

The script also writes ID and SG into small files in ~/lab-data/serving/wrk/, so the later steps can read them back.

The three sha256 fingerprints match the ones lesson 2 recorded for its files, so the box serves the same model. These are the first commands of the step:

  S=(ssh -i "$KEY" "${SSHO[@]}" ubuntu@"$IP")
  C=(scp -q -i "$KEY" "${SSHO[@]}")
  "${S[@]}" "curl -LsSf https://astral.sh/uv/install.sh | sh > /dev/null 2>&1 && \
    ~/.local/bin/uv python install 3.13.15 > /dev/null 2>&1 && \
    ~/.local/bin/uv venv --allow-existing --python 3.13.15 $B/venv > /dev/null 2>&1"
  "${C[@]}" "$HERE/examples/lat_requirements.txt" ubuntu@"$IP":$B/
  "${S[@]}" "~/.local/bin/uv pip install --python $B/venv/bin/python -q -r $B/lat_requirements.txt"

The full list of files to copy is in the copy step of wrk_aws_steps.sh, which also sets B, the folder on the box, and HERE, the lab folder. I ran this step twice on the same box. The first recording had too few rows and lost its first lines, so I ran the step again for the recording you see. The step is safe to repeat, because --allow-existing keeps the venv and the copies only overwrite files.

Step 6: stop the timers and look at the box. Ubuntu runs small jobs on a timer, such as checking for updates. One of those starting in the middle of a run would land straight in the results, so the script stops every timer first. Then it prints what the box is (lscpu), how many cores it has (nproc), how busy it has been (uptime), and every installed package with its version (pip freeze).

A real terminal recording of step 6, with the address replaced by a placeholder. Over ssh: stop every active timer and the unattended-upgrades service, after which systemctl prints 0 timers listed; lscpu prints architecture aarch64, 1 CPU, vendor ARM, model name Neoverse-V1, 1 thread per core, 1 core per socket; nproc prints 1; uptime prints up 3 min with load average 0.12, 0.09, 0.03. Then the box info is saved, and Python 3.13.15 and the full pip freeze are printed, including fastapi 0.142.2, numpy 2.5.3, scikit-learn 1.9.1, threadpoolctl 3.7.0, uvicorn 0.54.0 and uvloop 0.23.0.

The box reports one CPU, a Neoverse-V1, and the same package versions as lessons 2 to 4. These are the commands that stop the timers and look:

  ssh -i ~/lab-data/serving/keys/ai-research-course-wrk.pem "${SSHO[@]}" ubuntu@"$IP" \
    'for t in $(systemctl list-units --type=timer --state=active --no-legend --plain | cut -d" " -f1); do
       sudo systemctl stop "$t"; done; sudo systemctl stop unattended-upgrades.service
     systemctl list-timers --no-pager | tail -2'
  ssh -i ~/lab-data/serving/keys/ai-research-course-wrk.pem "${SSHO[@]}" ubuntu@"$IP" 'lscpu | head -12; nproc; uptime'

Check three things: systemctl says 0 timers listed, nproc prints 1, and the versions match lat_requirements.txt.

Step 7: start the schedule, then leave. The schedule starts in the background with setsid nohup, so it keeps running after SSH disconnects. Then I log out, and nobody logs in again until it ends.

A real terminal recording of step 7, with the address replaced by a placeholder. Over ssh: cd wrk, then setsid nohup bash wrk_schedule.sh 5 8 8 with its output sent to raw/schedule.log, in the background; two seconds later the log shows schedule start K=5 sat=8 open=8.

The log's first line appeared two seconds after the start, and then the SSH session closed. This is the command that starts the schedule:

  ssh -i ~/lab-data/serving/keys/ai-research-course-wrk.pem "${SSHO[@]}" ubuntu@"$IP" \
    "cd wrk && (setsid nohup bash wrk_schedule.sh $K $SAT_S $OPEN_S > raw/schedule.log 2>&1 < /dev/null &); \
     sleep 2; cat raw/schedule.log"

The three numbers are the rounds, and the seconds of each closed-loop and each open-loop run: 5, 8 and 8.

A real terminal recording of step 12, with the address replaced by a placeholder. Over ssh, the box records its info after the schedule and prints the size of its raw folder, 9.6M; then rsync copies the box's wrk/raw folder to my laptop; ls counts 519 files; du prints 9.5M for the copy.

The 519 files came to 9.5 MB, small because the client already stored every run in a compact form. This is the copy command:

  rsync -a -e "ssh -i $KEY ${SSHO[*]}" ubuntu@"$IP":wrk/raw/ ~/lab-data/serving/wrk/raw/

Step 13: delete the box, the key and the door rule, and check. A box you forget keeps costing money. The script deletes the box, waits until AWS says "terminated", then deletes the key pair and the security group. Then it asks AWS again. The box must say terminated. The counts of key pairs, security groups, disks with this lesson's tag, and course boxes still alive must all be 0.

A real terminal recording of step 13, with the instance and group ids replaced by placeholders. terminate-instances prints shutting-down; wait instance-terminated; delete-key-pair prints true; delete-security-group prints true. Then the checks: describe-instances prints terminated; the count of key pairs named ai-research-course-wrk is 0; the count of security groups of that name is 0; the count of volumes tagged Lesson=wrk is 0; the count of course instances not terminated is 0. Last, the key file is removed.

Every count printed 0 and the box printed terminated, so nothing from this lesson was left in the account. These are all the teardown commands:

  aws ec2 terminate-instances --instance-ids "$ID" --query "TerminatingInstances[0].CurrentState.Name" --output text
  aws ec2 wait instance-terminated --instance-ids "$ID"
  aws ec2 delete-key-pair --key-name ai-research-course-wrk --query Return
  until aws ec2 delete-security-group --group-id "$SG" --query Return 2>/dev/null; do sleep 10; done
  aws ec2 describe-instances --instance-ids "$ID" --query "Reservations[0].Instances[0].State.Name" --output text
  aws ec2 describe-key-pairs --filters "Name=key-name,Values=ai-research-course-wrk" --query "length(KeyPairs)"
  aws ec2 describe-security-groups --filters "Name=group-name,Values=ai-research-course-wrk" \
    --query "length(SecurityGroups)"
  aws ec2 describe-volumes --filters "Name=tag:Lesson,Values=wrk" --query "length(Volumes)"
  aws ec2 describe-instances --filters "Name=tag:Project,Values=ai-research-course" \
    "Name=instance-state-name,Values=pending,running,stopping,stopped,shutting-down" \
    --query "length(Reservations[].Instances[])"
  rm -f ~/lab-data/serving/keys/ai-research-course-wrk.pem

The security group can only be deleted once the box is fully gone, so the script retries that command every ten seconds until AWS accepts it.

One more check: every score at 2, 4 and 8 threads was equal to the bit to the score with 1 thread. That held one row at a time and all rows at once. Each row is scored by one thread from start to end, so the thread count only decides who does the work.

During a pure-Python loop it kept 0.498.
switch interval

Here is a quick reasoned check on sorted(), not a measurement. If the side thread got one 0.005-second turn per 37-millisecond sort, its share would be about 0.005 / 0.042, or 0.12. I measured 0.17, the same size.

Why did one predict call take five times as long? This part is my reading of the numbers, worked out after the results; the lab did not design a test for it.

A hand-drawn timeline against time, with two rows. The predict row repeats a short box marked tree followed by a long box marked waits for lock. The side thread row has a long box marked runs Python under each wait. Below: the model has 54 trees, and predict_proba scores them one tree at a time; it lets the GIL go for each tree's work, then takes it back; when it takes it back while the side thread is running Python, it waits for the side thread's switch interval, 5 ms by default; 54 x 5 ms = 270.0 ms; one call measured 270.2 ms; short boxes are drawn wider than to scale. Caption: the model's own work was short; almost all of the call was waiting to get the lock back.

predict_proba lets the lock go for each tree's work and takes it back again and again. Each time it takes it back while the side thread is busy running Python, it has to wait. A busy Python thread hands the lock over only when its switch interval runs out, about 5 ms. This kind of queue for a lock is often called a lock convoy.

On the box it added up to almost exactly one interval per tree. The model has 54 trees, 54 x 5 ms is 270 ms, and one call took 270.2 ms. The share it predicts for the side thread, 1 - 52.6 / 270.2 = 0.805, matches the 0.797 I measured.

I checked the direction of this, not its timing, on my 10-core laptop, where other programs were running. With a busy Python thread beside it, predict_proba on all rows got slower as I raised the switch interval, even though there were free cores. It took 120.5, 204.5, 846.6 and 3,147.1 ms at switch intervals of 0.5, 1, 5 and 20 ms, against 34.5 ms alone. The script is wrk_convoy_check.py, written by the lesson's reviewer, and the numbers are in results/wrk-convoy-check.json. There the slowdown was about three switch intervals per tree, not one, so "one per tree" is what the box showed, not a rule.

So the GIL did matter, but not the way I first thought. On one core, letting the lock go cannot make a second core appear. On a machine with several cores, the tree work can overlap on other cores, but every retake of the lock still waits behind a busy Python thread. So a busy thread beside the model slows every call, on any number of cores.

The in-handler predict time tells the same story from the inside. Each answer carried how long its own predict_proba call took. At 0.3 times S, the median was 604 us for 1x1 and 1,739 us for 1x4, 2.88 times as long. For 4x8 it was 25,041 us: a one-row call that should take half a millisecond took 25, because its threads kept waiting behind the other workers' threads.

3: 1,048
3: 1,049
4x23: 3832: 4643: 3813: 3802: 464
4x43: 2563: 2534: 2403: 2523: 254
4x82: 1922: 1942: 1913: 1642: 196

With one thread per worker, this hardly changed the answers per second, because the workers all wait for the same core anyway. With several threads per worker it did: every worker that got connections brought its helper threads, and the rounds with fewer busy workers gave more answers. On a machine with several cores it would matter: an idle worker is an idle core. If you use several workers, check that each one actually gets work. The answer's process id is a cheap way to see it.

On this box, with 2 GiB of memory, four workers used almost half of it. On one core, those four copies bought nothing.

Memory grows with workers on any machine. Four workers held about four copies here, and would on a bigger box too.

  • "S3. 1x2 and 1x4 closed-loop answers per second below 1x1's, told apart; 1x4 at most 0.8 x 1x1." Right: 0.654 and 0.454.

  • "S4. the lowest closed-loop answers per second of the 11 settings is 4x4's or 4x8's, at most 0.6 x 1x1's." Right: 4x8, 0.179.

  • "S5. at 0.6 x S, 1x1 has the lowest median p99 of all 11, and every setting with T at least 2 has a higher p99 than 1x1, told apart." Right: 8 of 8.

  • "S6. every setting keeps up at 0.3 x S in all 5 rounds." Wrong: 2x4, 4x4, 1x8 and 4x8 did not.

  • "S7. at 0.3 x S, the median in-handler predict time of 1x4 is at least 1.5 x 1x1's." Right: 2.88 times.

  • "M1. RSS per worker process between 80 and 250 MB; total RSS of 4x1 (all its processes) at least 3.5 x 1x1's." Right: 218.2 to 228.3 MB, and 4.14 times.

  • "M2. the sum of Pss is below the sum of RSS for every setting with W > 1." Right.

  • "C1. every returned score equals the box's single-row score to the bit." Right: 1,109,009 of 1,109,009.

  • predict_proba
    r"""Processes, threads, or both? Try OpenMP threads, Python threads and worker processes on YOUR machine.
    
    Lesson 5 of 'Serving and Inference Basics'. It uses lesson 2's Python environment, made once inside this folder
    (Python 3.13 is what the lab used):
        python3 -m venv venv-sv
        source venv-sv/bin/activate              # Windows: venv-sv\Scripts\activate
        pip install -r lat_requirements.txt      # numpy, pandas, pyarrow, scikit-learn, threadpoolctl, ...
    It also needs the shop data from the features chapter: run python ../../features/fetch_data.py once first. Then:
        python wrk_demo.py                  # train, then print parts 2 to 4 below, numbered 1 to 3
        python wrk_demo.py --save out.json  # and save the numbers
    If a package is missing, it prints one line saying why and stops.
    The timings it prints come from YOUR machine, with its number of cores, as it is right now: other programs move
    them. The lab measured a 1-core box; your machine probably has more cores, so its numbers answer a different
    question. Read them for what they show about your machine, not as a check on the lab.
    It writes nothing except out.json if you ask.
    
    Design, written 2026-10-04 after the lab's design (workers_and_threads.py) and before this file first ran.
    Fixed 2026-10-04 AFTER the run on the box: the GIL check (step 3) built its list with a new random generator for every number,
    so all 200,000 numbers were the same and sorting them was quick; it now draws them from one generator. The box's
    stored run (results/wrk-demo-run.txt) was made before this fix.
      1. Train the features chapter's model (it must score test AP 0.5450), as lesson 4's bat_demo.py does.
         OMP_NUM_THREADS is set to 1 at the top, before numpy or scikit-learn load, so that every thread starts with
         one OpenMP thread and threadpoolctl can then raise it for part 2 (scikit-learn only goes above the core count
         when OMP_NUM_THREADS is set).
      2. OpenMP threads inside one call: predict_proba on 1 row (500 calls) and on all 26,851 rows (10 calls) with
         1, 2 and 4 OpenMP threads (threadpoolctl). Median time per call.
      3. The GIL: a side thread counts in pure Python for 1 s while the main thread (a) sleeps, (b) calls
         predict_proba on all rows again and again, (c) calls sorted() on 200,000 floats again and again. The side
         thread's count as a share of its count while (a).
      4. Python threads against worker processes: 3,000 one-row predict_proba calls shared out over 1, 2 and 4 Python
         threads (one OpenMP thread each), then over 1, 2 and 4 worker processes (a multiprocessing pool, started and
         warmed before the clock). Rows per second.
    
    Author: Roni Das
    Created: 2026-10-04
    """
    import os
    
    os.environ["OMP_NUM_THREADS"] = "1"   # before numpy and scikit-learn load: see the design, step 1
    
    import importlib.util  # noqa: E402
    import json  # noqa: E402
    import multiprocessing as mp  # noqa: E402
    import pickle  # noqa: E402
    import random  # noqa: E402
    import sys  # noqa: E402
    import threading  # noqa: E402
    import time  # noqa: E402
    from concurrent.futures import ThreadPoolExecutor  # noqa: E402
    from pathlib import Path  # noqa: E402
    from time import perf_counter, perf_counter_ns  # noqa: E402
    
    NEEDED = ("numpy", "pandas", "pyarrow", "sklearn", "threadpoolctl")
    missing = [m for m in NEEDED if importlib.util.find_spec(m) is None]
    if missing:
        sys.exit(f"wrk_demo.py needs {', '.join(missing)}: make the environment in its docstring "
                 f"(pip install -r lat_requirements.txt), then run it with that environment's python.")
    
    import numpy as np  # noqa: E402
    from threadpoolctl import threadpool_limits  # noqa: E402
    
    CALLS = 3000
    
    
    def build():
        """The features chapter's model and its 26,851 test rows (lesson 4's bat_demo.py build step)."""
        from sklearn.metrics import average_precision_score
        here = Path(__file__).resolve().parent
        sys.path.insert(0, str(here.parents[1] / "features"))
        import task
        from what_a_feature_is import HAND_COLS, hgb, joined
    
        ev = task.load_events()
        lab_tr, _, lab_te = task.splits(ev)
        tr = joined(ev, lab_tr, task.TRAIN_CUTOFFS)
        te = joined(ev, lab_te, task.TEST_CUTOFFS)
        model = hgb(0).fit(tr[HAND_COLS].to_numpy(float), tr["label"].to_numpy())
        X = np.ascontiguousarray(te[HAND_COLS].to_numpy(np.float64))
        p = model.predict_proba(X)[:, 1]
        y, cut = te["label"].to_numpy(), te["cutoff"].to_numpy()
        ap = float(np.mean([average_precision_score(y[cut == c], p[cut == c]) for c in np.unique(cut)]))
        assert round(ap, 4) == 0.5450, ap
        return model, X, ap
    
    
    def omp_threads(model, X) -> dict:
        out = {}
        for t in (1, 2, 4):
            with threadpool_limits(limits=t, user_api="openmp"):
                for name, rows, calls in (("1 row", X[:1], 500), ("all rows", X, 10)):
                    for _ in range(5):
                        model.predict_proba(rows)
                    ts = []
                    for _ in range(calls):
                        a = perf_counter_ns()
                        model.predict_proba(rows)
                        ts.append(perf_counter_ns() - a)
                    out[f"{t} {name}"] = float(np.median(ts)) / 1000
        return out
    
    
    def gil(model, X) -> dict:
        count, stop = [0], threading.Event()
    
        def side():
            n = 0
            while not stop.is_set():
                for _ in range(1000):
                    n += 1
                count[0] = n
    
        rng = random.Random(0)
        data = [rng.random() for _ in range(200_000)]   # one generator: 200,000 different numbers
        th = threading.Thread(target=side, daemon=True)
        th.start()
        time.sleep(0.2)
        rate = {}
        for name in ("alone", "predict", "sort"):
            c0, t0 = count[0], perf_counter()
            while perf_counter() - t0 < 1.0:
                if name == "alone":
                    time.sleep(0.05)
                elif name == "predict":
                    model.predict_proba(X)
                else:
                    sorted(data)
            rate[name] = (count[0] - c0) / (perf_counter() - t0)
        stop.set()
        return {k: rate[k] / rate["alone"] for k in ("predict", "sort")}
    
    
    def _chunk(args):
        model, rows = args
        with threadpool_limits(limits=1, user_api="openmp"):
            for r in rows:
                model.predict_proba(r[None, :])
        return len(rows)
    
    
    _M = {}
    
    
    def _init(blob):
        _M["m"] = pickle.loads(blob)
    
    
    def _proc_chunk(rows):
        for r in rows:
            _M["m"].predict_proba(r[None, :])
        return len(rows)
    
    
    def threads_vs_processes(model, X) -> dict:
        rows = X[:CALLS]
        out = {}
        for n in (1, 2, 4):
            parts = np.array_split(rows, n)
            with ThreadPoolExecutor(n) as ex:
                list(ex.map(_chunk, [(model, p[:20]) for p in parts]))
                a = perf_counter()
                list(ex.map(_chunk, [(model, p) for p in parts]))
                out[f"threads {n}"] = CALLS / (perf_counter() - a)
        blob = pickle.dumps(model)
        for n in (1, 2, 4):
            parts = np.array_split(rows, n)
            with mp.get_context("spawn").Pool(n, initializer=_init, initargs=(blob,)) as pool:
                pool.map(_proc_chunk, [p[:20] for p in parts], chunksize=1)
                a = perf_counter()
                pool.map(_proc_chunk, parts, chunksize=1)
                out[f"processes {n}"] = CALLS / (perf_counter() - a)
        return out
    
    
    def main() -> None:
        try:
            load = f"{os.getloadavg()[0]:.2f}"
        except (AttributeError, OSError):
            load = "not available on this system"
        print(f"Every time below was measured on THIS machine ({os.cpu_count()} cores, 1-minute load {load}).")
        model, X, ap = build()
        print(f"model trained: test AP {ap:.4f}; {len(X):,} test rows")
    
        o = omp_threads(model, X)
        print("\n1. OpenMP threads inside one predict_proba call, median microseconds (us) per call")
        print("   threads      1 row    all rows")
        for t in (1, 2, 4):
            print(f"   {t:7d} {o[f'{t} 1 row']:10.1f} {o[f'{t} all rows']:11.1f}")
    
        g = gil(model, X)
        print("\n2. the GIL: a side thread's pure-Python count, as a share of its count when the main thread sleeps")
        print(f"   main thread in predict_proba(all rows): {g['predict']:.2f}")
        print(f"   main thread in sorted(200,000 floats):  {g['sort']:.2f}")
    
        tp = threads_vs_processes(model, X)
        print(f"\n3. {CALLS:,} one-row predict_proba calls, shared out; rows per second")
        print("   how many     Python threads   worker processes")
        for n in (1, 2, 4):
            print(f"   {n:8d} {tp[f'threads {n}']:16.0f} {tp[f'processes {n}']:18.0f}")
        if "--save" in sys.argv:
            out = {"load": load, "cores": os.cpu_count(), "test_ap": ap, "omp_us": o, "gil_share": g, "rows_per_s": tp}
            Path(sys.argv[sys.argv.index("--save") + 1]).write_text(json.dumps(out, indent=1))
    
    
    if __name__ == "__main__":
        main()
    

    You saw the demo running on the lab's box in step 11, after the schedule. That exact run is stored in results/wrk-demo-run.txt.

    One mistake, found after the box was gone. In that run, the GIL check sorted a list of 200,000 numbers that were all the same. My code made a new random generator for every number, and each new generator gave the same first number. A list of equal numbers sorts very quickly, so the side thread got its turn often and kept 0.48 of its speed.

    The lab's own check used different numbers and got 0.173. I fixed the demo after the box was deleted. The code above draws all the numbers from one generator, and the stored box run is from the earlier version. The other lines of the box run are from code that did not change.

    On the box, the rest told the same story as the lab. One row took 509.2 us with 1 thread and 1,653.4 with 4. With one core, one worker process gave 1,977 rows a second and four gave 1,411, and Python threads gave 1,985 with one and 1,843 with four.

    And here is the same demo in VS Code on my laptop, run with python wrk_demo.py in the examples folder with venv-sv active.

    A real screenshot of VS Code's terminal on my laptop, in the venv-sv environment, after running python wrk_demo.py. The first line says every time was measured on this machine, with 10 cores and a 1-minute load of 5.39; the model trained to test AP 0.5450 on 26,851 test rows. Part 1, median microseconds per predict_proba call with 1, 2 and 4 OpenMP threads: one row 171.1, 218.5 and 722.9; all rows 38775.4, 21550.1 and 16153.4. Part 2, the side thread's share of its speed: 1.00 while the main thread ran predict_proba, 0.26 while it ran sorted(). Part 3, rows per second with 1, 2 and 4: Python threads 4888, 4220 and 3052; worker processes 4535, 5905 and 6316.

    This run is from my laptop, not the box. My laptop has 10 cores, and other programs were busy on it: its 1-minute load was 5.39. So these numbers answer a different question from the lab's, and I do not compare them with the box one by one.

    Look instead at the directions, which change once there are free cores. With more OpenMP threads, all rows got faster, because each helper had a core of its own, but one row still got slower. The side thread kept all of its speed during predict_proba, 1.00, because it had its own core. The demo does not print how long predict_proba itself took meanwhile, and it slows a lot, as the GIL slide explains. During sorted(), which holds the lock, the side thread still dropped to 0.26.

    More worker processes gave more rows a second, because each process had its own core and its own lock, while more Python threads in one process gave fewer. Your own machine will give other numbers; the directions are what to look at.

    /proc

    wrk_files/wrk_predict.py is Part P, and wrk_files/wrk_gil.py is Part G. wrk_files/wrk_run_once.sh starts a fresh service with W workers and T threads, and waits for every worker and for an idle core. Then it records the load and the system log, runs the client and stops the service. wrk_files/wrk_schedule.sh runs everything in order, five rounds, each in its own shuffled order.

    wrk_report.py does not import any of the above. It reads the small copy of the raw files with its own code and recomputes every number. It also checks the guesses, the demo's stored run, the fact-check, the teardown and the redaction, and it writes the playground.

  • Check that every worker gets work. Here, with four workers, one or two often answered nothing.