Serving And Inference

Batching Requests: Does It Help a Small CPU Model?

0 of 31 complete

0%

Contents

Back|Serving And InferenceBatching Requests: Does It Help a Small CPU Model?
1/31
86 min left
  1. Home
  2. AI Engineering: Data, RAG and Agents
  3. Serving and Inference Basics
  4. Batching Requests: Does It Help a Small CPU Model?
Prerequisites
Tail Latency: What Sits in the Slowest 1% of RequestsrequiredLatency Anatomy: Where the Time in One Prediction Request GoesrequiredWhat a Feature Is: A Better Model or a Better Feature?required
Related Topics
Model Signatures: The Right Numbers in the Wrong Shape, and What a Schema Check CatchesPackaging, Registry and Versioning
1 of 31
Previous lessonTail Latency: What Sits in the Slowest 1% of Requests

System Design

  • Foundation
  • Intermediate
  • Advanced
  • Capstone

AI Engineering

  • Foundation
  • Data, RAG and Agents
  • Evaluation, LLM Ops and Security

systemdesign.academy

  • Home
  • Glossary
  • Interview prep
  • Reviews
  • About
  • Privacy
  • Terms

The Shuttle That Waits for More People

Let me start at an airport.

Outside the terminal, a small shuttle bus takes people to the hotels. It has 20 seats. The driver has one rule: leave when all 20 seats are full, or 10 minutes after the first person sits down, whichever comes first.

A flat illustration of a young man in a sage hoodie and dark trousers, standing with his hands in his pockets, like someone waiting for a ride. Below the picture: the airport shuttle leaves when its 20 seats are full, or 10 minutes after the first person sits down, whichever comes first; one trip carries many people for about the cost of one; the first person in pays with waiting.

Think about what this rule does. On a busy evening the bus fills in two minutes and leaves. One trip carries 20 people for about the cost of driving one. Without the rule, the driver would make 20 trips, and a long line would form at the door.

On a quiet morning it is different. The first person sits down and waits. Nobody else comes. After 10 minutes the bus leaves with one person in it. That person waited 10 minutes for nothing, and the trip cost the same as a full one.

There is a third way to run the bus. The driver does not wait at all. When the bus comes back, everyone standing at the door gets on, and it leaves at once. On a quiet morning it leaves with one or two people. On a busy evening it leaves full, because people gathered while it was away.

A prediction service can follow the same kinds of rules. Instead of asking the model for one answer at a time, it can collect several requests and ask the model once for all of them. In this lesson I build those rules into a real service and measure both sides: how much work they save, and how long they make people wait.

Where This Lesson Starts

This is lesson 4 of the chapter on serving a model. Lesson 2, latency anatomy, took one request apart and found that the model's predict_proba was the largest step. Lesson 3, tail latency, looked at the slowest requests. Both lessons sent one request at a time, so no request ever waited for another one.

This lesson lets requests overlap. Batching cannot work without this: a request can only share a model call with another request if both are waiting at the same moment.

I use the same model, the same service code and the same kind of rented machine as lessons 2 and 3, so I do not explain them again. In short: the model is the features chapter's gradient boosted tree model. It scores one customer of a real online shop at the start of a month, and it scored a test AP of 0.5450. The service is a small web service built with FastAPI and run by uvicorn, and it reads the customer's six features from SQLite. If any of that is new, please read lesson 2 first.

The question is short. Does batching help a small model that runs on a CPU, and what does it cost in ? I answer it in two parts. Part A times the model alone, with no web server, on batches of different sizes. Part B puts a batcher inside the service and sends it real traffic.

The Words You Need First

Please read this slide slowly if any word is new. Every slide after it uses these words.

A hand-drawn grid of twelve cards defining batch, batch size, per-call time, per-row time, micro-batcher, W the wait, B the cap, throughput, saturation (S), open loop, Poisson arrivals and keeps up, each with the one-line meaning used in the text. Below: request, latency, p50, p99 and microsecond (us) mean what they meant in lessons 2 and 3; 1 ms = 1,000 us.

A batch is a group of rows given to the model in one call. Think of a photocopier. Copying one page means walking to the machine, lifting the lid, pressing start, and walking back. Copying ten pages costs the same walk and only a little more machine time. The walk is the fixed cost of a call. The machine time for each page is the per-row cost.

The batch size is how many rows are in one call. The per-call time is how long one call takes. The per-row time is the per-call time divided by the batch size.

A micro-batcher is a small part of the service that collects requests as they arrive and scores them together. Mine has two dials. B is the cap: the most requests in one batch, like the 20 seats of the shuttle. W is the wait: the longest the first request in a batch waits for more to come, like the 10 minutes. With W = 0, the batcher never waits on purpose. It takes whatever has gathered while it was busy, like the third way to run the bus.

is how many answers per second the service actually gives. Saturation is the most it can give. I call the saturation of the service with no batching S.

An open loop test sends requests on their own clock, like customers who walk in when they want, whether or not the shop is busy. Poisson arrivals are arrivals at random moments with a steady average rate. Sometimes three come almost together, sometimes there is a long gap. A service keeps up when it answers at least 95% as fast as requests arrive. So up to 5% of the requests may still pile up in a run that keeps up. If it answers slower than that, a queue grows and never empties.

The Headline: Batching Kept the Service Up, and Helped Even When It Was Quiet

Here is the main result first. Each line is one way of running the service. Each point is the p99 at one arrival rate, the median of five rounds. Latency is counted from the moment a request arrived, so every kind of waiting counts.

A line chart of p99 latency on a log scale, from 1 ms to 100 s, against the arrival rate as a multiple of S: 0.5x, 0.8x, 1x, 1.25x, 1.6x, 2x and 2.5x, not to scale. Three lines: no batching; B=32 with W=0 ms; B=32 with W=2 ms. The no-batching line starts at about 5 ms, rises steeply after 0.8x and reaches about 13 s at 2.5x. The two batching lines stay between about 2.5 ms and 13 ms across all rates. Below: p99 at 0.5x S (549 a second): no batching 5.26 ms, B=32 W=0 2.50 ms; at 1x S (1,098 a second): 721 ms against 3.73 ms; at 2.5x S: 12,782 ms against 13.0 ms; S = 1,098 answers a second. Caption: the same model, the same box, the same arrivals: only the batcher differs.

With no batching, the service fell behind at S. At 1,098 arrivals a second its p99 was 721 milliseconds, and at 2.5 times that rate it was 12.8 seconds. Its queue never emptied, so those numbers only grow if the test runs longer.

With batches of up to 32 and no timed wait, the service kept up at every rate I tried, up to 2,746 arrivals a second. Its p99 stayed between 2.50 ms and 13.0 ms.

Even at half of S, that batcher cut the p99 in half, from 5.26 ms to 2.50 ms. I did not expect that, and a later slide explains it.

A timed wait cost latency at a low rate. With W = 2 ms the median was 3.11 ms against 1.18 ms. Its p99 (4.82 ms) was still lower than no batching's, but higher than W = 0's 2.50 ms.

So for this small model on one core, batching helped a lot, and the setting with no timed wait was the safe choice at every rate. The rest of the lesson shows where those numbers come from.

How the Lab Was Built

I wrote the lab's design into the docstring of scripts/labs/serving/batching_requests.py on 2026-10-04, before any timed run. It names every measurement, every setting, the rule for "told apart from noise", and eleven guesses.

A page in five labelled zones, titled two questions, one box, 605 runs. Part A, the model alone: predict_proba on 1, 2, 4, up to 1,024 real test rows, 2,000 timed calls per size, 5 runs, no web server. Same answers: every test row scored one at a time, then in batches of every size: equal to the bit? Saturation first: the most answers per second the service gives with no batching (S), 5 closed-loop runs. Part B, the service: 17 settings: no batching, a queue with B=1, and B 8, 32, 128 with W 0, 1, 2, 5, 10 ms; 7 rates from 0.5x to 2.5x S; random arrivals, 8 s each, 5 rounds. The rule: a setting differs from no batching only if all 5 paired rounds agree in sign and the values do not overlap. Caption: eleven guesses written first, checked as 13 statements; the main one: below S, waiting for a batch costs latency; above S, batching keeps the service up.

Part A times predict_proba alone, with no web server, on batches of 1, 2, 4 and so on up to 1,024 real test rows. Each size gets 2,000 timed calls, and the whole thing runs five times, each time in a fresh Python process with the sizes in a new random order.

The saturation runs find S. Sixteen connections each send a new request the moment their last answer arrives, for 8 counted seconds, with no batching. This is a closed loop: the client waits for answers before sending more, so it measures the most the service can do.

Part B puts the batcher in the service. It tries 17 settings. None is lesson 2's handler, which scores its one row itself. A queue with B = 1 sends every request through the batcher machinery but never groups two, so it measures what the machinery alone costs. The other 15 are B = 8, 32 or 128 with W = 0, 1, 2, 5 or 10 ms. B = 1 with a W above 0 would make no sense: a batch of one is full the moment its only request arrives, so it never waits.

Each setting meets seven arrival rates: 0.5, 0.8, 1, 1.25, 1.6, 2 and 2.5 times S. Each run lasts 8 seconds of arrivals and starts a fresh service, and there are five rounds. Inside a round the 119 runs go in a shuffled order. So I can compare each setting with no batching on the very same traffic, round by round.

The Machine, and the Client That Shares It

Like lessons 2 and 3, every timing comes from a small machine I rented from Amazon Web Services. None comes from my laptop, which is always busy with other work.

Five rows, each with a logo. EC2 c7g.medium, us-east-1: 1 vCPU (Neoverse-V1), 2 GiB, not burstable; on for 2.735 hours, $0.0993. Ubuntu 24.04, arm64: core at least 99.5% idle before every one of 605 runs; host steal at most 0.00%; systemd timers stopped (cron still ran). Python 3.13.15, scikit-learn 1.9.1: numpy 2.5.3; the model has 54 trees of at most 31 leaves; OpenMP left at its default, one thread on one core. FastAPI 0.142.2, uvicorn 0.54.0: one worker, uvloop 0.23.0, httptools 0.8.0; the batcher is one asyncio task. SQLite, lesson 2's feature table: 26,851 rows; the same model file and request list as lessons 2 and 3, same sha256. Caption: every timing I quote comes from this box, never from my laptop.

The box is an EC2 c7g.medium. AWS's own API lists it with 1 virtual CPU, 2,048 MiB of memory and no bursting, and its AWS page says C7g instances use Graviton3 processors. My account allows only one of these at a time, so, as in lessons 2 and 3, the client and the service share that one core.

That matters more here than before. In lessons 2 and 3 the client sent one request and waited. Here the client sends requests on its own clock, so it works at the same moments the service does. Every request costs some core time in the client too. So S, the most the service answered with no batching, is really "the most the service and its client together could do on one core". The client's work per request is the same in every setting, so the comparisons between settings are fair. But the absolute numbers are lower than a service would reach with its client on another machine.

Three words from the figure. Steal is time the physical machine under my rented box gives to other customers' boxes; it was 0.00% in every run. OpenMP is a library that lets a program split work across several cores; scikit-learn uses it inside predict_proba, and on one core it uses one thread. An asyncio task is one job inside Python's asyncio system, which runs many small jobs on one thread by switching between them while each waits.

Before every run, the core had to be at least 95% idle over two seconds. I did not wait for the one-minute load average to fall, as lesson 3 did, because it takes minutes to decay and there were 605 runs. The load is still recorded before and after every run.

The Whole Lab in One Picture

Before the recordings, here is the whole lab as one numbered sequence. The numbers on the arrows match the numbers on the recordings that follow.

A sequence diagram with five lifelines: my laptop, AWS, the box, client and service, and thirteen numbered arrows. 1, my laptop asks AWS whether any box is running. 2, key, group, SSH rule. 3, run-instances. 4, my laptop to the box: wait, then SSH. 5, Python, files. 6, stop timers, look. 7, start, log out. 8, my laptop to AWS: read the console. 9, client to service: requests at random. 10, service to client: answers. 11, my laptop to the box: the demo. 12, the box to my laptop: raw files. 13, my laptop to AWS: terminate, delete. Below: arrows 9 and 10 repeat 7,179,151 times in the open-loop runs, with nobody connected; the client and the service are two processes that share the box's one core; step 8 reads the box's serial console through AWS while 9 and 10 run, and never touches the box. Caption: steps 1 to 8 and 11 to 13 are the recordings that follow, numbered the same; all on the box that measured the numbers.

Steps 1 to 7 happen once: check, rent the box, set it up, and start the schedule. Step 8 is how I watched the schedule without touching the box. Arrows 9 and 10 are the timed work, millions of times, with nobody logged in. Steps 11 to 13 come after the schedule: the student demo, the copy of the results back to my laptop, and the deletion of everything.

Is anything pinned to the core? No. With one core there is nothing to choose. The client's two threads and the service take turns on it, and the operating system switches between them. That switching is part of every here, in every setting alike.

Build the Lab Box Yourself

The rented box is part of the lab, so I show every step of it. You can rent the same kind of box and run the same schedule. The flow figure numbers the steps, and each recording below carries the same number.

Which box you see. Every recording comes from the box that measured this lesson's numbers. I planned the order so that no recording happened during a timed run. Steps 1 to 7 come before the schedule. Step 8 reads the box from outside without logging in. Steps 11 to 13 come after the schedule finished.

The path to follow. All the steps live in one file, scripts/labs/serving/bat_files/bat_aws_steps.sh. From the scripts/labs/serving folder, run bash bat_files/bat_aws_steps.sh check, then access, launch, wait, copy, quiet, start and watch; after the schedule, demo, fetch and, at the end, teardown. The code blocks on these slides are that file's lines, copied exactly, so you can read what each step does. If you would rather paste them by hand, paste this variables block first, in the same terminal. Each step also saves the group id, the box's id and its address to a small file, so the steps work in separate terminals too.

Steps 1 to 3: Check, Lock the Door, Rent the Box

Step 1: is any other course box running? My account allows one small box of this kind at a time, and another lesson may be using it. So the first command counts the course's boxes that are not yet deleted. It must print 0.

A real terminal recording of step 1. The shell prints the command aws ec2 describe-instances with filters for the tag Project=ai-research-course and the states pending, running, stopping, stopped and shutting-down, asking for the number of instances; the answer is 0.

aws ec2 describe-instances --filters "Name=tag:Project,Values=ai-research-course" \
  "Name=instance-state-name,Values=pending,running,stopping,stopped,shutting-down" \
  --query "length(Reservations[].Instances[])"

Step 2: a key and a locked door. A key pair is how you prove to the box that you are allowed in. AWS keeps one half, and you keep the other half in a file that only you can read. A security group is a door rule for the box. Mine opens only port 22, the SSH port, and only to my own address.

A real terminal recording of step 2, with the address, network and group ids replaced by placeholders. create-key-pair for ai-research-course-bat with its tags writes the key material to the key file without printing it; chmod 600 on the key file; curl reads my address; describe-vpcs finds the default network; create-security-group makes the group ai-research-course-bat with its tags; authorize-security-group-ingress prints a small table: from my address /32, port 22. Last line: the key file, readable only by its owner, 388 bytes.

mkdir -p "$(dirname "$KEY")"
aws ec2 create-key-pair --key-name "$NAME" --key-type ed25519 \
  --tag-specifications "ResourceType=key-pair,Tags=[{Key=Project,Value=ai-research-course},{Key=Lesson,Value=bat}]" \
  --query KeyMaterial --output text > "$KEY"
chmod 600 "$KEY"
MYIP=$(curl -sf https://checkip.amazonaws.com)
VPC=$(aws ec2 describe-vpcs --filters Name=isDefault,Values=true --query "Vpcs[0].VpcId" --output text)
SG=$(aws ec2 create-security-group --group-name "$NAME" --vpc-id "$VPC" \
  --description "lesson 4 batching, ssh from one address" \
  --tag-specifications "ResourceType=security-group,Tags=[{Key=Project,Value=ai-research-course},{Key=Lesson,Value=bat}]" \
  --query GroupId --output text)
aws ec2 authorize-security-group-ingress --group-id "$SG" --protocol tcp --port 22 --cidr "$MYIP/32" \
  --query "SecurityGroupRules[].{port:FromPort,from:CidrIpv4}" --output table
echo "SG=$SG" >> "$STATE/vars"
ls -l "$KEY"

Keep the key file outside any git folder. Mine lives in , and the last step deletes it.

Steps 4 to 7: Wait, Set Up, Look, Start

Step 4: wait until it runs. A new box takes a short while to start. The AWS tool can wait for you, and then the script tries SSH every five seconds until the box answers.

A real terminal recording of step 4, with the instance id and address replaced by placeholders. aws ec2 wait instance-running, then describe-instances prints running, c7g.medium, us-east-1b and the address. Then ssh to the box, which answers: ssh works, aarch64, Ubuntu 24.04.5 LTS.

aws ec2 wait instance-running --instance-ids "$ID"
IP=$(aws ec2 describe-instances --instance-ids "$ID" \
  --query "Reservations[0].Instances[0].PublicIpAddress" --output text)
echo "IP=$IP" >> "$STATE/vars"
until ssh "${SSH[@]}" -o ConnectTimeout=5 ubuntu@"$IP" true 2>/dev/null; do sleep 5; done
ssh "${SSH[@]}" ubuntu@"$IP" 'echo "ssh works:" $(uname -m) $(lsb_release -ds)'

Step 5: Python and the lab files. The box gets the same Python, 3.13.15, and the same pinned library versions as lessons 2 and 3, through uv, a fast installer for Python. Then the lab files go over with scp, a copy command that works through SSH: lesson 2's model file, feature table, request list and service, this lesson's files, and the files the demo needs. The last command builds the SQLite table on the box and prints the files' sha256 fingerprints, so you can check they match the ones on your computer.

A real terminal recording of step 5, with the address replaced by a placeholder and my home folder shortened to a tilde. Over ssh: make the folders; install uv, Python 3.13.15 and a venv; install the pinned packages from lat_requirements.txt. Then scp copies lesson 2's model, data and service files, this lesson's bat_files, the features chapter's files and data, and the demo. Last: build_store.py prints 26851 rows; the sha256 of model.pkl starts bec4d11e, of store.parquet 62aadc30, of requests.parquet 407d62e6; then the folder's list, including folders named raw-attempt1 and raw-attempt2.

ssh "${SSH[@]}" ubuntu@"$IP" "mkdir -p bat/raw bat/features/results bat/serving/examples lab-data/features"
ssh "${SSH[@]}" ubuntu@"$IP" "curl -LsSf https://astral.sh/uv/install.sh | sh > /dev/null 2>&1 && \
  ~/.local/bin/uv python install 3.13.15 > /dev/null 2>&1 && \
  ~/.local/bin/uv venv --allow-existing --python 3.13.15 bat/venv > /dev/null 2>&1"
scp -q "${SSH[@]}" examples/lat_requirements.txt ubuntu@"$IP":bat/
ssh "${SSH[@]}" ubuntu@"$IP" "~/.local/bin/uv pip install --python bat/venv/bin/python -q -r bat/lat_requirements.txt"
scp -q "${SSH[@]}" ~/lab-data/serving/lat/model.pkl ~/lab-data/serving/lat/store.parquet \
  ~/lab-data/serving/lat/requests.parquet lat_files/service.py lat_files/build_store.py lat_files/box_info.sh \
  bat_files/bat_*.py bat_files/bat_*.sh ubuntu@"$IP":bat/
scp -q "${SSH[@]}" ../features/task.py ../features/what_a_feature_is.py ubuntu@"$IP":bat/features/
scp -q "${SSH[@]}" ../features/results/data-manifest.json ubuntu@"$IP":bat/features/results/
scp -q "${SSH[@]}" ~/lab-data/features/retail.parquet ubuntu@"$IP":lab-data/features/
scp -q "${SSH[@]}" examples/bat_demo.py examples/lat_requirements.txt ubuntu@"$IP":bat/serving/examples/
ssh "${SSH[@]}" ubuntu@"$IP" "cd bat && OMP_NUM_THREADS=1 venv/bin/python build_store.py store.parquet store.sqlite \
  && sha256sum model.pkl store.parquet requests.parquet && ls"

Steps 8 and 11 to 13: Watch, Demo, Bring It Home, Delete

Step 8: watch without touching. While it runs, the schedule writes one progress line to the box's serial console after every run. The serial console is a log that AWS keeps for the box, and you can read it through AWS without logging in. So I could follow the runs and never disturb them.

A real terminal recording of step 8, with the instance id replaced by a placeholder. aws ec2 get-console-output with --latest, piped through grep BAT-PROGRESS and tail -12, prints progress lines with times: schedule start K=5 calls=2000 sat=8 open=8 at 05:35:25; load before the schedule 0.09 0.06 0.01 after waiting 45s; predict-r1 to predict-r5 done; bits done; sat-r1 and sat-r2 done at 05:39:15.

aws ec2 get-console-output --instance-id "$ID" --latest --output text | grep BAT-PROGRESS | tail -12

This recording comes from the first start of the schedule on this box, which later stopped working. The command is the same one I used to follow the run that produced the numbers. I explain the restart on the limits slide.

Step 11: the student demo, on the same box. After the schedule, I ran the demo you will meet later, so you can see what it prints on a quiet box. Its output went into a file at the same moment it went to the screen.

A real terminal recording of step 11, with the address replaced by a placeholder: python bat_demo.py on the box, its output also written to bat-demo-run.txt. It prints the machine, 1 CPU, load 0.24, and test AP 0.5450. Part 1: predict_proba per call from 509.4 us at 1 row to 2596.2 us at 1,024 rows, per row 2.54 us at 1,024. Part 2: all 26,851 rows equal to the bit in batches of 32 and in one call. Part 3: a tiny batcher with no web server, S = 1,963 calls a second; at 1.5 x S, one predict each had a p50 of 764293 us, batches of up to 32 with a 2 ms wait 3119 us.

ssh "${SSH[@]}" ubuntu@"$IP" 'cd bat/serving/examples && ../../venv/bin/python bat_demo.py \
  --save ~/bat/raw/bat-demo.json | tee ~/bat/raw/bat-demo-run.txt'

Step 12: bring the results home first. Lesson 3 lost a whole box of results because it was deleted before anyone copied them. So the rule is: the moment the runs finish, copy the raw files back, before anything else.

A real terminal recording of step 12, with the address replaced by a placeholder. Over ssh, the box records its info after the schedule and prints the size of its raw folder, 279M; then rsync copies the box's bat/raw folder to my laptop; ls counts 1826 files; du prints 278M for the copy.

The Lab's Report, Running

This is a real recording of the report script, bat_report.py. It ran on my laptop, but it times nothing. It reads the raw files the box measured and does all the arithmetic again with its own code.

A terminal recording of bat_report.py in eight numbered sections. 1: the box, c7g.medium, 1 vCPU, 2.73 h at $0.0363 an hour = $0.0993, and everything deleted; 605 runs, the core at least 99.5% idle before each. 2: predict_proba per call, 1 row 501.7 us, 1,024 rows 2628.8 us, 5.24 times one row, 195 times cheaper per row, all batched scores equal to the bit. 3: S = 1098.4 answers a second and the seven rates. 4: a table of p50 and p99 in ms for ten settings at 0.5x, 1x, 1.6x and 2.5x of S, with a star where a setting fell behind; no batching and the queue with B=1 are starred at the three high rates. 5: 7,179,151 of 7,179,151 service scores equal; the waits for W of 0, 1, 2, 5 and 10 ms. 6: 11 guesses right, 2 wrong. 7 and 8: the demo and the playground match. Last line: all 3693 checks agree with the stored lab.

The report reads the small copy of the raw files kept in the repo. It computes its own percentiles, its own , its own paired differences and its own verdicts. Then it compares every one with the lab's results file. If any number disagreed, it would stop with an error.

Section 1 is the quiet check. Before every one of the 605 runs, the core was at least 99.5% idle over two seconds, and the machine under my box took no measurable time from it. One SSH session was open at the start of the very first run of part A, the one that started the schedule as it logged out.

The system log had 153 lines during runs. Counted after the review: 86 came from the AWS management agent (SSM) in 5 runs, which tried to reach a web address and logged the error page it got back; 51 from cron, the older job scheduler, in 16 runs; 8 from sshd in 2 runs; 4 from the login manager in 2; 3 from systemd in 2; and 1 from the kernel.

I had stopped the systemd timers but not cron, so cron still ran.

Part A: One Row or a Thousand, Almost the Same Price

First, the model alone, with no web server. How long does one call to predict_proba take on 1, 2, 4 and up to 1,024 rows?

A line chart with log scales on both axes. The x axis is the rows in one call: 1, 2, 4, 8, 16, 32, 64, 128, 256, 512, 1024. The y axis is time from 1 us to 10 ms. The per-call line is almost flat near 500 us up to 32 rows, then rises to about 2.6 ms at 1,024. The per-row line falls steadily from about 500 us at 1 row to about 2.6 us at 1,024. Below: per call, 1 row 501.7 us, 32 rows 550.6 us, 1,024 rows 2628.8 us; per row, 501.7 us at 1 row, 2.57 us at 1,024; the 5 runs' medians never differed by more than 40.3 us at any size. Caption: 1,024 rows cost 5.24x the time of one row; per row, 195 times cheaper.

One row took 501.7 microseconds. 1,024 rows took 2,628.8 microseconds, only 5.24 times as long. So one row inside a batch of 1,024 cost 2.57 microseconds, 195 times less than one row alone.

The photocopier explains it. Most of a call's time is the walk to the machine: checking the input, setting up arrays, starting the model's code. That part costs about the same for 1 row or 100. The trees themselves are fast. The model has 54 trees, each with at most 31 leaves, and walking one row through all of them takes only a couple of microseconds.

Notice something odd at the left of the chart: 2 rows (468.9 us) and 4 rows (476.3 us) were faster per call than 1 row (501.7 us). I measured it in all five runs, but I did not find out why, so I only report it.

An isometric drawing of eleven cylinders in two rows, one per batch size, each as tall as the square root of the per-row time, with the time in microseconds printed under it: b=1 502, b=2 234, b=4 119, b=8 61, b=16 32, b=32 17; second row b=64 9.78, b=128 6.07, b=256 4.13, b=512 3.11, b=1024 2.57. Below: each cylinder's height grows with the square root of the per-row time, so the small ones stay visible; even from 512 to 1,024 rows, doubling still saved more than 10% per row. Caption: almost all of one row's cost was the fixed price of making a call.

Where does the per-row cost stop falling? I fixed the rule before the runs: the first size at which doubling the batch saves less than 10% per row. It never happened. From 512 to 1,024 rows, the per-row time still fell from 3.11 to 2.57 microseconds, 17.5% less. So for this model, bigger batches kept getting cheaper per row all the way to the largest size I tried.

Same Answers, to the Last Bit

A batch is only useful if it gives the same answers. Does a row scored inside a batch get the same score as when it is scored alone?

Three panels. Model alone: all equal; 26,851 rows, in batches of every size from 1 to 1,024 and all at once: equal to the bit. Through the service: 7,179,151 of 7,179,151 answers equal to the bit; 6,092,432 of them were scored in a batch of more than one row. Box against my laptop: 26,818 of 26,851 single-row scores equal; the rest differ by at most 1.1e-16. Caption: on the box, one row or a thousand, every score came out the same.

Yes, exactly. On the box I scored all 26,851 test rows one at a time. Then I scored them again in batches of every size from 2 to 1,024, and all 26,851 in one call. Every score was equal to the bit, which means every one of the 64 bits that store the number was the same. The tolerance was zero.

The service runs agree too. Every one of the 7,179,151 answers the service sent back, in all 595 open-loop runs, was equal to the bit to the box's own single-row score for that customer. That count includes the no-batching runs and batches of one row. Counted after the review, 6,092,432 of those answers were scored in a batch of more than one row.

The box and my laptop are a different story, and a familiar one from lesson 3. 26,818 of the 26,851 single-row scores matched my laptop's to the bit. The other 33 differed by at most 0.00000000000000011. Two different processors and builds of the same libraries can round the last bit differently. That has nothing to do with batching.

How Many Answers Can the Service Give With No Batching?

Before Part B, I needed one number: S, the most answers per second the service gives with no batching. Every arrival rate in Part B is a multiple of it.

A hand-drawn bar chart of five runs of the service with no batching, 16 connections each sending again the moment its answer arrives, 8 seconds counted: run 1 1094.8, run 2 1098.4, run 3 1100.5, run 4 1091.3, run 5 1098.5 answers a second, the bars nearly the same height. Below: S, the median of the 5 runs, 1098.4 answers a second, about 910 us of the core per answer; the client sends and reads on the same core, so part of that time is the client's; the open-loop rates were 549, 879, 1,098, 1,373, 1,757, 2,197, 2,746 a second. Caption: every rate in part B is a multiple of this one number.

The five runs gave 1,091.3 to 1,100.5 answers a second, and S, their median, was 1,098.4. This works out to about 910 microseconds of the core for each answer, of which predict took 587 to 597 microseconds inside the service (its median at half and at 1.6 times S) and the rest is , checking the request, the SQLite lookup, the answer, and the client.

So the seven rates in Part B were 549, 879, 1,098, 1,373, 1,757, 2,197 and 2,746 requests a second.

One thing S does not tell you. S came from a closed-loop client, which is cheap to run. The open-loop client in Part B uses two threads and a pool of 512 kept-alive connections, connections that stay open between requests instead of being made fresh for each one, and it spends more of the shared core on each request. With that client, no batching could only give about 1,034 to 1,045 answers a second, a little below S.

This is most likely why no batching already fell behind at exactly S. I did not test it. Another possible reason is the 512 connections themselves: the service has to watch many more open connections than the 16 of the saturation runs.

How the Batcher Works

Here is what one batch does, using a real example: B = 32, W = 2 ms, at half of S, 549 requests a second.

A hand-drawn timeline with an arrow labelled time. Three boxes left to right: 1st arrives; wait 2.09 ms; predict 564 us. Below, one small box for a request that joins. Text: about 1.4 more join, on average; then every request in the batch is answered. Below: median of 5 rounds: batches held 2.4 requests on average (the first plus about 1.4 more); a batch that did not fill waited 2.09 ms from its first request to being taken; predict on the whole batch took 564 us; box lengths are not to scale. Caption: at a low rate the batch rarely fills, so the first request pays the whole wait.

Each request first goes through the same steps as in lesson 2: read the body, check it, look up the six features in SQLite. Then, instead of calling the model, the handler puts its row in a list and waits. One batcher task, a small loop running inside the same server, watches the list.

When the first row arrives, the batcher starts a timer for W. If B rows arrive before the timer ends, it takes them at once. If not, it takes whatever is there when the timer ends. Then it calls predict_proba once for the whole batch, and hands each request its own score.

At 549 requests a second, about one request arrives every 1.8 ms. So in a 2 ms wait, only one or two more come. These batches held 2.4 requests on average, never close to 32. The first request in each batch waited about 2 ms for almost nothing.

With W = 0, the batcher does not start a timer. It takes whatever rows have gathered by the time it gets to run. At half of S those batches held only 1.2 requests on average. But, as the next slide shows, those few small batches still mattered.

At a Low Rate: What Batching Costs and What It Saves

At half of S, the service with no batching keeps up easily. So what does a batcher do there? For each setting, I subtracted the no-batching median from the setting's median latency, in the same round, on the same arrivals.

A dot chart of five rounds per setting at 0.5x S: the setting's p50 minus the no-batching p50, in ms, with a dashed line at zero. Columns are grouped by B: the queue with B=1, then B=8, B=32 and B=128, each with W = 0, 1, 2, 5 and 10. The queue sits just above zero; every W=0 column sits on the zero line; within each B group the dots climb with W, to about 1 ms at W=1, 2 ms at W=2, 4 ms at W=5 and 5.5 to 7 ms at W=10. Below: each group of five columns is one B, with W = 0, 1, 2, 5 and 10 ms; q is the queue with B=1; told apart from no batching (all 5 rounds one sign, ranges apart): 13 of 16; no batching's p50 here: 1.18 ms. Caption: the higher the W, the higher the p50: below S, a batch mostly waits.

Every timed wait raised the median. No batching had a p50 of 1.18 ms. With B = 32, W = 1 gave 2.20 ms, W = 2 gave 3.11, W = 5 gave 4.99 and W = 10 gave 7.80. All twelve settings with a W of 1 ms or more had a higher p50 than no batching, told apart in all five rounds. The cost was roughly the wait itself, because at this rate batches rarely fill.

W = 0 cost almost nothing at the median. For B = 8, 32 and 128 with W = 0, the p50 was 1.24 ms against 1.18 ms, and that difference cannot be told apart from run-to-run noise. In round 4, all three went the other way, by 6, 8 and 18 microseconds.

But W = 0 cut the p99 in half. At the same rate, no batching had a p99 of 5.26 ms, and B = 32 with W = 0 had 2.50 ms. All five rounds agreed, by 2.4 to 3.3 ms.

Why would batching help when the service is only half busy? Random arrivals come in clumps. Sometimes four requests arrive within a millisecond. With no batching, they are scored one after another, about 0.9 ms each, so the fourth waits behind three others. With W = 0, the requests that gathered while the model was busy are scored together in one call, for about the price of one. The clump clears quickly, so the slowest requests are not so slow. The median request rarely meets a clump, so the median barely moved.

I checked the clump idea after the review, so it is a labelled addition. With no batching at half of S, the slowest 1% of requests had 4.5 to 4.9 arrivals in the 3 ms before them, against 1.60 to 1.69 for all requests, in all five rounds. So the slowest requests did arrive in clumps. That fits the explanation, but it does not prove that scoring the clump together is what helped.

The Wait I Asked For, and the Wait I Got

Did a batch with W = 1 ms really wait 1 ms? I measured, for every batch that did not fill, the time from its first request joining the list to the batcher taking the rows.

Five pairs of bars for B=32 at 0.5x S: an outlined bar for the W asked and a filled bar for the median wait measured, for batches that did not fill. W = 0 ms: asked 0 ms, waited 0.016 ms. W = 1 ms: asked 1 ms, waited 1.085 ms. W = 2 ms: asked 2, waited 2.089. W = 5 ms: asked 5, waited 5.155. W = 10 ms: asked 10, waited 10.165. Below: the outlined bar is W; the filled bar is the measured median wait, from the batch's first request being queued to the batch being taken. Caption: each wait ran 0.085, 0.089, 0.155, 0.165 ms past its W; one possible reason: after the timer fires, the batcher waits its turn in the event loop.

Each wait ran a little longer than asked: 1.085 ms for W = 1, 2.089 for W = 2, 5.155 for W = 5 and 10.165 for W = 10. With W = 0 there was no timer, and the median wait was 16 microseconds.

The batcher runs inside the service's event loop, the part of the server that switches between all its waiting jobs on one thread. The batcher's timer uses loop.call_later. Python's documentation describes it as "Schedule callback to be called after the given delay number of seconds". The service runs on uvloop, a faster event loop for Python, which is built on a library called libuv, and libuv's documentation says its timers' "timeout and repeat are in milliseconds".

So I guessed, before the runs, that a 1 ms wait would come out somewhere between 1 and 2 ms, and it did. The extra was 85 to 165 microseconds, not a whole millisecond. One possible reason for the extra is that after the timer fires, the batcher still waits its turn among the other tasks in the event loop. I did not test that.

The lesson for a real service: the W you set is a floor, not an exact time.

At Higher Rates: Who Kept Up

Now the busy side. This chart follows the median of five settings as the arrival rate rises.

A line chart of p50 latency on a log scale, from 1 ms to 10 s, against the arrival rate as a multiple of S, not to scale. Five lines: no batching; queue, B=1; B=32, W=0 ms; B=32, W=2 ms; B=128, W=10 ms. The no-batching line and the queue line stay low up to 0.8x, then jump to hundreds of milliseconds at 1x and to several seconds at 2.5x. The three batching lines stay between about 1 ms and 13 ms at every rate, with the W=10 line highest of the three. Below: no batching kept up to 0.8x; queue, B=1 kept up to 0.8x; B=32, W=0 ms kept up to 2.5x; B=32, W=2 ms kept up to 2.5x; B=128, W=10 ms kept up to 2.5x; kept up means at least 95% of the arrival rate answered, in all 5 rounds. Caption: a line that jumps is a queue that never emptied: its latency grows with the length of the run.

No batching kept up to 0.8 times S and no further. At 0.8x its p50 was 2.77 ms and its p99 16.4 ms. At 1x the p50 jumped to 291 ms, and at 2.5x to 6.9 seconds. The queue with B = 1 broke at the same place, and was worse at every busy rate.

Every setting with B of 32 or 128 kept up at every rate, in all five rounds. So did B = 8, except at 2.5 times S, where it kept up in only 2 of 5 rounds. A cap of 8 rows per call was too small for that rate: even in the two rounds that kept up, its p50 was about 310 and 360 ms, because up to 5% of a run's requests may pile up and still count as keeping up.

Look at the numbers above S with care. A setting that does not keep up has a queue that grows for the whole run. If my runs had lasted 16 seconds instead of 8, its latencies would have been larger still. So for no batching above S, "721 ms" really means "the queue never emptied". It is not a stable number.

A line chart of answers per second actually given against the arrival rate as a multiple of S, with a dotted line for the arrivals themselves. The no-batching line and the queue line follow the arrivals up to 0.8x, then flatten near 1,000 to 1,045 answers a second. The B=32 W=0 and B=128 W=10 lines follow the arrivals all the way to about 2,740 at 2.5x. Below: at 2.5x S (2746 arrivals a second), no batching gave 1,045 answers a second; the best batched setting gave 2,738, 2.62 times as many. Caption: where a line leaves the dotted arrivals line, that setting has stopped keeping up.

tells the same story in answers per second. No batching topped out at about 1,034 to 1,045 answers a second. The queue with B = 1 topped out lower, at about 972 to 978. The best batched settings gave 2,738 answers a second at 2.5 times S, at least what no batching gave. Batching's own limit was not reached: I stopped at 2.5 times S, so I do not know where batching itself would have run out of room.

Every Setting at One Busy Rate

Which batched setting was best? Here is every one of them at 1.6 times S, 1,757 requests a second.

A grid of 15 cells, rows B=8, 32 and 128, columns W=0, 1, 2, 5 and 10 ms, each cell holding the p99 in ms at 1.6x S and the words kept up, with darker cells for higher p99. B=8: 6.23, 6.41, 6.02, 7.58, 8.35. B=32: 6.00, 5.51, 6.39, 10.1, 17.8. B=128: 5.85, 5.47, 6.41, 10.1, 17.6. Every cell kept up. Below: for comparison at the same rate, no batching p99 5501 ms (fell behind); the queue with B=1 6377 ms; a darker cell has a higher p99; a framed cell fell behind in at least one round. Caption: lowest p99 that kept up in all 5 rounds: B=128, W=1 ms.

At this rate, every batched setting kept up, and the p99s ran from 5.47 ms to 17.8 ms. No batching was at 5,501 ms. Next to that gap, the differences between batched settings are small.

The lowest p99 was B = 128, W = 1 ms, at 5.47 ms. But compared directly, round by round, it could not be told apart from B = 32 with W = 0 or 1, or from B = 128 with W = 0. So I cannot say it was better than those three.

The large waits were clearly worse. With W = 10 ms the p99 was about 17.6 to 17.8 ms. And at 2 and 2.5 times S, W = 10 got much worse: B = 32, W = 10 had a p99 of 50.1 ms at 2 times S and 70.4 ms at 2.5 times S, against 8.18 ms and 13.0 ms for W = 0. I did not find out what made the long waits so much worse at those rates.

When did a timed wait beat no timed wait? After seeing the results, I checked each setting with a W of 1 ms or more directly against all three W = 0 settings with a cap of 8 or more. This check was not in the design, so I label it as added later.

A timed wait had a lower p99 than all three, told apart, only twice: B = 32 with W = 1 at 1.6 times S, and B = 32 or 128 with W = 2 at 2 times S. At 2.5 times S, the setting with the lowest p99, B = 32 with W = 2 at 10.5 ms, could not be told apart from B = 32 or 128 with W = 0, or from B = 128 with W = 2.

So on this box, a timed wait sometimes helped a little at high rates, and cost a lot at low rates. W = 0 with a cap of 32 or 128 was never far from the best, at any rate.

Where the Time Inside the Service Went

The service also recorded three moments for every request: when the handler started, when the row was ready for the model, and when the score came back. For four of the runs, the repo keeps those readings.

Four hand-drawn bars, one per setting and rate, each split into three parts: before predict, predict or the wait for the batch, and building the answer. No batching at 0.5x and 1.6x: short bars. B=32 W=2 at 0.5x and 1.6x: bars about four times as long, almost all in the middle part. Below: none 0.5x: 53 / 587 / 12; B=32 W=2 0.5x: 53 / 2,578 / 12; none 1.6x: 54 / 597 / 12; B=32 W=2 1.6x: 50 / 2,667 / 11 us; these are medians of each part, so they do not add up to a request's latency, and they leave out time spent before the handler. Caption: batching changes only the middle part; with W = 2 ms it is longer, because it now includes the wait; reading and writing stay per request.

Inside the handler, reading the body, checking it and the SQLite lookup took a median of about 50 microseconds, and building the answer about 12. With no batching, the middle part, the model call itself, took about 590 microseconds. With B = 32 and W = 2, the middle part was the wait for the batch plus the batch's model call: about 2,600 microseconds.

The handler's own work is small. But the whole core spends about 910 microseconds per answer at S, so most of a request's cost is outside the handler: the web server reading and writing , and the client. Batching does nothing about that part. So the batched service's throughput, though 2.62 times higher, is not 195 times higher like the model's per-row cost.

Why a GPU or a Big Model May Answer Differently

Everything above is one small tree model on one CPU core. I did not measure any other machine, so I give no numbers for one. But the measurements here show which parts of the answer would change, and why.

A two-column ledger titled this box, measured, and elsewhere, reasoned. Left: here, one row cost 502 us per call, 1,024 rows 2,629 us; the fixed part of a call was most of it. Here, one core, so the rows of a batch are still scored one after another; batching only skipped the fixed part of each call. Here, reading, checking and the lookup stayed per request, so they set a ceiling batching cannot lift. Right: a GPU runs many rows at the same moment, and each call also pays to copy data to the card: batching usually matters more. A big model spends far longer per row, so the fixed part of a call is a smaller share: batching saves a smaller share of the time. Measure on that hardware; NVIDIA Triton's dynamic batcher has a wait, max_queue_delay_microseconds, and a cap, max_batch_size.

What batching saved here was the fixed cost of a call. On one core, the rows of a batch are still scored one after another, and each row still costs its couple of microseconds. Batching only stopped me paying the 500-microsecond walk to the photocopier for every row.

A GPU changes the per-row part. A GPU is a chip that does thousands of small calculations at the same moment. Many rows can be scored at once, not one after another, so a bigger batch can cost almost the same as a small one even when the model is large. Each call also has to send its data to the GPU's own memory and back, another fixed cost. Both of those make batching matter more on a GPU, which is why serving systems for GPUs usually batch.

A big model changes the balance. If one row took 10 milliseconds of real work, the fixed 500 microseconds would be a small share of it. Batching would save a smaller share of the time on a CPU, and the extra of a timed wait would be smaller compared with the work.

These are reasons, not measurements. If you serve on other hardware, measure Part A on it first. Batching servers built for GPUs expose the same dials I used. For example, NVIDIA's Triton Inference Server documents a dynamic batcher with a max_queue_delay_microseconds setting, like my W, and a max_batch_size setting, the largest batch, like my B. Its setting is a preference for certain sizes, not a cap.

My Eleven Guesses Before the Run, Checked

I wrote eleven guesses into the lab before it ran. Two of them had two parts, so the report checks thirteen statements. Eleven were right and two were wrong.

  1. "predict on 1 row: median per call between 450 and 700 us." Right: 501.7 us.

  2. "1,024 rows cost less than 10 x the time of 1 row per call." Right: 5.24 times.

  3. "Per-row cost still falling by 10% or more from 512 to 1,024." Right: it fell by 17.5%.

  4. "Every batched score equals the single-row score to the bit, for every b." Right.

  5. "S between 800 and 1,300 answers per second." Right: 1,098.4.

  6. "At 0.5 x S, every setting with W at least 1 ms has a higher p50 than none, told apart; every W = 0 setting's p50 cannot be told apart from none's." First part right: 12 of 12. Second part wrong: the queue with B = 1 had a p50 244 us higher, told apart. The three W = 0 settings with a real cap could not be told apart from no batching.

  7. "At 1.25 x S and above, none does not keep up; b32w2 keeps up at 1.6 x S and 2.0 x S, with a lower p99 than none, told apart." Both parts right. No batching kept up in 0 of 5 rounds at every one of those rates.

  8. "The highest of any setting at 2.5 x S is at least 1.5 x none's." Right: 2.62 times.

  9. "b1w0 has a p50 within 50 us of none's at 0.5 x S." Wrong: 244 us. The batcher's machinery costs more than I thought.

Try It Yourself

The full lab needs a rented box and about two and a half hours. The demo, bat_demo.py, runs on your own computer. It times the model on batches, checks the scores, and runs a tiny batcher with no web server.

A page in four labelled zones, headed bat_demo.py, designed before it ran. Train: the features chapter's model; it must score test AP 0.5450. Raw cost: predict_proba on 1 to 1,024 rows, 300 timed calls each, per call and per row. Same answers: every test row alone, in batches of 32, and all at once: equal to the bit? A tiny batcher: predict only, no web server; random arrivals at 0.5x and 1.5x of this machine's own single-row rate, one predict each against batches of up to 32 with a 2 ms wait. Caption: on the box it printed 509.4 us for 1 row and 2596.2 us for 1,024.

I wrote the demo's design into its docstring after the lab's design and before the demo first ran. I did not change it after the results.

A real screenshot of VS Code with bat_demo.py open at the top of the file, showing its docstring: what it needs, how to run it, and the design written before it first ran.

Before you run this lab. The demo uses lesson 2's Python environment, with the same library versions as the lab's box. If you made it for lesson 2 or 3, use it again. If not, make it once, inside the scripts/labs/serving/examples folder:

python3 -m venv venv-sv
source venv-sv/bin/activate        # on Windows: venv-sv\Scripts\activate
pip install -r lat_requirements.txt

I used Python 3.13. Then run python ../../features/fetch_data.py once, which downloads the shop data. Now run python bat_demo.py. It needs no GPU, no web server and no cloud account. If a package is missing, it prints one line saying why and stops.

The first line it prints says the times are from your machine, with its number of CPUs and its load at that moment. Read them for their shape, not their size. On a machine with many cores, predict_proba may use several of them in one call, so the per-call times can look quite different. Does the per-row cost fall as the batch grows? Are the batched scores the same? Does the tiny batcher keep up where one predict each falls behind?

Pick B, W and a Rate

This box holds the real results from the lab. It needs nothing but Python, so it runs in your browser. It has no model and no clock: it only looks up what the box measured.

Press Run. Then try B = 1 to see the queue with batches of one. Try W = 10 at RATE = 0.5, then at RATE = 2.5. Try B = 8, W = 0 at RATE = 2.5, where a cap of 8 was too small.

The report script writes this box from the lab's raw files, runs it with six settings, and checks that each run prints the lab's numbers. Every line shows no batching first, at the same rate, so you always compare with it.

The Lab's Code, Piece by Piece

The lab has one file that runs on my laptop, a few small files that run on the box, and the step script you saw above.

batching_requests.py, on my laptop, holds the design in its docstring, written before the first timed run, with dated notes added later. Its commands are check (lesson 2's files match by sha256), status (reads the box's serial console), collect, cost and factcheck. collect calls bat_stats.py, which makes a small copy of the raw files in results/bat-raw/ and does all the arithmetic on it.

bat_files/bat_predict.py is Part A. It loads the model and the 26,851 real test rows and times predict_proba on 1 to 1,024 rows. Its bits command scores every row alone and in batches and compares the scores bit by bit.

bat_files/bat_service.py is the service. It imports lesson 2's service file unchanged, so the model, the SQLite table and the request checks are lesson 2's. With no batching, the handler calls predict_proba on its one row, as in lesson 2. With batching, the handler puts its row and a in a list and waits. A future is an empty box for an answer: the handler waits on it, and someone else fills it in later. The batcher task takes up to B rows off the list, calls once, and fills in every future.

How to Decide Whether Your Service Should Batch

Here is the order I would follow, using only what this lab did.

A flowchart. Time predict on 1, 8, 32 and more rows, alone, leads to a diamond: much cheaper per row? No leads to batching cannot help; stop here. Yes leads to measure S; try W = 0 with a cap B of 8 or more against no batching, same arrivals. That leads to a diamond: traffic near or above S? No leads to keep W = 0: a timed wait only adds latency. Yes leads to raise W only if a direct comparison at that rate shows a lower p99. Below: here, at 0.5x S, W = 0 with B = 32 had p99 2.50 ms against 5.26 ms with no batching; a timed wait had a lower p99 than all three W = 0 settings with B of 8 or more, each compared directly, only at 1.6x S (B=32, W=1 ms) and 2x S (B=32, W=2 ms; B=128, W=2 ms); that check was added after the results. Caption: measure the model alone first, then compare on the same arrivals.

The chart starts with the cheapest test, the model alone, and only then brings in the service and its traffic. Here are the same steps in words, each with the number from this lab behind it.

  1. Time the model alone first. Call it on 1, 8, 32 and more rows. If one row in a batch is not much cheaper than one row alone, no batcher can help, and you can stop. Here it was 195 times cheaper at 1,024 rows.

  2. Measure S, with no batching. It tells you where the service will fall behind. Here no batching fell behind at S and gave at most about 1,045 answers a second.

  3. Test with random arrivals, and count from arrival. A test that waits for each answer before sending the next can never show a queue. Lesson 7 of this chapter measures what that kind of test hides.

  4. Start with W = 0 and a cap. Here it never did worse than no batching on p99, and from 0.8 times S up it was better on p50 too.

  5. Compare settings on the same arrivals, directly. Two settings that each differ from no batching are not ranked by that. Compare them with each other.

When Batching Helps, and When It Does Not

Batching helps when the fixed cost of a call is large compared with the cost of a row. Here the fixed part was about 500 microseconds and a row about 2.6, so batching saved almost everything the model spent.

It helps most when traffic comes near or above S. Here, no batching fell behind at S. Batches of up to 32 or 128 kept up at 2.5 times S.

It can help a little even when the service is quiet, if it does not wait on purpose. At half of S, W = 0 halved the p99 by scoring the clumps of random arrivals together.

A timed wait is a cost at low traffic. Every W of 1 ms or more raised the median at half of S, by about the length of the wait. If your traffic is quiet most of the day and busy for an hour, a fixed W makes every quiet request slower to help the busy hour.

A batcher that never groups requests is pure cost. The queue with B = 1 had a higher p50 than no batching at every rate, and once both fell behind it gave fewer answers a second.

Do not batch blindly into a model you have not checked. Batching changed no score here, to the bit. Check yours the same way before you trust it.

Do not copy my handler as it is into a busy service. Like lesson 2's, it is an async def that calls SQLite and predict_proba, which block the event loop while they run. On one core with one worker, that is what I measured on purpose. With more cores, lesson 5 measures workers and threads.

What This Lab Cannot Tell You

Two columns titled shows and cannot show. Shows: one small tree model on one 1-vCPU box; client and service share the core. Random arrivals at 7 steady rates, 8 seconds each, 5 rounds per setting. Batching of predict only; reading and the lookup stay per request. Cannot show: a GPU, a big model, or a box with many cores (lesson 5 measures workers and threads). Traffic that rises and falls within seconds, or stays above S for longer than 8 seconds. A service that also batches the feature lookup, or a different web server.

One small model, one core, one shared client. The client took part of the core in every run. A service with its client on another machine would answer more per second, with or without batching.

Steady rates for 8 seconds. Real traffic rises and falls. Above S, my latencies for no batching grow with the run's length, so they are not stable numbers.

The schedule was started three times on the same box. The first start ran Part A, the bit check and the saturation runs, then its first open-loop run hung. uvicorn closes a kept-alive connection after 5 seconds with no request, and my client read the closed connections in a loop instead of failing. I stopped it, made the client fail loudly, started uvicorn with --timeout-keep-alive 3600, and copied the files again.

That copy step stopped early, because uv venv refuses a folder that already has a venv, and I only noticed after I had started the schedule again with the old files. I stopped that second start within its first minute, before any run, fixed the step with --allow-existing, and started the whole schedule from the beginning.

Every number in this lesson comes from that third start.

Before the first start, a short functional check ran on the box with me connected: Part A with 20 calls per size, a 1-second saturation run, and two 1-second open-loop runs at 300 a second. Before the third start, one 8-second run of B = 128 with W = 0 at 547 a second, the run that had hung, checked the fix. I printed only that they finished, plus the saturation run's count of 2,201 answers, and deleted their files. From the first start I saw its saturation numbers (S = 1,094.6) before I threw them away. The watch recording in step 8 is from the first start.

What to Do on Monday

A hand-drawn grid of six cards, titled five habits for batching. 1, time the model alone: per call and per row, for batch sizes 1 to 1,024. 2, find S first: the most answers a second with no batching. 3, use open-loop arrivals: and count latency from arrival, so waiting counts. 4, compare on the same arrivals: each setting against no batching, and settings against each other directly. 5, check the scores: batched against single, to the bit. The reason: here, W = 0 batching halved p99 even at half of S; every W of 1 ms or more raised p50; no batching fell behind at S. Caption: batching is a tool for the busy hour, not a free speed-up.

If you take one thing to work on Monday, time your model on 1, 8, 32 and 1,024 rows. It takes a few minutes, and it tells you whether batching can help at all.

If it can, find the most answers a second your service gives with no batching, then try a batcher with no timed wait and a cap of 32, on the same random arrivals. On this box that one change kept the service up at 2.5 times its old limit and halved its p99 even when it was half busy.

The one idea to keep: batching pays for the fixed cost of a call once instead of many times. Waiting on purpose to fill a batch is a separate choice, and it costs every quiet request. Measure both on your own model and hardware.

Knowledge Check

Knowledge Check

4 questions - Score 80% to pass

Q1

In Part A, how long did predict_proba take on 1,024 rows compared with one row?

Q2

At half of S, what did batching with W = 0 and B = 32 do compared with no batching?

Q3

With the open-loop client, how many answers a second did no batching give at most?

Q4

The lowest p99 at 1.6 times S was B = 128 with W = 1 ms. What does the lesson conclude?

Within one round, every setting at one rate gets exactly the same arrival times and the same customers.

The rule, written before the runs, is lesson 3's rule. A setting is "told apart from run-to-run noise" only if two things hold. All five paired differences have the same sign. AND its five values and the other setting's five values do not overlap. Otherwise it "cannot be told apart". When I say one setting beat another, I compared those two directly. I never conclude it from "this one differed from no batching and that one did not".

cd scripts/labs/serving
export AWS_DEFAULT_REGION=us-east-1 AWS_PAGER=""
NAME=ai-research-course-bat
STATE=~/lab-data/serving/bat
KEY=~/lab-data/serving/keys/$NAME.pem
SSH=(-i "$KEY" -o LogLevel=ERROR -o StrictHostKeyChecking=accept-new -o "UserKnownHostsFile=$STATE/known_hosts")
K=${K:-5}; CALLS=${CALLS:-2000}; SAT_S=${SAT_S:-8}; OPEN_S=${OPEN_S:-8}
mkdir -p "$STATE"
touch "$STATE/vars"; . "$STATE/vars"

What you need first. An AWS account, the AWS command line tool (aws) set up with your keys for the region us-east-1, and ssh, scp, rsync and curl. The region needs a default VPC, the ready-made private network that new accounts get; step 2 puts the box's door rule in it. Step 5 copies two things that must already be on your computer: lesson 2's prepared files in ~/lab-data/serving/lat/, made by python latency_anatomy.py prepare in this folder, and the shop data in ~/lab-data/features/retail.parquet, made by python ../features/fetch_data.py.

What it costs. A c7g.medium costs $0.0363 an hour on demand in us-east-1, from AWS's price list, and AWS bills it by the second. My box was on for 2.735 hours, from launch to terminated, which is $0.0993 for compute. Its 16 GiB disk was deleted with it, and adds a few cents a month pro rata for the hours it existed. These numbers are in results/bat-cost.json. The schedule itself took about 2 hours and 13 minutes.

The file changed a little after the recordings. After a review I made the file easier to copy from, without renting a box again. Step 3 now looks up today's Ubuntu image instead of using a fixed id; the recorded run used ami-0bec8cef5313300ad. The ids and the address are kept in shell variables instead of small files. Paths on the box are written from its home folder. Step 6 prints two lines of the timer list instead of one. Step 13 also deletes the small file of variables. The recordings show the earlier version, and the file's header lists the same changes.

What the recordings hide. Every line passed through a small filter, tail_redact.py, which I wrote for lesson 3 and use again here. It hides my account number and every IP address. It also hides every name AWS makes up for a resource, such as the box's id, my user name, and anything from a key file. Where you see <ip> or i-<id>, your terminal shows the real value. Lines that start with + are the shell printing each command just before it runs it.

I then checked every recording again with a separate scanner, which looks for my account number, my address and the real ids, and with my own eyes.

~/lab-data/serving/keys/

Step 3: rent the box. run-instances asks for one c7g.medium with Ubuntu 24.04 and a 16 GiB disk that is deleted with the box. It also sets two tags, labels that say which project and lesson the box belongs to. The tags are how step 1 and step 13 find it again.

A real terminal recording of step 3, with the group and instance ids replaced by placeholders. The command aws ec2 run-instances with the Ubuntu image ami-0bec8cef5313300ad, type c7g.medium, the key ai-research-course-bat, a 16 GiB gp3 disk deleted on termination, and the tags Project, Lesson and Name. It prints the new instance's id, c7g.medium and pending.

AMI=$(aws ssm get-parameters \
  --names /aws/service/canonical/ubuntu/server/24.04/stable/current/arm64/hvm/ebs-gp3/ami-id \
  --query "Parameters[0].Value" --output text)
ID=$(aws ec2 run-instances --image-id "$AMI" --instance-type c7g.medium --count 1 \
  --key-name "$NAME" --security-group-ids "$SG" \
  --block-device-mappings '[{"DeviceName":"/dev/sda1","Ebs":{"VolumeSize":16,"VolumeType":"gp3","DeleteOnTermination":true}}]' \
  --tag-specifications \
  "ResourceType=instance,Tags=[{Key=Project,Value=ai-research-course},{Key=Lesson,Value=bat},{Key=Name,Value=$NAME}]" \
  "ResourceType=volume,Tags=[{Key=Project,Value=ai-research-course},{Key=Lesson,Value=bat}]" \
  --query "Instances[0].InstanceId" --output text)
echo "ID=$ID" >> "$STATE/vars"

The first command asks AWS for the id of today's Ubuntu 24.04 image for arm processors, from a public list that Canonical keeps up to date, so it works in any region. On the day I recorded, it was ami-0bec8cef5313300ad in us-east-1, the id you see in the recording.

The full list of files to copy is in the copy step of bat_aws_steps.sh. The two folders named raw-attempt1 and raw-attempt2 at the end of the list hold the files of two earlier starts of the schedule, which I explain on the limits slide. None of their numbers is used.

Step 6: stop the timers and look at the box. Ubuntu runs small jobs on a timer, such as checking for updates. One of those starting in the middle of a run would land straight in the results, so the script stops every timer first. Then it prints what the box is (lscpu), how many cores it has (nproc), how busy it has been (uptime), and every installed package with its version (pip freeze).

A real terminal recording of step 6, with the address replaced by a placeholder. Over ssh: stop every active timer and the unattended-upgrades service, after which systemctl list-timers prints only its last line, the hint Pass --all to see loaded but inactive timers, too; lscpu prints architecture aarch64, 1 CPU, vendor ARM, model name Neoverse-V1, 1 thread per core, 1 core per socket; nproc prints 1; uptime prints up 26 min with load average 0.06, 0.38, 0.49. Then the box info is saved, and Python 3.13.15 and the full pip freeze are printed, including fastapi 0.142.2, numpy 2.5.3, pandas 3.0.6, scikit-learn 1.9.1, starlette 1.7.0, uvicorn 0.54.0 and uvloop 0.23.0.

ssh "${SSH[@]}" ubuntu@"$IP" 'for t in $(systemctl list-units --type=timer --state=active --no-legend --plain \
  | cut -d" " -f1); do sudo systemctl stop "$t"; done; sudo systemctl stop unattended-upgrades.service; \
  systemctl list-timers --no-pager | tail -2'
ssh "${SSH[@]}" ubuntu@"$IP" 'lscpu | head -12; nproc; uptime'
ssh "${SSH[@]}" ubuntu@"$IP" 'cd bat && bash bat_box_info.sh before > raw/box-info-before.txt 2>&1; \
  venv/bin/python --version; ~/.local/bin/uv pip freeze --python venv/bin/python'

Check three things: nproc prints 1, the load is falling towards 0, and the versions match lat_requirements.txt. The 5- and 15-minute load averages here, 0.38 and 0.49, are still high from my earlier attempts on the same box. The schedule waits until the 1-minute average is at most 0.10 before its first run.

Step 7: start the schedule, then leave. The schedule starts in the background with setsid nohup, so it keeps running after SSH disconnects. Then I log out, and nobody logs in again until it ends.

A real terminal recording of step 7, with the address replaced by a placeholder. Over ssh: cd bat, then setsid nohup bash bat_schedule.sh 5 2000 8 8 with its output sent to raw/schedule.log, in the background; two seconds later the log shows schedule start K=5 calls=2000 sat=8 open=8, and load before the schedule 0.05 0.37 0.49 after waiting 0s.

ssh "${SSH[@]}" ubuntu@"$IP" "cd bat && (setsid nohup bash bat_schedule.sh $K $CALLS $SAT_S $OPEN_S \
  > raw/schedule.log 2>&1 < /dev/null &); sleep 2; cat raw/schedule.log"

K, CALLS, SAT_S and OPEN_S, set in the variables block, are the rounds, the timed calls per batch size, and the seconds of each saturation run and each open-loop run. The lab used 5, 2,000, 8 and 8.

ssh "${SSH[@]}" ubuntu@"$IP" 'cd bat && bash bat_box_info.sh after > raw/box-info-after.txt 2>&1; du -sh raw'
rsync -a -e "ssh ${SSH[*]}" ubuntu@"$IP":bat/raw/ "$STATE/raw/"
ls "$STATE/raw/" | wc -l
du -sh "$STATE/raw/"

Step 13: delete the box, the key and the door rule, and check. A box you forget keeps costing money. The script deletes the box, waits until AWS says "terminated", then deletes the key pair and the security group. Then it asks AWS again: the box must say terminated, and the counts of key pairs, security groups, disks with this lesson's tag, and course boxes still alive must all be 0.

A real terminal recording of step 13, with the instance and group ids replaced by placeholders. terminate-instances prints shutting-down; wait instance-terminated; delete-key-pair prints true; delete-security-group prints true. Then the checks: describe-instances prints terminated; the count of key pairs named ai-research-course-bat is 0; the count of security groups of that name is 0; the count of volumes tagged Lesson=bat is 0; the count of course instances not terminated is 0. Last, the key file is removed.

aws ec2 terminate-instances --instance-ids "$ID" --query "TerminatingInstances[0].CurrentState.Name" --output text
aws ec2 wait instance-terminated --instance-ids "$ID"
aws ec2 delete-key-pair --key-name "$NAME" --query Return
until aws ec2 delete-security-group --group-id "$SG" --query Return 2>/dev/null; do sleep 10; done
date -u +%Y-%m-%dT%H:%M:%S+00:00 > "$STATE/terminated"
aws ec2 describe-instances --instance-ids "$ID" --query "Reservations[0].Instances[0].State.Name" --output text
aws ec2 describe-key-pairs --filters "Name=key-name,Values=$NAME" --query "length(KeyPairs)"
aws ec2 describe-security-groups --filters "Name=group-name,Values=$NAME" --query "length(SecurityGroups)"
aws ec2 describe-volumes --filters "Name=tag:Lesson,Values=bat" --query "length(Volumes)"
aws ec2 describe-instances --filters "Name=tag:Project,Values=ai-research-course" \
  "Name=instance-state-name,Values=pending,running,stopping,stopped,shutting-down" \
  --query "length(Reservations[].Instances[])"
rm -f "$KEY" "$STATE/vars"

The security group can only be deleted once the box is fully gone, so the script retries that command every ten seconds until AWS accepts it.

The five runs agreed closely. The run medians for 1 row ranged from 499.9 to 507.0 microseconds, and for 1,024 rows from 2,594.9 to 2,635.2. The biggest spread at any size was 40.3 microseconds.

The queue with B = 1 was worse than no batching. Its p50 was 244 microseconds higher, told apart. Sending every request through the list and the batcher task, without ever grouping two, only added steps. This is the cost of the batcher's machinery on its own, and it is why the capped batch sizes above it need more than one row in a batch to win.

2.62 times
preferred_batch_size

"For W = 1 ms, a batch that does not fill waits longer than 1 ms: median between 1.0 and 2.0 ms at 0.5 x S." Right: 1.085 to 1.088 ms for the three caps.

  • "Every returned score equals the box's single-row score to the bit." Right: 7,179,151 of 7,179,151.

  • What I did not guess at all is the most useful result: that batching with no timed wait would halve the p99 at half of S. I had expected batching to matter only when the service was busy.

    r"""Does batching help a small CPU model? Time predict_proba on 1 to 1,024 rows, then try a tiny micro-batcher.
    
    Lesson 4 of 'Serving and Inference Basics'. It uses lesson 2's Python environment, made once inside this folder
    (Python 3.13 is what the lab used):
        python3 -m venv venv-sv
        source venv-sv/bin/activate              # Windows: venv-sv\Scripts\activate
        pip install -r lat_requirements.txt      # numpy, pandas, pyarrow, scikit-learn, ...
    It also needs the shop data from the features chapter: run python ../../features/fetch_data.py once first. Then:
        python bat_demo.py                  # print the three parts below
        python bat_demo.py --save out.json  # and save the numbers
    If a package is missing, it prints one line saying why and stops.
    The timings it prints come from YOUR machine, as it is right now: other programs move them, and a machine with
    many cores lets one predict_proba call use several of them. Read them for their shape, not their size.
    It writes nothing except out.json if you ask.
    
    Design, written 2026-10-04 after the lab's design (batching_requests.py) and before this file first ran:
      1. Train the features chapter's model (it must score test AP 0.5450), as lesson 3's tail_demo.py does.
      2. Raw cost: for b = 1, 2, 4, ..., 1024, 20 calls not timed, then 300 timed calls of predict_proba on b
         consecutive test rows from a random start. Print the median time per call and per row.
      3. Same answers? Score every test row one at a time, then in batches of 32, then all in one call; count the rows
         whose score is equal to the bit.
      4. A tiny micro-batcher, in this one process, with NO web server (so it shows the predict part only; the lab's
         service also reads HTTP and SQLite for every request). First the most single-row calls per second this machine
         makes back to back (S). Then requests arrive at random (Poisson) for 3 seconds at 0.5 x S and at 1.5 x S, and
         are answered two ways: one predict per request, and a batcher that waits up to 2 ms or until 32 are queued.
         Latency is counted from each request's arrival time, so waiting counts. Print p50, p99 and answers per second.
    
    Author: Roni Das
    Created: 2026-10-04
    """
    import asyncio
    import importlib.util
    import json
    import os
    import sys
    from pathlib import Path
    from time import perf_counter, perf_counter_ns
    
    NEEDED = ("numpy", "pandas", "pyarrow", "sklearn")
    missing = [m for m in NEEDED if importlib.util.find_spec(m) is None]
    if missing:
        sys.exit(f"bat_demo.py needs {', '.join(missing)}: make the environment in its docstring "
                 f"(pip install -r lat_requirements.txt), then run it with that environment's python.")
    
    import numpy as np  # noqa: E402
    
    SIZES = [2 ** k for k in range(11)]
    CALLS = 300
    
    
    def build():
        """The features chapter's model and its 26,851 test rows (lesson 3's tail_demo.py build step, shortened)."""
        from sklearn.metrics import average_precision_score
        here = Path(__file__).resolve().parent
        sys.path.insert(0, str(here.parents[1] / "features"))
        import task
        from what_a_feature_is import HAND_COLS, hgb, joined
    
        ev = task.load_events()
        lab_tr, _, lab_te = task.splits(ev)
        tr = joined(ev, lab_tr, task.TRAIN_CUTOFFS)
        te = joined(ev, lab_te, task.TEST_CUTOFFS)
        model = hgb(0).fit(tr[HAND_COLS].to_numpy(float), tr["label"].to_numpy())
        X = np.ascontiguousarray(te[HAND_COLS].to_numpy(np.float64))
        p = model.predict_proba(X)[:, 1]
        y, cut = te["label"].to_numpy(), te["cutoff"].to_numpy()
        ap = float(np.mean([average_precision_score(y[cut == c], p[cut == c]) for c in np.unique(cut)]))
        assert round(ap, 4) == 0.5450, ap
        return model, X[np.random.default_rng(0).permutation(len(X))], ap
    
    
    def raw_cost(model, X) -> dict:
        rng = np.random.default_rng(1)
        X2 = np.ascontiguousarray(np.concatenate([X, X[:1024]]))
        out = {}
        for b in SIZES:
            starts = rng.integers(0, len(X), size=CALLS + 20)
            for s in starts[:20]:
                model.predict_proba(X2[s:s + b])
            t = []
            for s in starts[20:]:
                v = X2[s:s + b]
                a = perf_counter_ns()
                model.predict_proba(v)
                t.append(perf_counter_ns() - a)
            out[b] = float(np.median(t)) / 1000
        return out
    
    
    def same_answers(model, X) -> dict:
        one = np.array([model.predict_proba(X[i:i + 1])[0, 1] for i in range(len(X))])
        b32 = np.concatenate([model.predict_proba(X[i:i + 32])[:, 1] for i in range(0, len(X), 32)])
        full = model.predict_proba(X)[:, 1]
        return {"rows": len(X), "batch32_equal_bits": int((b32 == one).sum()), "all_equal_bits": int((full == one).sum())}
    
    
    async def serve(model, X, rate: float, seconds: float, batch: bool, seed: int) -> dict:
        """Poisson arrivals at `rate` for `seconds`. Each request is one row. Latency from the arrival time."""
        loop = asyncio.get_running_loop()
        rng = np.random.default_rng(seed)
        at = np.cumsum(rng.exponential(1.0 / rate, size=int(rate * seconds * 1.5) + 50))
        at = at[at < seconds]
        lat = np.zeros(len(at))
        queue: list = []
        have, full = asyncio.Event(), asyncio.Event()
        t0 = perf_counter() + 0.05
    
        async def batcher():
            while True:
                if not queue:
                    have.clear()
                    await have.wait()
                if len(queue) < 32:
                    full.clear()
                    left = queue[0][2] + 0.002 - perf_counter()
                    if left > 0:
                        h = loop.call_later(left, full.set)
                        await full.wait()
                        h.cancel()
                take = queue[:32]
                del queue[:32]
                ps = model.predict_proba(np.array([r for r, _, _ in take]))[:, 1]
                for (_, f, _), p in zip(take, ps):
                    f.set_result(p)
    
        async def one(i: int):
            row = X[i % len(X)]
            if batch:
                f = loop.create_future()
                queue.append((row, f, perf_counter()))
                have.set()
                if len(queue) >= 32:
                    full.set()
                await f
            else:
                model.predict_proba(row[None, :])
            lat[i] = perf_counter() - (t0 + at[i])
    
        task = loop.create_task(batcher()) if batch else None
        jobs = []
        for i, a in enumerate(at):
            wait = t0 + a - perf_counter()
            if wait > 0:
                await asyncio.sleep(wait)
            jobs.append(loop.create_task(one(i)))
        await asyncio.gather(*jobs)
        end = perf_counter()
        if task:
            task.cancel()
        us = lat * 1e6
        return {"p50_us": float(np.percentile(us, 50)), "p99_us": float(np.percentile(us, 99)),
                "answers_per_s": len(at) / (end - t0)}
    
    
    def main() -> None:
        try:
            load = f"{os.getloadavg()[0]:.2f}"
        except (AttributeError, OSError):
            load = "not available on this system"
        print(f"Every time below was measured on THIS machine ({os.cpu_count()} CPUs, 1-minute load {load}).")
        model, X, ap = build()
        print(f"model trained: test AP {ap:.4f}; {len(X):,} test rows")
    
        cost = raw_cost(model, X)
        print("\n1. predict_proba on b rows, median of 300 calls, microseconds (us)")
        print("       b    per call     per row   per row vs b=1")
        for b in SIZES:
            print(f"  {b:6d} {cost[b]:11.1f} {cost[b] / b:11.2f}   {cost[b] / b / cost[1]:8.3f}")
    
        eq = same_answers(model, X)
        print(f"\n2. same answers? of {eq['rows']:,} rows scored one at a time, equal to the bit:")
        print(f"  in batches of 32: {eq['batch32_equal_bits']:,}    all in one call: {eq['all_equal_bits']:,}")
    
        s = 1e6 / cost[1]
        print(f"\n3. a tiny micro-batcher, predict only (no web server); S = {s:,.0f} single-row calls per second")
        print("   rate        answered by       p50 us     p99 us   answers/s")
        runs = {}
        for k, mult in enumerate((0.5, 1.5)):
            for batch in (False, True):
                r = asyncio.run(serve(model, X, mult * s, 3.0, batch, seed=k))
                name = "batch up to 32, 2 ms" if batch else "one predict each"
                runs[f"{mult}x {'batch' if batch else 'single'}"] = r
                print(f"  {mult:3.1f} x S   {name:21s} {r['p50_us']:9.0f} {r['p99_us']:10.0f} {r['answers_per_s']:10.0f}")
        if "--save" in sys.argv:
            out = {"load": load, "cpus": os.cpu_count(), "test_ap": ap, "median_call_us": cost, "same": eq,
                   "S": s, "runs": runs}
            Path(sys.argv[sys.argv.index("--save") + 1]).write_text(json.dumps(out, indent=1))
    
    
    if __name__ == "__main__":
        main()
    

    You saw the demo running on the lab's box in step 11, after the schedule. That exact run is stored in results/bat-demo-run.txt.

    The demo told the same story on the box, with no web server in the way. One row took 509.4 microseconds a call and 1,024 rows 2,596.2, and every batched score was equal to the bit. The tiny batcher's S was 1,963 single-row calls a second, higher than the service's 1,098.4 because there was no . At 1.5 times that, one predict each fell behind (p50 764,293 us), while batches of up to 32 with a 2 ms wait answered with a p50 of 3,119 us. At half of S, the 2 ms wait raised the p50 from 1,327 to 3,094 us.

    One demo run is one data point, so I read no more into it than that.

    And here is the same demo in VS Code on my laptop, run with python bat_demo.py in the examples folder with venv-sv active.

    A real screenshot of VS Code's terminal on my laptop, in the venv-sv environment, after running python bat_demo.py. Its first line says every time was measured on this machine, 10 CPUs, 1-minute load 2.65. Part 1, the per-call time for b = 1, 2, 4, 8, 16, 32, 64, 128, 256, 512 and 1024 rows: 729.9, 1415.9, 695.0, 682.0, 720.2, 707.0, 709.4, 752.9, 779.1, 867.8 and 1012.7 us. Part 2: 26,851 of 26,851 rows equal to the bit, both in batches of 32 and in one call. Part 3: S = 1,370 single-row calls a second. At 0.5 x S, one predict each: p50 2915 us, p99 137155 us, 679 answers a second; batches of up to 32 with a 2 ms wait: p50 3604 us, p99 14424 us, 678 a second. At 1.5 x S, one predict each: p50 1609453 us, p99 3087863 us, 1008 a second; the batcher: p50 2750 us, p99 7541 us, 2049 a second.

    I ran this on my laptop while it was busy with other work. Its first line says so: 10 CPUs, and a 1-minute load of 2.65. So read only the shape, as you would with your own run. The per-row time still fell as the batch grew, from 729.9 us at one row to under 1 us at 1,024.

    One number looks odd: 2 rows took 1,415.9 us a call, slower than 4 rows. On a loaded machine with many cores that is noise, not a finding, and I read nothing into it. The same goes for the 137 ms p99 of one predict each at half of S.

    Every batched score matched to the bit. And the tiny batcher kept up at 1.5 times this machine's own S, giving 2,049 answers a second, while one predict each fell behind at 1,008. These times come from a different, busy, many-core machine, so they are not comparable with any other number in this lesson.

    future
    predict_proba

    bat_files/bat_client.py is the client. In its open-loop mode it draws all the arrival times first. Then a sender thread sleeps until each arrival and sends, while a receiver thread reads the answers the moment they arrive. It keeps 512 connections open, because uvicorn answers one request at a time on each connection. Its source has the comment "Callback for pipelined requests to be started." next to the code that holds back a second request on a connection until the first is answered.

    bat_files/bat_run_once.sh waits until the core is idle, records the load and the CPU counters, starts a fresh service with the setting, runs the client, and stops the service. bat_files/bat_schedule.sh runs everything in order: Part A, the bit check, the saturation runs, then every setting at every rate, five rounds, each round in its own shuffled order.

    bat_report.py does not import any of the above. It reads the small copy of the raw files with its own code and recomputes every number. It also checks the guesses, the demo's stored run, the fact-check and the teardown, and it writes the playground.

  • Check the scores. Batched against single, to the bit. Here they matched, but a model that pads or groups rows differently might not.

  • One mistake during the runs. About 16 minutes into the third start, I opened one short SSH session to the box by mistake. Six runs were going on in the minute around it, all in round 1. Three of them sat outside the range of their own four other rounds, which by chance alone would happen to about two in five.

    The biggest was B = 8 with W = 1 ms at 2.5 times S (open-r1-b8w1-m6): its p99 was 243.6 ms, against 94.7 to 122.3 ms in the other four rounds. I kept all six. Dropping round 1 at their rates changes none of the comparisons with no batching.

    After the review I also checked every comparison between two settings at those rates: 8 of 1,085 change with round 1 dropped. One of them is a pair I quote: at 1.6 times S, B = 128 with W = 1 against B = 32 with W = 0. With only four rounds the rule is easier to pass, so I keep the five-round verdicts. These checks are in results/bat-result.json, the last one marked as added after the review.

    Labelled additions. The check of timed waits against all three W = 0 settings, on the grid slide, was added after the results. The journal sources, the count of answers scored in real batches, the clump check and the pairwise check of round 1 were added after the review, as was the redaction of the box's logs before they are stored in the repo. So were the shortened report lines for the recording, and the step in bat_stats.py that keeps the service's own split only for the four runs drawn on the "where the time went" slide, so the repo stays small. Neither changes a number.