Serving And Inference

Cost per Prediction: What Do 1,000 Answers Really Cost?

0 of 32 complete

0%

Contents

Back|Serving And InferenceCost per Prediction: What Do 1,000 Answers Really Cost?
1/32
82 min left
  1. Home
  2. AI Engineering: Data, RAG and Agents
  3. Serving and Inference Basics
  4. Cost per Prediction: What Do 1,000 Answers Really Cost?
Prerequisites
One Box, Many Models: How Many Fit, and What Happens When They Don't?requiredBatching Requests: Does It Help a Small CPU Model?requiredWorkers and Threads: Processes, Threads, or Both on One Core?requiredLoad Testing Honestly: Is Your Load Test Lying?requiredBatch or Online Scoring: A Month-Old Score Cost Nothing I Could Measure Hererequired
Related Topics
Model Signatures: The Right Numbers in the Wrong Shape, and What a Schema Check CatchesPackaging, Registry and VersioningFeature Cost and Selection: Two Fifths of the Lines, No Measurable LossFeatures and Feature Stores
1 of 32
Previous lesson
One Box, Many Models: How Many Fit, and What Happens When They Don't?
Next lessonA Serving Checklist: The Whole Chapter as One Service

System Design

  • Foundation
  • Intermediate
  • Advanced
  • Capstone

AI Engineering

  • Foundation
  • Data, RAG and Agents
  • Evaluation, LLM Ops and Security

systemdesign.academy

  • Home
  • Glossary
  • Interview prep
  • Reviews
  • About
  • Privacy
  • Terms

The Shopkeeper's Scale

Let me start in a small corner shop.

The shopkeeper pays rent for her shop every month. She pays it whether ten people come in or a hundred. At the end of the month she wants to know one thing: how much did each sale really cost her?

A flat illustration of a small corner shop with shelves of parcels and a window onto a sunny street. A shopkeeper in a navy apron stands behind a wooden counter. With one hand she drops a coin onto one pan of an old brass balance; on the other pan sits a small stack of wrapped parcels tied with string.

The rent alone cannot tell her. She also needs to know how many sales the rent paid for. In a busy month each sale is cheap. In a quiet month each sale is dear, even though the rent did not change. Coins go on one side of her scale, and the work they paid for goes on the other.

A machine learning service has this problem too. You rent a computer by the hour, and it answers questions with a model. In this lesson I measure how many answers one hour of a rented computer can give, on five different computers, and work out what 1,000 answers really cost.

Where This Lesson Starts

This is lesson 11 of the chapter on serving a model. Serving a model means running it as a small program that answers questions while people wait.

The model is the features chapter's model. It looks at six numbers about one customer of a real online shop, such as the days since their last order. It gives a score between 0 and 1: how likely they are to buy again in the next 30 days.

Lessons 2 to 10 measured this service on one kind of rented computer. They asked how fast it is, how it behaves when busy, and how much memory it uses. None of them asked about money.

So here is my question: what do 1,000 predictions cost? A prediction is one answer from the model, one score for one customer. To answer, I need two numbers for each computer. The first is its price per hour, which Amazon publishes. The second is how many answers it gives in that hour while still answering quickly. Nobody publishes that one, so I measured it.

The Words You Need First

Please read this slide slowly if any word is new. Each word comes with an everyday picture.

A hand-drawn street. A yellow taxi, seen from the side, with a rounded cab, two windows, two wheels and a white roof sign reading TAXI; a small dark meter in its windscreen reads $0.0363/h. Four coloured speech bubbles with question marks walk towards it along the road, and a stopwatch hangs in the sky. Handwritten red words with thin lines point at the parts: box = the taxi, a rented computer, one core; price per hour = the meter, it runs, busy or not; request = one rider's question, rate = riders a second; and, by the stopwatch, the slowest 1 in 100 within 25 ms.

A box is my short name for a rented computer. Amazon Web Services (AWS) rents them through a service called EC2. The CPU is the chip that does a computer's work. Inside it, a core is one worker that runs one step of a program at a time. AWS rents you a share of a chip and calls one core's worth a vCPU. Every box in this lesson has one.

On-demand means renting with no promise: start the box when you like, stop it when you like, and pay a fixed price per hour. Think of a taxi meter. It runs while the taxi is yours, whether or not anyone rides.

A request is one question sent to the service, such as "score customer 12346". The rate is how many requests arrive each second.

Waiting times are in milliseconds (ms). A millisecond is a thousandth of a second, so 25 ms is a fortieth of a second. A blink of an eye takes about 100 to 400 ms, so 25 ms is a small part of one blink.

Users notice the slowest answers most, so I watch those. Line up every request's waiting time, shortest to longest. The p99 is the time that 99 out of every 100 requests beat. I will mostly say "the slowest 1 in 100", which means the same thing. My target in this lesson: the slowest 1 in 100 must answer within 25 ms.

A hand-drawn picture in two rows, made with the Excalidraw library's own drawings. Top row, labelled nightly batch job: a grey database drum marked every customer's numbers, an arrow to a yellow page marked one program, one big call, and an arrow to a stack of papers marked a file of scores, read later. Bottom row, labelled online service: an envelope marked a request, an arrow to a blue cloud marked always on, scoring now, and an arrow to a notepad marked the answer.

There are two ways to hand out scores. A nightly batch job scores every customer once a night and keeps the scores in a file. Online serving runs a service all the time and scores each request the moment it comes. Lesson 1 compared the two on how old the scores get. This lesson compares their price.

One more word. A call to the model is one time the program asks the model for scores. Batching means scoring several waiting requests in one call. Lesson 4 found that one call for 32 rows costs little more than one call for one row. I use its best setting: up to 32 requests in one call, and the service never waits on purpose for more to arrive. I call the other way "one by one".

The Five Machines

My AWS account may run only one core at a time, so I could rent only small one-core boxes, one after another. I picked five. Their names look like codes, so here is how to read them. The first letter says the family: "c" is AWS's family for computing work. The number is the generation. A "g" after it means AWS's own Graviton chip, and an "a" means an AMD chip. ".medium" is the smallest size, with one core. The a1.medium is an older family with its own name.

Five chip dies seen from above, each a square with pins along its four edges and a logo in the middle. Four stand on a line marked AWS Graviton, Arm designs, oldest to newest: a1, Graviton 1st, Cortex-A72, $0.0255/h; c6g, Graviton 2nd, Neoverse-N1, $0.0340/h; c7g, Graviton 3rd, Neoverse-V1, $0.0363/h; c8g, Graviton 4th, Neoverse-V2, $0.0399/h. Each carries the Arm logo. Behind a dashed line, marked x86, a red die with the AMD logo: c7a, AMD EPYC 9R14, $0.0513/h.

AWS makes its own chips, called Graviton. They use designs from a company called Arm, whose designs also run most phones. My four Graviton boxes cover four generations. The a1.medium has the first Graviton, the c6g.medium has Graviton2, the c7g.medium has Graviton3, and the c8g.medium has Graviton4. The c7g.medium is the box of lessons 2 to 10.

The fifth box, the c7a.medium, has an AMD EPYC chip. It is of the x86 kind, used in most Windows laptops and many servers. The two kinds understand different machine instructions, so each needs its own copy of Python. A package is a ready-made box of code, such as numpy. The command pip freeze prints every installed package and its version, and its list was identical on all five boxes.

Each box reported its own core design through lscpu, a command that describes the processor. The names in the picture, such as Neoverse-N1, are Arm's and AMD's names for those designs.

Every box has 2 GiB of memory, about two billion bytes; one byte holds one letter. None is burstable, a kind of box whose speed drops after a busy spell. All ran Linux, the system that runs most servers, in us-east-1, Amazon's group of data centres in Virginia, USA. Prices per hour come from AWS's own price list, read on 2026-10-11.

The Headline: Read the Unit Price

A good supermarket shows two prices on a shelf label: the price of the packet, and the unit price, such as the price per kilogram. The unit price tells you which packet is the better deal. Here are my five boxes as shelf labels.

Five supermarket shelf labels, one per box, cheapest per hour first. Each label shows the box's name on a coloured top strip, then per million answers in large type for batched, then the one-by-one price in smaller type, and below a dotted line the price per hour. a1: $0.0084 batched, $0.0385 one by one, $0.0255 an hour. c6g: $0.0046, $0.0187, $0.0340 an hour. c7g: $0.0037, $0.0110, $0.0363 an hour. c8g: $0.0025, $0.0078, $0.0399 an hour, with a dark top strip. c7a: $0.0030, $0.0098, $0.0513 an hour.

Each label's big number is the cost of a million answers with batching, with the box at its own limit. A box's limit is the most requests a second that still kept the slowest 1 in 100 under 25 ms. Under it is the cost one by one, and at the bottom the sticker price per hour. The numbers are the middle of 5 rounds of testing.

The a1.medium has the cheapest sticker, $0.0255 an hour, and the dearest unit price. One by one, a million answers on it cost $0.0385, five times the c8g.medium's $0.0078. It answered too slowly to make its low price pay.

The newest chip won. The c8g.medium was cheapest per answer both ways, and batched it gave a million answers for $0.0025, about a quarter of a cent.

Batching mattered on every box. It cut the cost per answer to between 22% and 34% of the one-by-one cost, because each hour gave 3.0 to 4.6 times as many answers.

Two warnings come with these labels. First, my test program shared each box's one core with the service: it took 4 to 9% of the core one by one, and 13 to 19% batched. So every rate here is a little low, the batched ones more. Second, the labels assume each box is busy at its limit every second. A real service is quiet for much of the day, and the meter still runs. Later slides price both of those.

How the Lab Was Built

Before any run, I wrote the lab's plan at the top of its program file, scripts/labs/serving/cost_per_prediction.py. That note at the top of a Python file is called a docstring. It names every setting and the rule for "clearly different", and it lists twelve guesses, written first so they could be checked later and could be wrong.

The service is lesson 4's, copied unchanged, with lesson 2's model and lesson 2's small database of customer numbers. It runs as one program with one thread, one line of work inside a program, which lesson 5 found fastest on one core. I ran it two ways and call each way an arm, like one recipe in a cooking test: one by one, and batched.

A tester is a program that pretends to be many users sending questions. Mine is lesson 7's honest tester. It plans every request's arrival time before the run, at random, like customers walking into a shop. Then it sends each request at its planned time, even if earlier answers are late, and times each answer from the planned arrival. Time spent waiting in a queue therefore counts. Each counted run lasted 12 seconds.

To find a box's limit, I climbed. A quick first climb, which I call the scout, found roughly where the limit was. Then I fixed ten rates around it, each about 6% higher than the last, like the steps of a staircase. A run passed if its slowest 1 in 100 took at most 25 ms and the service kept up with the arrivals.

Each box ran 5 rounds. A round tries every rate of both arms once, plus one nightly batch job. Each round runs in a new random order, so the time of day could not favour one arm. In each round, an arm's max rate is the highest rate that passed with every lower rate passing too. When I say "middle round", I mean the middle of the five values.

Here is my rule for calling two arms on one box clearly different. One arm beat the other in all 5 rounds, and even its worst round beat the other's best. Boxes ran at different hours, so for two boxes I only ask that their two sets of 5 values do not overlap at all. Anything else is .

One Machine's Day on Its Own Clock

Here is the first box's day, the c6g.medium, as one numbered sequence. The numbers on the arrows match the numbers on the recordings that follow.

A step diagram with three columns, my laptop, AWS with its logo and the c6g with the Ubuntu logo, and thirteen numbered arrows, each starting with its clock time. 1, 08:12, anything running? 2, 08:12, key, door rule. 3, 08:12, rent it. 4, 08:12, is it up? 5, 08:12, Python, files. 6, 08:14, timers off. 7, 08:15, start, leave. 8, 08:17, read console. 9 and 10 are loops on the c6g: 08:15 tests begin, and 08:50 tests end. 11, 08:52, the demo. 12, 08:53, raw files, from the c6g back to my laptop. 13, 08:53, delete it all, in red. On the right, a tall bar headed the clock, drawn to scale from 08:12 to 08:54: a short pale part to 08:15, a long dark part to 08:50 marked 35 min tests, and a short pale part to the end.

Steps 1 to 7 rent the box, set it up and start the schedule, a program on the box that runs every test in order, by itself. Step 8 is how I watched without logging in. Arrows 9 and 10 are the half hour of tests, the long dark part of the clock bar. Steps 11 to 13 came after: the student demo, the copy of the results to my laptop, and the deletion of everything. The whole day took 42 minutes.

The other four boxes went through steps 1 to 10, 12 and 13 in that order, driven by a script. On them, step 8 was a status command that reads that console, and step 11, the demo, did not run.

Build the Lab Machine Yourself

The rented boxes are part of the lab, so I show every step. You can rent this kind of box and run my schedule.

Every recording comes from the first box, the c6g.medium, the box that measured its own numbers. I did not record the other four boxes. The script that drove them is at the end of this slide.

What you need first. An AWS account, and the AWS command line tool (aws) set up with your keys. You also need four small programs you type commands into. ssh logs in to another computer, scp and rsync copy files to and from it, and curl downloads a web page. You also need lesson 2's model and data files in ~/lab-data/serving/lat, which lesson 2's lab makes.

All the steps live in one file, scripts/labs/serving/cost_files/cost_aws_steps.sh. You run it from the scripts/labs/serving folder, one step at a time, such as TYPE=c7g.medium bash cost_files/cost_aws_steps.sh check, then access, and so on. Every code block below is copied word for word from that file.

If you paste the blocks by hand, set TYPE first, such as TYPE=c6g.medium, then run these lines from the scripts/labs/serving folder. All of them come from the file, except : the file works out that folder from its own location, which a pasted line cannot do.

Steps 1 to 3: Is the Account Free? Lock the Door, Rent the Machine

Step 1: is any other box running? My account allows one core at a time, and another project may be using it. So the first command counts every box in the account that is starting, running or stopping. It must print 0.

Step 2: a key and a locked door. A key pair is how you prove to the box that you may come in. AWS keeps one half, and you keep the other half in a file only you can read. A security group is a door rule. A port is a numbered door on a computer, and SSH, the secure remote login, uses door 22. My rule opens only door 22, and only to my own address. Both get tags, small labels that say which project and lesson they belong to.

Two real terminal recordings under a strip of thirteen numbered circles with 1 and 2 lit, headed 08:12 on the c6g. Step 1: aws ec2 describe-instances with a filter for the states pending, running, stopping and shutting-down, asking for the number of instances; the answer is 0. Step 2, with the address, network and group ids hidden: create-key-pair for ai-research-course-cost with the tags Project and Lesson writes the key to the key file without printing it; chmod 600; curl reads my address; describe-vpcs finds the default network; create-security-group makes the group with its tags; authorize-security-group-ingress prints a small table, from my address only, port 22; last, ls shows the key file, readable only by its owner.

Step 1 printed 0, so the account was free. In step 2 the key went straight into its file and was never printed. This is step 1's command:

  aws ec2 describe-instances \
    --filters "Name=instance-state-name,Values=pending,running,stopping,shutting-down" \
    --query "length(Reservations[].Instances[])"

These are step 2's commands:

  aws ec2 create-key-pair --key-name ai-research-course-cost --key-type ed25519 \
    --tag-specifications "ResourceType=key-pair,Tags=[{Key=Project,Value=ai-research-course},{Key=Lesson,Value=cost}]" \
    --query KeyMaterial --output text > ~/lab-data/serving/keys/ai-research-course-cost.pem
  chmod 600 ~/lab-data/serving/keys/ai-research-course-cost.pem
  MYIP=$(curl -sf https://checkip.amazonaws.com)
  VPC=$(aws ec2 describe-vpcs --filters Name=isDefault,Values=true --query "Vpcs[0].VpcId" --output text)
  SG=$(aws ec2 create-security-group --group-name ai-research-course-cost --vpc-id "$VPC" \
    --description "lesson 11 cost per prediction, ssh from one address" \
    --tag-specifications "ResourceType=security-group,Tags=[{Key=Project,Value=ai-research-course},{Key=Lesson,Value=cost}]" \
    --query GroupId --output text)
  aws ec2 authorize-security-group-ingress --group-id "$SG" --protocol tcp --port 22 --cidr "$MYIP/32" \
    --query "SecurityGroupRules[].{port:FromPort,from:CidrIpv4}" --output table
  ls -l ~/lab-data/serving/keys/ai-research-course-cost.pem
  echo "$SG" > "$STATE/sg"

Steps 4 to 6: Wait, Set Up, Look

Step 4: wait until it runs. The AWS tool waits until the box is running. Then the script reads its public address into IP and tries SSH every five seconds until the box answers. It writes down the moment SSH first worked, because the nightly job's bill needs to know how long a box takes to start.

A real terminal recording under the strip with 4 lit, headed 08:12 on the c6g, with the instance id and address hidden. aws ec2 wait instance-running; describe-instances reads the public address into IP; describe-instances prints running, c6g.medium, us-east-1a. Then ssh with ConnectTimeout=5 runs true; it fails twice, each time followed by sleep 5, and the third try answers. The last line prints: ssh works: aarch64 Ubuntu 24.04.5 LTS.

SSH failed twice while the box was still starting, then worked. "aarch64" is Linux's name for an Arm chip. These are the commands:

  aws ec2 wait instance-running --instance-ids "$ID"
  IP=$(aws ec2 describe-instances --instance-ids "$ID" \
    --query "Reservations[0].Instances[0].PublicIpAddress" --output text)
  aws ec2 describe-instances --instance-ids "$ID" \
    --query "Reservations[0].Instances[0].[State.Name,InstanceType,Placement.AvailabilityZone]" --output text
  until ssh -i ~/lab-data/serving/keys/ai-research-course-cost.pem "${SSHO[@]}" -o ConnectTimeout=5 \
    ubuntu@"$IP" true 2>/dev/null; do sleep 5; done
  ssh -i ~/lab-data/serving/keys/ai-research-course-cost.pem "${SSHO[@]}" ubuntu@"$IP" \
    'echo "ssh works:" $(uname -m) $(lsb_release -ds)'
  date -u +%Y-%m-%dT%H:%M:%S+00:00 > "$STATE/ssh_ready"
  echo "$IP" > "$STATE/ip"
  aws ec2 describe-instances --instance-ids "$ID" --query "Reservations[0].Instances[0].LaunchTime" \
    --output text > "$STATE/launched"

Step 5: Python, the model and the lab files. The box gets Python 3.13.15 and the library versions of lessons 2 to 10, through uv, a fast installer for Python. A venv is a private toolbox folder for one project's Python, with exact versions written down. Then scp sends lesson 2's model and data, lesson 4's service and this lesson's files.

Last, the box builds its small database and scores all 26,851 request rows once, one at a time. Those are its reference scores: the box's own score for every row, kept to check every later answer against. Then it prints three fingerprints. A fingerprint is a long code worked out from a file's bytes; it changes if even one byte changes.

Steps 7, 8 and 11 to 13: Start, Watch, Demo, Bring It Home, Delete

Step 7: start the schedule, then leave. The schedule starts in the background with setsid nohup, a command that keeps a program running after I log out. Nobody logs in again until it ends.

Step 8: watch without touching. After every run, the schedule writes one progress line to the box's serial console, a log that AWS keeps for the box. You can read it through AWS without logging in.

Two real terminal recordings under the strip with 7 and 8 lit, headed 08:15 on the c6g. Step 7, with the address hidden: over ssh, cd cost, then setsid nohup venv/bin/python cost_schedule.py 5 with its output sent to raw/schedule.log, in the background; two seconds later the log shows COST-PROGRESS 08:15:04 schedule start K=5. Step 8, with the instance id hidden: aws ec2 get-console-output, piped through grep COST-PROGRESS and tail -12, prints the start line, load before the schedule 0.09 0.33 0.17 after waiting 105s, scout-none-200 with its slowest 1 in 100 at 4.95 ms, pass, and scout-none-250 at 7.35 ms, pass.

The number 5 is the count of rounds. These are the two commands:

  ssh -i ~/lab-data/serving/keys/ai-research-course-cost.pem "${SSHO[@]}" ubuntu@"$IP" \
    "cd cost && (setsid nohup venv/bin/python cost_schedule.py $K > raw/schedule.log 2>&1 < /dev/null &); \
     sleep 2; cat raw/schedule.log"
  aws ec2 get-console-output --instance-id "$ID" --latest --output text | grep COST-PROGRESS | tail -12

Step 11: the student demo, on that box. After the schedule, I ran the demo you will meet later, so you can see what it prints on one core.

A real terminal recording under the strip with 11 lit, headed 08:52 on the c6g, with the address hidden: python cost_demo.py on the box, its output also written to cost-demo-run.txt. It prints: every number below was measured on THIS machine, 1 core, 1-minute load 0.48; the chapter's model, test AP 0.5450 on 26,851 test rows; price per hour $0.1000, example price. A table of three ways of scoring, with rows, middle seconds, rows a second and the cost per 1,000 at 100% and at 30% busy: one at a time, 2,000 rows, 1,064 a second, $0.00002611 and $0.00008704; 32 at a time, 6,400 rows, 32,545 a second, $0.00000085 and $0.00000285; all at once, 26,851 rows, 339,309 a second, $0.00000008 and $0.00000027. Then: scores that differ between the three ways, to the bit: 0, and a note that a real service does fewer a second than the first row.

The Lab's Report, Running

This is a real recording of the report script, cost_report.py. It ran on my laptop, but it measures nothing. It reads the raw files the five boxes wrote and does all the arithmetic again with its own code.

A terminal recording of cost_report.py in nine numbered sections. 1: the five boxes with their chips, hours and bills, each terminated with 0 left, their runs, idle at least 0.990 and steal at most 0.0008, ssh 0; all boxes $0.1478. 2: every run's slowest 1 in 100 rechecked from the stored slowest 2%, per box. 3: each box's two staircases of rates, written as ladders, and its max rate per round. 4: the cost per million at the limit and at 30%, an example, for every box and arm. 5: batched over one by one per box from the middle rates, with the range over rounds, all clearly different, and 20 of 20 box pairs clearly different. 6: the nightly job's time per box, scores equal to the c7g.medium's on 26,851 of 26,851 rows on every box, the c6g's night as built and with a ready-made copy, and the shop's average hour at 25.3% of its busiest. 7: the twelve guesses graded, 12,700,169 served scores checked with 0 different, the demo's numbers and 8 of 8 quotes found. 8: the leak scan, 0 of each kind. 9: the playground's real numbers. Last line: all 3095 checks agree with the stored lab.

The report does not import the lab. For every run, it works out the slowest 1 in 100 again from the stored slowest 2% of waits. It also redoes each pass or fail, each max rate, each cost and each verdict, then compares every one with the lab's results file. If any number disagreed, it would stop with an error. I tried that: I changed one stored cost by a millionth of itself, and the report printed "MISMATCHES" and stopped. Then I put the file back.

Section 8 scans every stored file, even inside the packed .npz number files and the terminal recordings. It looks for my account number, IP addresses, host names, AWS ids, secret-looking codes such as key fingerprints and request ids, and my home folder. It found none of any kind.

The Meter's Sum

Every cost in this lesson comes from one small sum. Let me work it through once, slowly, with the chapter's own box, the c7g.medium, scoring one by one.

A hand-drawn taxi meter: a dark box with a slanted top and a round yellow FOR HIRE flag, holding three rows of rolling number drums. Top row: $0.0363, this hour. Written between, divided by. Middle row: 3,290,400, answers this hour. Written between, times 1,000. Bottom row: $0.0000110, per 1,000.

The box costs $0.0363 an hour. In the middle round it kept its slowest 1 in 100 under 25 ms at up to 914 requests a second. An hour has 3,600 seconds, so that is 3,290,400 answers in one hour. Divide the price by the answers and multiply by 1,000: $0.0000110 for 1,000 answers.

That number is too small to read easily, so from here on I give the cost of a million answers. For this box, one by one, a million answers cost $0.0110, about one cent.

Read this as the best case. The sum assumes the box is busy at its limit every second of the hour, and a meter keeps running when nobody rides.

Climbing the Staircase

How did I find "up to 914 a second"? The picture shows a search like it, for the c7g.medium with batching, one step at a time.

A hand-drawn staircase of ten shaded steps rising to the right, one per rate of the c7g.medium's batched search: 1,907, 2,026, 2,151, 2,285, 2,426, 2,577, 2,737, 2,907, 3,087 and 3,278 requests a second. Above each step are five marks, one per round: dots on the first six steps; on 2,737, three dots and two crosses; crosses only on the last three. A key reads: dot, under 25 ms; cross, over, or fell behind.

Each step is one rate, lowest on the left. On each step, each round left one mark. A dot means the slowest 1 in 100 stayed under 25 ms and the service kept up; a cross means it did not. The scout had placed the staircase around the limit, so the low steps passed every time and the high steps never did.

The limit moved a little between rounds. In 3 rounds the highest passing step was 2,737 a second, and in the other 2 it was 2,577, the step below. Steps are about 6% apart, so that is as finely as this search can see: a real limit anywhere between two steps lands on the lower one.

Why does the floor give way so suddenly? Below the limit, requests sometimes arrive a little faster than average, and the service catches up afterwards. Above it, the service cannot catch up; the line of waiting requests grows, and every later request waits behind it. On the limit's own step, the slowest 1 in 100 took 13 to 16 ms. On the next step up, only about 6% faster, it took 63 to 505 ms.

Every Machine Has a Knee

Now all five boxes, one by one.

A chart with requests a second along the bottom, 0 to 2,200, and the slowest 1 in 100 up the side, where each step up is ten times the last: 1, 10, 100, 1,000 ms. A dashed line marks 25 ms. One line per box through the middle round's value at each rate of its search, with small dots for the 5 rounds, each line labelled at its end. Each line stays low, then bends up sharply through the 25 ms line: the a1 near 180 a second, the c6g near 500, the c7g near 900, and the c8g and c7a close together near 1,450. Above the bend the lines climb past 1,000 ms.

Every line has one shape. It stays low and nearly flat, then bends up sharply and crosses 25 ms. I call the bend the knee. Only where it happens differs, from about 184 a second on the a1.medium to about 1,450 on the c8g.medium and c7a.medium.

The c7g.medium's knee sat at 861, 861, 914, 914 and 914 a second over the 5 rounds. Lesson 4 measured this service on a different c7g.medium, rented on another day: there 879 a second passed and 1,098 fell behind. The two boxes agree.

Batching Multiplied the Rate

Next, the two arms side by side.

A dot chart with one row per box, slowest at the bottom, and requests a second at the limit along the bottom, 0 to 6,000. Each row has five dots for one by one on the left and five for batched on the right, joined by a line between the middle rates. Beside each row, in red, the middle batched rate divided by the middle one-by-one rate, and under it the range over the 5 rounds: a1 4.6 times, rounds 4.1 to 4.9; c6g 4.1, rounds 3.8 to 4.3; c7g 3.0, rounds 2.8 to 3.2; c8g 3.1, rounds 2.8 to 3.1; c7a 3.3, rounds 3.1 to 3.3.

The red number is the middle batched rate divided by the middle one-by-one rate, so you can check it from the two dots it sits between. It was 4.6 on the a1.medium, 4.1 on the c6g.medium, 3.0 on the c7g.medium, 3.1 on the c8g.medium and 3.3 on the c7a.medium. Under it is the range of the five per-round ratios. On every box, batching won in all 5 rounds and the two sets of 5 did not overlap, so it is clearly different everywhere.

Why so much? Near the limit, requests pile up, so each call to the model took a handful of them at once. A call's fixed cost was shared.

A hand-drawn pair of seat plans seen from above. On the left, a yellow taxi with a driver's seat and one passenger seat, filled, marked one by one, 1 seat a trip, at most 914 a second. On the right, a long green bus with 32 seats in four rows of eight, 13 seats filled and 19 empty, marked batched, up to 32 seats a trip, 13 filled here; at most 2,737 a second, 12.7 a trip on average.

Think of a taxi and a bus. One by one is a taxi with one seat: each trip carries one rider. Batched is a bus with 32 seats that leaves the moment the driver is free, with whoever is waiting. On the c7g.medium at its batched limit, a trip carried 12.7 requests on average in the middle round. The bus was well under half full and still did three times the work.

The c8g.medium and c7a.medium were too close to call in rate, both one by one and batched: their sets of 5 overlapped each time.

Five Piles of Coins

The shelf labels gave the headline. This slide shows how the five boxes compare with each other.

Five pairs of coin piles, drawn in 3D, in a row, cheapest box per hour on the left: a1, c6g, c7g, c8g, c7a. In each pair the left pile is one by one and the right pile batched, and one coin is $0.0020 per million answers, rounded to whole coins. Under each pair, the two costs per million: a1 $0.0385 and $0.0084; c6g $0.0187 and $0.0046; c7g $0.0110 and $0.0037; c8g $0.0078 and $0.0025; c7a $0.0098 and $0.0030. The a1's left pile is by far the tallest.

Read left to right, the sticker price per hour goes up, but the piles do not shrink in step. The a1.medium's left pile towers over the rest.

The c7a.medium costs the most per hour, $0.0513. One by one it reached 1,453 a second against the c8g.medium's 1,428, too close to call. Its price is 29% higher, though, so per answer it came out dearer, and that difference was clear in every round.

For every pair of boxes, both arms, the two sets of 5 costs did not overlap, so all 20 comparisons are clearly different. Even so, batching moved costs more than the choice of box did. The dearest box batched, the a1.medium at $0.0084, cost only a little more per answer than the cheapest box one by one, the c8g.medium at $0.0078.

These costs share the tester's handicap from the headline: 4 to 9% of the core went to the tester one by one, 13 to 19% batched.

Who Used the One Core

The tester and the service shared one core, as in every lesson of this chapter. So how much of the core did each get?

A hand-drawn tall slab marked 1 core, cut into three shaded parts that fill it top to bottom: the service 80%, the tester 9%, the rest 11%.

This is the c7g.medium, one by one, at its limit, in the middle round. Processor time is the time the chip spent running a program. I read the tester's and the service's processor time during the counted 12 seconds and divided each by the 12 seconds. The service used 80% of the core and the tester 9%. The rest, 11%, went to the system, other programs, or nothing at all. The three parts add up to the whole core.

I had guessed the tester would take at least a quarter. That guess was wrong: one by one it took 4 to 9% on the five boxes. Batched it took more, 13% on the a1.medium up to 19% on the c7g.medium, because it sent three to four and a half times as many requests.

The limit came before the core was full. Over each whole run, Linux counted the core as idle about 10% of the time on the c7g.medium at its one-by-one limit, and about 30% on the a1.medium. So the slowest 1 in 100 went over 25 ms with time to spare. Requests arrive at random and sometimes bunch up, and on one core a bunch must wait in line, as lesson 3 showed.

In a real setup the users would be on their own computers, and the service would get the tester's share too. I did not measure how much faster it would then be, so I add nothing. My costs are pessimistic, more so for batching.

Why Real Services Run Below Their Limit

So far every cost assumed a box busy at its limit every second. No real service runs like that. I can show two reasons with this lab's own data, and add a third that is common sense rather than measurement.

A 24-hour clock face with midnight at the top, 06:00 on the right, 12:00 at the bottom and 18:00 on the left. Bars stand out from an inner ring, one per hour, as long as the shop's invoices in that hour from July to November 2011: none from 21:00 to 05:00, a few from 06:00, rising to the longest bar at 12:00, marked busiest hour, 12:00 to 13:00, then falling to almost none by 20:00. A dashed ring is marked: dashed ring, the all-day average, 25% of the busiest hour.

The first reason is that traffic has a shape. An invoice is the bill for one order, and this clock counts the shop's real invoices from July to November 2011 by hour of the day. Almost nothing happens at night. The busiest hour, 12:00 to 13:00, had 11.4 invoices on an average day. Over all 24 hours the average was 2.9 an hour, 25% of the busy hour.

A service must be big enough for its busiest hour. Say you size a box for this shop's average busiest hour. Then over the whole day it does on average only a quarter of the work it could do. Real busy days, and busy minutes inside the busy hour, run higher than that average, so the true share is lower still. And the shop's actual numbers are tiny, 11 invoices an hour is far below any box's limit; what carries over is the shape of the day.

The second reason comes from the staircase. On the c7g.medium's batched search, the scout's quick first climb, not a counted run, saw the slowest 1 in 100 at 2.3 ms at 400 requests a second. On the limit's own step, in counted runs, it was 13 to 16 ms. A service run near its limit is one burst away from the cliff.

The third is common sense: boxes break and traffic grows. With two boxes, if one fails the other must carry everything, which only works if each normally runs below half its limit.

A chart of dollars per million answers, up the side from 0 to 0.08, against the share of its full work the box does, along the bottom from 0 to 100%, for the c7g.medium batched. One curve falls steeply from the left and flattens to the right. Three points are marked: full, $0.0037; 30%, my example, $0.0123; the shop's shape, 25%, $0.0146.

The price per hour does not change when the box is idle. So the cost per answer is the cost at the limit divided by the share of work the box does. Half the work means twice the cost per answer. This curve is arithmetic on the measured cost, not a new measurement.

Always On, or Once a Night

Lesson 1 offered the other way: score everyone once a night and keep the scores. What does that cost? I priced it on the first box, the c6g.medium, because its setup was recorded and timed.

A Gantt chart, bars laid along a line of time, of one run of the c6g's nightly job, in seconds along the bottom from 0 to about 1.6. Three bars follow each other: start and stop Python 317 ms, from 0; import libraries 1,285 ms, the longest bar by far; load, read, score, write 42 ms, a short bar at the end.

A fresh Python program started and imported its libraries, which means loading their code into memory before it can be used. It loaded the model, read all 5,722 customers' numbers for 1 November 2011 from the small database, scored them all in one call, and wrote the scores. In its middle run the whole program took 1.64 seconds, timed from outside. Importing took 1,285 ms of that, as lesson 6 found, and the scoring itself 22 ms. Every one of the job's scores was equal to the box's reference scores, to the last bit, in every round on every box.

Now the bill. AWS's on-demand page says you pay "by the hour or second (minimum of 60 seconds)". A box rented only for this job must also start, and the c6g.medium took 32 seconds from launch until SSH answered. My boxes started from plain Ubuntu, so before the job they also needed Python, the packages and the files: step 5, which took 111.5 seconds. That step also made this lab's reference scores, which a nightly job would not need. So a real job's setup would be a little shorter, but I did not time it separately. Boot, setup and job together come to about 145 seconds: $0.00137 a night.

A ready-made copy of the box could already hold Python, the model and the data. Then boot plus the job would be under a minute, and the bill would be the 60-second minimum, $0.00057 a night. I did not build such a copy, so that number is worked out, not measured, and it leaves out the small monthly charge for storing the copy.

Two side-by-side day columns for the c6g, one day, on-demand price only, each a tall strip from 00:00 at the top to 24:00 at the bottom. Left, always on: the strip is completely filled, marked paid all 24 hours, with about 60 small white ticks between 06:00 and 20:00, marked white ticks, about 60 real visits a day, spread like the shop's invoices; under it, $0.816 a day. Right, once a night: an empty strip with one dark line at 02:00, marked 02:00, one short job, and paid for about two and a half minutes; under it, a zoomed bar of three minutes with a pale part, a long striped part and a thin dark line, marked 02:00 zoomed: boot 32 s, setup 112 s (striped), job 1.6 s; then $0.0014 a night, and with a ready-made copy, not built here, $0.0006.

The Same Answers on Every Chip

A cheaper box is only cheaper if it gives the same answers, so I checked.

Six rows of barcodes, each drawn from the sha256 fingerprint of one computer's 26,851 one-at-a-time scores, with the first six letters of the fingerprint on the right. The five rented boxes, c6g, c7a, c7g, c8g and a1, have identical dark barcodes, all a256a6. The last row, my laptop, has a different barcode in another colour, 2b0d36.

On each box, I scored all 26,851 request rows one at a time, the service's way, and again all at once in one call. Then I compared each box's one-at-a-time scores with the c7g.medium's.

All five boxes gave identical scores, to the last bit. They matched the c7g.medium on every row, so their fingerprints match. On every box, the all-at-once scores matched the one-at-a-time scores. And every answer the service gave in every run matched its box's reference: 12,700,169 answers, none different.

I had guessed that the AMD box would differ from the Arm boxes in the last digits, because the two kinds of chip can round some sums differently. That part of my guess was wrong. The one difference I found was elsewhere: all five boxes differed from lesson 2's laptop scores on 33 rows, as step 5 showed.

Bigger Boxes, Priced but Not Measured

Four stacks of cubes, drawn in 3D,, one cube per core, for the c7g family: medium, 1 core, $0.0363 an hour, measured; large, 2 cores, $0.0725, not measured; xlarge, 4 cores, $0.1450, not measured; 2xlarge, 8 cores, $0.2900, not measured.

My account can run only one core at a time, so I could not test bigger boxes. AWS sells each family in sizes, and each size up has twice the cores. Here are the next three after the one I used, with their published on-demand prices. None of the bigger ones was measured.

The price doubles with the cores: $0.0725 for 2 cores, $0.1450 for 4 and $0.2900 for 8. So one core has almost exactly one price at every size. A bigger box is only cheaper per answer if your service keeps all its cores busy, for example with one copy of the service for each core. Lesson 5 could not test that on one core, and I do not guess a number.

A Month's Bill, Line by Line

An account book split down the middle, for one c7g.medium on all month, 730 hours. Left, the lines; right, the amounts: the machine, 730 h x $0.0363, $26.50; its internet address, 730 h x $0.005, $3.65; its 16 GiB disk, 16 x $0.08, $1.28; a double line; total for the month, $31.43.

What would one always-on c7g.medium cost for a month? A month is about 730 hours, so the box itself is 730 hours times $0.0363. AWS also charges for its public internet address, $0.005 an hour, and for its 16 GiB disk, $0.08 per GB per month. All three prices come from AWS's price list.

The three lines add up to $31.43 a month, and the box itself is 84% of that. My costs per million so far used the box's price alone. With the address and the disk added, the c7g.medium's cost per million one by one rises from $0.0110 to $0.0131. The results file holds the full figure for every box. None of this counts traffic sent out of AWS, which this lab did not use.

My Twelve Guesses Before the Run, Checked

I wrote twelve guesses into the lab before it ran. Each is quoted word for word from the docstring, with a short note on its words. Seven were right, three were partly right and two were wrong. In the guesses, "none" means one by one, "batch" means batched, and a "box" is a rented computer.

  1. "G1. c7g.medium, none: max rate between 850 and 1,050 a second in every round (lesson 4: 879 passed, 1,098 failed)." Right: 861, 861, 914, 914 and 914.

  2. "G2. On every box, batch's max rate is at least 2 times none's, in every round." Right: 2.77 to 4.88 times, round by round.

  3. "G3. c7g.medium, batch: max rate between 2,500 and 4,000 a second in every round." Right: 2,737, 2,577, 2,737, 2,737 and 2,577.

  4. "G4. none's max rate orders the boxes c7a > c8g > c7g > c6g > a1, each neighbour pair clearly different." A neighbour pair is two boxes next to each other in that order. Partly right. The order held in the middle rounds: c7a 1,453, c8g 1,428, c7g 914, c6g 506, a1 184. But the c7a.medium and c8g.medium were too close to call; the other three pairs were clearly different.

  5. "G5. The cheapest box per hour (a1.medium) is NOT the cheapest per 1,000 predictions, for either arm; a1.medium costs more per 1,000 than c7g.medium, clearly different, for both arms." Right. It was the dearest of all five per answer.

  6. "G6. c7g.medium, none, at the limit: between $0.0000080 and $0.0000140 per 1,000 predictions (median of rounds)." The median is the middle value. Right: $0.0000110.

  7. "G7. Every served score equals its box's single-row reference to the bit, and all.npy equals ref.npy on every box. The four Arm boxes' ref.npy are identical to the bit; c7a.medium (x86) differs from c7g.medium in at least one row, by at most 1e-15." Here all.npy and ref.npy are the scores made all at once and one at a time, and 1e-15 is a millionth of a billionth. Partly right. Every served score matched, all at once matched one at a time on every box, and the four Arm boxes matched each other. The wrong part: the AMD box did not differ at all.

How I Would Price a Model Service

If I had to price a prediction service tomorrow, this is the order I would work in, with this lab's numbers as the example.

  1. Write down the target. Mine was: the slowest 1 in 100 within 25 ms. A different target gives a different rate, and so a different cost.

  2. Measure the rate on the box you will rent, with an honest tester. Send requests at planned random times, time each answer from its planned arrival, climb in small steps, and repeat the climb. The knee is sharp: on the c7g.medium batched, one step past the limit took the slowest 1 in 100 from 13 to 16 ms to at least 63 ms.

  3. Divide. Price per hour, divided by answers per hour, times 1,000. Then read the unit price, not the sticker.

  4. Turn on batching before you buy a bigger box. Here it gave 3.0 to 4.6 times the answers for no extra money.

  5. Divide again by how much of its work the box will really do. You must size for the busy hour. For a day shaped like this shop's, the box does about a quarter of its work, so the cost per answer is about four times the cost at the limit.

  6. Price the nightly job too. If a score a day old is good enough, one short job a night costs about $0.0014 on the c6g.medium as I built it, and less with a ready-made copy.

Try It Yourself

The full lab needs five rented boxes for about three and a half hours in all. The demo, cost_demo.py, runs on your own computer. It times the model alone in three ways and turns your own speed into a cost per 1,000, at a price you type in.

I wrote the demo's design into its docstring after the lab's design and before the demo first ran.

A real screenshot of VS Code with cost_demo.py open at the top of the file, showing its docstring: what it needs, how to run it, with and without --price, and the design written before it first ran.

Before you run this lab. The demo uses lesson 2's venv, with the library versions of the lab's boxes. If you made it for lessons 2 to 10, use it again. If not, make it once, inside the scripts/labs/serving/examples folder:

python3 -m venv venv-sv
source venv-sv/bin/activate        # on Windows: venv-sv\Scripts\activate
pip install -r lat_requirements.txt

I used Python 3.13. Then run python ../../features/fetch_data.py once, which downloads the shop data. Now run python cost_demo.py. To use your own price, add --price and the dollars per hour, such as python cost_demo.py --price 0.25. It needs no graphics card, no cloud account and no administrator rights. If a package is missing, it prints one line saying why and stops.

Its first line says the numbers come from your computer, with its number of cores and its load at that moment. The demo times the model alone, with no web service around it, so its numbers are not the lab's numbers. Read its three rows against each other.

r"""What would 1,000 predictions cost on YOUR machine, at a price YOU type in?

Lesson 11 of 'Serving and Inference Basics'. It uses lesson 2's Python environment, made once inside this folder
(Python 3.13 is what the lab used):
    python3 -m venv venv-sv
    source venv-sv/bin/activate              # Windows: venv-sv\Scripts\activate
    pip install -r lat_requirements.txt      # numpy, pandas, pyarrow, scikit-learn, ...
It also needs the shop data from the features chapter: run python ../../features/fetch_data.py once first. Then:
    python cost_demo.py                      # train, time three ways of scoring, print the cost table
    python cost_demo.py --price 0.25         # the same, with YOUR price per hour in US dollars
    python cost_demo.py --save out.json      # and save the numbers
If a package is missing, it prints one line saying why and stops. It needs no root, no GPU and no cloud account.
Every number it prints comes from YOUR machine, as it is right now: other programs move the times. Read the rows
against each other. They are not numbers to set beside the lab's rented machines, because this demo times the model
alone, with no web service, no feature lookup and no network around it.

Design, written 2026-10-11 after the lab's design (cost_per_prediction.py) and before this file first ran:
  1. Train the features chapter's model (it must score test AP 0.5450), as lesson 10's mm_demo.py does.
  2. Time three ways of scoring the test rows, each 5 times, keeping the middle (median) of the 5:
       one at a time   2,000 rows, one predict_proba call per row (like a service with no batching)
       32 at a time    6,400 rows, 32 rows per call (like lesson 4's batching, with full batches)
       all at once     all 26,851 rows in one call (like a nightly batch job)
     Rows per second = rows / seconds. One thread (OMP_NUM_THREADS=1), as in the lab.
  3. Check that the scores are the same in all three ways, to the bit.
  4. Cost per 1,000 predictions = your price per hour / (rows per second x 3,600) x 1,000, at 100% busy and at 30%
     busy (a service usually waits for traffic most of the time). --price sets the price; without it, the demo uses
     0.10 dollars an hour, a round example and not any real machine's price.

Author: Roni Das
Created: 2026-10-11
"""
import os

os.environ.setdefault("OMP_NUM_THREADS", "1")       # one thread per predict, as in the lab (lesson 5's advice)

import importlib.util  # noqa: E402
import json  # noqa: E402
import sys  # noqa: E402
from pathlib import Path  # noqa: E402
from time import perf_counter_ns  # noqa: E402

NEEDED = ("numpy", "pandas", "pyarrow", "sklearn")
missing = [m for m in NEEDED if importlib.util.find_spec(m) is None]
if missing:
    sys.exit(f"cost_demo.py needs {', '.join(missing)}: make the environment in its docstring "
             f"(pip install -r lat_requirements.txt), then run it with that environment's python.")

import numpy as np  # noqa: E402

REPEATS = 5
BUSY = 0.30


def build():
    """The features chapter's data and model (lesson 10's mm_demo.py build step)."""
    from sklearn.metrics import average_precision_score
    here = Path(__file__).resolve().parent
    sys.path.insert(0, str(here.parents[1] / "features"))
    import task
    from what_a_feature_is import HAND_COLS, hgb, joined

    ev = task.load_events()
    lab_tr, _, lab_te = task.splits(ev)
    tr = joined(ev, lab_tr, task.TRAIN_CUTOFFS)
    te = joined(ev, lab_te, task.TEST_CUTOFFS)
    X, y = tr[HAND_COLS].to_numpy(float), tr["label"].to_numpy()
    Xt = np.ascontiguousarray(te[HAND_COLS].to_numpy(np.float64))
    yt, ct = te["label"].to_numpy(), te["cutoff"].to_numpy()
    model = hgb(0).fit(X, y)
    p = model.predict_proba(Xt)[:, 1]
    ap = float(np.mean([average_precision_score(yt[ct == c], p[ct == c]) for c in np.unique(ct)]))
    assert round(ap, 4) == 0.5450, ap
    return model, Xt, ap


def timed(fn) -> tuple[float, np.ndarray]:
    """Median seconds of REPEATS runs of fn(), and the scores of the last run."""
    secs, out = [], None
    for _ in range(REPEATS):
        t0 = perf_counter_ns()
        out = fn()
        secs.append((perf_counter_ns() - t0) / 1e9)
    return float(np.median(secs)), out


def main() -> None:
    price = float(sys.argv[sys.argv.index("--price") + 1]) if "--price" in sys.argv else 0.10
    try:
        load = f"{os.getloadavg()[0]:.2f}"
    except (AttributeError, OSError):
        load = "not available on this system"
    n = os.cpu_count()
    print(f"Every number below was measured on THIS machine ({n} {'core' if n == 1 else 'cores'}, 1-minute load {load}).")
    model, Xt, ap = build()
    print(f"the chapter's model: test AP {ap:.4f} on {len(Xt):,} test rows")
    one = Xt[:2000]
    many = Xt[:6400]

    def one_at_a_time():
        return np.array([model.predict_proba(one[i:i + 1])[0, 1] for i in range(len(one))])

    def by_32():
        return np.concatenate([model.predict_proba(many[i:i + 32])[:, 1] for i in range(0, len(many), 32)])

    def all_at_once():
        return model.predict_proba(Xt)[:, 1]

    rows = {}
    s1, p1 = timed(one_at_a_time)
    s32, p32 = timed(by_32)
    sall, pall = timed(all_at_once)
    rows["one at a time"] = (len(one), s1)
    rows["32 at a time"] = (len(many), s32)
    rows["all at once"] = (len(Xt), sall)
    neq = int((p1 != pall[:len(one)]).sum() + (p32 != pall[:len(many)]).sum())

    unit = "your price" if "--price" in sys.argv else "example price (pass --price with yours)"
    print(f"\nprice per hour: ${price:.4f}, {unit}")
    print("                                          $ per 1,000 predictions")
    print("way of scoring     rows  middle s  rows a second   at 100% busy   at 30% busy")
    out = {"cores": os.cpu_count(), "load": load, "test_ap": ap, "price_per_hour": price, "repeats": REPEATS,
           "busy": BUSY, "ways": {}}
    for name, (n, s) in rows.items():
        rate = n / s
        c100 = price / (rate * 3600) * 1000
        c30 = c100 / BUSY
        print(f"{name:<14}{n:>8,}{s:>10.4f}{rate:>15,.0f}{c100:>15.8f}{c30:>14.8f}")
        out["ways"][name] = {"rows": n, "median_s": s, "rows_per_s": rate, "usd_per_1000": c100,
                             "usd_per_1000_at_30pct": c30}
    print(f"\nscores that differ between the three ways, to the bit: {neq}")
    print("This times the model alone. A real service also reads each request, looks up the features")
    print("and sends the answer, so it does fewer a second than the first row, and costs more per 1,000.")
    out["scores_neq"] = neq
    if "--save" in sys.argv:
        Path(sys.argv[sys.argv.index("--save") + 1]).write_text(json.dumps(out, indent=1))


if __name__ == "__main__":
    main()

Price a Day of Traffic

This box runs in your browser and needs nothing but Python. It is a calculator with a make-believe day of traffic. It uses the real rates each box reached in the lab and AWS's real prices. The day's shape follows the shop's invoices by hour. Everything else is made up: the size of the busy hour, 20,000 requests a second, and the random ups and downs.

Press Run. Then set FOLLOW_TRAFFIC = True, which changes the number of boxes every hour instead of keeping enough for the busy hour all day, and watch the bill. Try BOX = "a1.medium" and ARM = "none". One seed is luck, so try SEED = 1 to 5. The seed fixes the dice: each hour has its own dice, so changing BOX or ARM does not change the traffic.

The report script writes the boxes' real rates and prices, and the shop's invoices by hour, into this box from the lab's results. As shipped, the c7g.medium batched needs 9 boxes for the busy hour. Kept all day, they do 20 to 21% of their work over seeds 1 to 5. Following the traffic, the boxes do 66 to 67%, and with seed 1 the day's bill falls from $7.84 to $2.47.

The calculator is simpler than real life in three ways. It adds a tenth to every hour's need for bad luck, which is my choice and not a rule. It changes the number of boxes instantly, while a real box took about half a minute to start in this lab. And it assumes a box at its limit keeps its slowest 1 in 100 under 25 ms, which the lab measured only in 12-second runs.

The Lab's Code, Piece by Piece

The lab has files that run on my laptop, small files that run on the rented boxes, and the step script you saw above.

cost_per_prediction.py, on my laptop, holds the design and my guesses in its docstring, written before any run. Its commands: check makes sure every file is present; prices reads AWS's price list; status reads a box's console; cost works out one box's bill and checks that everything was deleted; collect gathers every box's raw files into the results; and factcheck looks up each quoted sentence in its source. collect calls cost_stats.py, which copies the raw files into results/cost-raw/, passes every log line through cost_redact.py to black out private details, keeps a small summary of each run, and does the arithmetic.

On the boxes, cost_files/cost_client.py is lesson 7's tester, with lesson 7's pauses taken out. cost_files/cost_ref.py makes the reference scores. cost_files/cost_batch.py is the nightly job in a fresh program. cost_files/cost_run.sh wraps every run: it waits for an idle core, starts a fresh service, and records the load and the system's log. does the scout, fixes the ten rates, and runs the 5 rounds, each in its own shuffled order. The service itself is lesson 4's , copied unchanged, which loads lesson 2's .

When to Use Each One, and When Not To

An always-on service is right when answers must be fresh, such as a score that depends on what the customer did a minute ago. It also needs traffic steady enough to keep the box busy. With thin traffic it is wrong: for this shop, one always-on box would cost $13.56 per 1,000 answers asked for.

A nightly batch job is right when a score a day old is good enough, as lesson 1 found for this shop. You must also be able to list everyone who needs a score. It is wrong when most scores are never read and the list is huge, or when a new customer needs a score before tonight.

Batching inside the service is right on any service that gets more than a few requests at a time. Here it gave every box three to four and a half times the answers for no extra money. When two requests almost never wait together, it neither helps nor hurts much.

Choosing a box by its hourly price alone is never right. Measure the rate on the real box with the real model, then divide.

What This Lab Cannot Tell You

The model is small. It is a tree model, which scores a customer by asking a chain of yes-or-no questions about their numbers, and it ran on a CPU. A neural network is a much bigger kind of model, usually run on a graphics card, and its cost per answer can be very different. Other boxes would win.

I rented one box of each type. Another box of the same type, on other hardware in the data centre, may be a few percent faster or slower. On the c7g.medium, the one-by-one limit here agrees with lesson 4's, measured on a different box.

Only one-core boxes could run, because my account allows one core at a time. Bigger boxes are priced but not measured.

The tester shared the core, so every rate includes its share and the costs are a little high, batched ones more.

I used on-demand prices only. AWS also sells cheaper ways to rent. One is a lower price for a one- or three-year promise; another is spare boxes that AWS may take back at short notice. I did not test those and do not quote their prices.

Traffic was steady and random for 12 seconds at a time. Real traffic rises and falls; the busy-hours slide prices that only with arithmetic.

The nightly bill with a ready-made copy of the box is worked out, not measured.

During the run I changed the demo's printing twice on the first box, as step 11 explains. Nothing else changed after the first box's schedule started. After an independent review, I added the setup time to the nightly bill and the idle share to the core slide. I also added the two scripts that drove the four unrecorded boxes.

What to Do on Monday

A metro line of six stops, left to right, each with a small picture above it and a word and a number below. 1, target: a stopwatch. 2, climb: a small staircase with a cracked top step. 3, divide: a coin, a division sign and a stack of bars. 4, batch: a bus. 5, the day: a clock with one shaded wedge. 6, night job?: a moon beside a sheet of paper.

On Monday, start at stop 1: write down the slowest wait your users will accept. Then climb, on the box you rent and with your own model, until you find the highest rate that keeps it. Divide the hourly price by the answers in an hour.

Stops 4 to 6 are cheap to try. Turn on batching if your service can get more than one request at a time. Look at how much of its work your box really does across a day. And if it is idle most of the day and the answers can wait until morning, price a nightly job instead.

The price of a box is the sticker; the answers it gives in an hour are the weight. Read the unit price.

Knowledge Check

Knowledge Check

4 questions - Score 80% to pass

Q1

The a1.medium had the lowest price per hour of the five boxes. Why was it not the cheapest per million answers?

Q2

On the c7g.medium, batching let the service take about three times as many requests a second. What made that possible?

Q3

A box rented only for the nightly job did about one second of work. Why could the bill be the 60-second minimum?

Q4

A box sized for the shop's average busiest hour does, over the day, a quarter of the work it could do. What does that do to its cost per answer?

Before every timed run, the core had to be at least 95% idle over two seconds. In every run on every box, it was at least 99.0% idle. AWS runs many rented boxes on one big computer, and steal is time that big computer gave to other customers' boxes during my runs. It was at most 0.08%. Nobody was logged in during any run.

too close to call

From a rate, the cost is a taxi meter's sum: the price per hour, divided by the answers in that hour, times 1,000.

HERE=$PWD
export AWS_DEFAULT_REGION=us-east-1 AWS_PAGER=""
NAME=ai-research-course-cost
TYPE=${TYPE:-c7g.medium}
AZ=${AZ:-us-east-1a}
case "$TYPE" in c7a.*|c8a.*|m7a.*|t2.*) ARCH=amd64 ;; *) ARCH=arm64 ;; esac
HERE=$PWD
STATE=${COST_STATE:-$HOME/lab-data/serving/cost/$TYPE}
KEY=$HOME/lab-data/serving/keys/$NAME.pem
K=${K:-5}
B=/home/ubuntu/cost
LAT=$HOME/lab-data/serving/lat
FEAT=$HERE/../features
mkdir -p "$STATE" "$HOME/lab-data/serving/keys"
SSHO=(-o LogLevel=ERROR -o StrictHostKeyChecking=accept-new -o "UserKnownHostsFile=$STATE/known_hosts")
[ -f "$STATE/id" ] && ID=$(cat "$STATE/id")
[ -f "$STATE/ip" ] && IP=$(cat "$STATE/ip")
[ -f "$STATE/sg" ] && SG=$(cat "$STATE/sg")

The line with case picks the right starter copy of Ubuntu for the chip: amd64 for AMD and Intel chips, arm64 for Arm chips. The last three lines read back small state files, where earlier steps saved the box's id, its address and its door rule's id.

What it costs. AWS bills each box by the second at its on-demand price. A box also has a public internet address, its number on the internet, like a phone number, at $0.005 an hour. And it has a 16 GiB disk, about 17 GB, billed at $0.08 per GB per month for the hours it existed. The c6g.medium was on for 0.698 hours, from launch to deleted, and cost $0.0284 in all. All five boxes together cost $0.1478. These numbers are in results/cost-cost.json and one results/cost-cost-<type>.json per box.

What the recordings hide. Every line of the AWS steps passed through a small filter, cost_redact.py, a copy of lesson 10's. It blacks out my account number and every IP address. It also hides every name AWS makes up for a resource, host names, my user name, my home folder and anything from a key file. Where you see <ip> or i-<id>, your screen shows the real value. Lines that start with + are the shell, the program that runs typed commands, printing each command just before it runs it.

The four boxes I did not record. After the first box was deleted, I started the c7a.medium with the script below. A few minutes later a small loop began and waited for that box to finish. Then it ran the other three in turn, ready to stop at the first that failed.

The script waits until no other box is running in my account, because another project shares the one-core limit. It stops if the launch fails, and it sends every line through the filter. It skips the demo. Both files are in cost_files/. Here they are, not recorded, copied from what ran with one change: the first cd line named my own folder.

# not recorded: the script that ran the other four machines, one at a time (cost_files/cost_runbox.sh)
set -uo pipefail
export TYPE=$1
cd "$(dirname "$0")/.."
S() { bash cost_files/cost_aws_steps.sh "$1" 2>&1 | python3 -u cost_redact.py; }
while true; do
  n=$(aws --region us-east-1 ec2 describe-instances --filters "Name=instance-state-name,Values=pending,running,stopping,shutting-down" --query "length(Reservations[].Instances[])" --output text)
  [ "$n" = "0" ] && break
  echo "$(date -u +%T) $n instance(s) running in the account; waiting"; sleep 60
done
S check; S access; S launch || exit 1; S wait; S copy; S quiet; S start
PY=$HOME/lab-data/venv/bin/python
until $PY cost_per_prediction.py status "$TYPE" | grep -q COST-SCHEDULE-DONE; do sleep 60; done
$PY cost_per_prediction.py status "$TYPE" | tail -2
S fetch; S teardown
$PY cost_per_prediction.py cost "$TYPE" | grep -E "total_usd|_left|state\""
echo "BOX-DONE $TYPE"
# not recorded: the loop over the four (cost_files/cost_chain.sh), run from the scripts/labs/serving folder
bash cost_files/cost_chain.sh

Keep the key file outside any git folder. Mine lives in ~/lab-data/serving/keys/, and the last step deletes it.

Step 3: rent the box. First the script asks AWS's parameter store, a noticeboard of current settings, for the newest Ubuntu 24.04 image for the box's kind of chip. Ubuntu is a popular version of Linux, and an image is a ready-made copy of a computer's software that a new box starts from. Then run-instances asks for one box of the chosen TYPE in one data centre, us-east-1a, with a 16 GiB disk that is deleted with the box.

A real terminal recording under the strip of thirteen circles with 3 lit, headed 08:12 on the c6g, with the image, group and instance ids hidden. aws ssm get-parameter reads the Ubuntu 24.04 arm64 image id into AMI; aws ec2 run-instances with that image, type c6g.medium, placement us-east-1a, the key ai-research-course-cost, a 16 GiB gp3 disk deleted on termination and the tags Project, Lesson and Name, asking only for the instance id; describe-instances prints the id, c6g.medium, pending and the image id.

The box started in the "pending" state, which is normal for the first few seconds. I fixed the data centre because the oldest type, a1.medium, is not offered in every one. The disk type, gp3, is a kind of Amazon disk. These are the commands:

  AMI=$(aws ssm get-parameter --query Parameter.Value --output text \
    --name "/aws/service/canonical/ubuntu/server/24.04/stable/current/$ARCH/hvm/ebs-gp3/ami-id")
  ID=$(aws ec2 run-instances --image-id "$AMI" --instance-type "$TYPE" --count 1 \
    --placement "AvailabilityZone=$AZ" --key-name ai-research-course-cost --security-group-ids "$SG" \
    --block-device-mappings '[{"DeviceName":"/dev/sda1","Ebs":{"VolumeSize":16,"VolumeType":"gp3","DeleteOnTermination":true}}]' \
    --tag-specifications \
    "ResourceType=instance,Tags=[{Key=Project,Value=ai-research-course},{Key=Lesson,Value=cost},{Key=Name,Value=ai-research-course-cost}]" \
    "ResourceType=volume,Tags=[{Key=Project,Value=ai-research-course},{Key=Lesson,Value=cost}]" \
    --query "Instances[0].InstanceId" --output text)
  aws ec2 describe-instances --instance-ids "$ID" \
    --query "Reservations[0].Instances[0].[InstanceId,InstanceType,State.Name,ImageId]" --output text
  echo "$ID" > "$STATE/id"
sha256

A real terminal recording under the strip with 5 lit, headed 08:12 on the c6g, with the address hidden and my home folder shortened to a tilde. Over ssh: make the folders; install uv, Python 3.13.15 and a venv; copy lat_requirements.txt and install it. Then scp copies lesson 2's model.pkl, store.parquet and requests.parquet, lesson 2's service.py, build_store.py and box_info.sh, lesson 4's bat_service.py and this lesson's cost_ files, then the features chapter's files, the shop data and the demo. Last, over ssh: store: 26851 rows, sqlite 3.53.1; then a line in braces with rows 26851, ref_equal_lesson2_expected 26818, all_equal_ref 26851 and a largest difference of 0.0; then three sha256 fingerprints for model.pkl, store.parquet and requests.parquet.

The file names do not matter much; the numbers do. The fingerprints match lesson 2's files. In the line in braces, 26,818 of the 26,851 scores were equal to the scores lesson 2 made on my laptop. The other 33 differed in the last digits, probably because a different maths library runs on macOS. I did not test why. That is the reason every box makes its own reference scores. This step took 111.5 seconds by its own recording, which matters for the nightly bill later. These are all its commands, run after step 4 has set IP:

  S=(ssh -i "$KEY" "${SSHO[@]}" ubuntu@"$IP")
  C=(scp -q -i "$KEY" "${SSHO[@]}")
  "${S[@]}" "mkdir -p $B/raw $B/features/results $B/serving/examples ~/lab-data/features"
  "${S[@]}" "curl -LsSf https://astral.sh/uv/install.sh | sh > /dev/null 2>&1 && \
    ~/.local/bin/uv python install 3.13.15 > /dev/null 2>&1 && \
    ~/.local/bin/uv venv --allow-existing --python 3.13.15 $B/venv > /dev/null 2>&1"
  "${C[@]}" "$HERE/examples/lat_requirements.txt" ubuntu@"$IP":$B/
  "${S[@]}" "~/.local/bin/uv pip install --python $B/venv/bin/python -q -r $B/lat_requirements.txt"
  "${C[@]}" "$LAT/model.pkl" "$LAT/store.parquet" "$LAT/requests.parquet" \
    "$HERE/lat_files/service.py" "$HERE/lat_files/build_store.py" "$HERE/lat_files/box_info.sh" \
    "$HERE/bat_files/bat_service.py" "$HERE"/cost_files/cost_*.py "$HERE"/cost_files/cost_*.sh ubuntu@"$IP":$B/
  "${C[@]}" "$FEAT/task.py" "$FEAT/what_a_feature_is.py" ubuntu@"$IP":$B/features/
  "${C[@]}" "$FEAT/results/data-manifest.json" ubuntu@"$IP":$B/features/results/
  "${C[@]}" "$HOME/lab-data/features/retail.parquet" ubuntu@"$IP":lab-data/features/
  "${C[@]}" "$HERE/examples/cost_demo.py" "$HERE/examples/lat_requirements.txt" ubuntu@"$IP":$B/serving/examples/
  "${S[@]}" "cd $B && OMP_NUM_THREADS=1 venv/bin/python build_store.py store.parquet store.sqlite \
    && OMP_NUM_THREADS=1 venv/bin/python cost_ref.py | tee raw/ref-build.json && sha256sum model.pkl store.parquet requests.parquet"

Step 6: stop the timers and look at the box. Ubuntu runs small jobs on a timer, such as checking for updates. One of them starting in the middle of a run would land in the results, so the script stops every timer first. Then it prints the chip, the load and every installed package with its version.

A real terminal recording under the strip with 6 lit, headed 08:14 on the c6g, with the address hidden. Over ssh: stop every active timer and the update service, after which the timer list says 0 timers listed; lscpu prints architecture aarch64, 1 CPU, vendor ARM, model name Neoverse-N1, 1 thread per core, 1 core per socket; nproc prints 1; uptime shows a load average of 0.65. Then Python 3.13.15 and the full pip freeze, including fastapi 0.142.2, numpy 2.5.3, scikit-learn 1.9.1, uvicorn 0.54.0 and uvloop 0.23.0.

The box reports one CPU, a Neoverse-N1. The load average is how many programs, on average, wanted the core over the last minute. A value of 0.65 is normal right after setup, and the schedule waits for it to settle. These are the commands:

  ssh -i ~/lab-data/serving/keys/ai-research-course-cost.pem "${SSHO[@]}" ubuntu@"$IP" \
    'for t in $(systemctl list-units --type=timer --state=active --no-legend --plain | cut -d" " -f1); do
       sudo systemctl stop "$t"; done; sudo systemctl stop unattended-upgrades.service
     systemctl list-timers --no-pager | tail -2'
  ssh -i ~/lab-data/serving/keys/ai-research-course-cost.pem "${SSHO[@]}" ubuntu@"$IP" 'lscpu | head -14; nproc; uptime'
  ssh -i ~/lab-data/serving/keys/ai-research-course-cost.pem "${SSHO[@]}" ubuntu@"$IP" \
    'cd cost && bash cost_box_info.sh before > raw/box-info-before.txt 2>&1; venv/bin/python --version; \
     ~/.local/bin/uv pip freeze --python venv/bin/python'

This is the command:

  ssh -i ~/lab-data/serving/keys/ai-research-course-cost.pem "${SSHO[@]}" ubuntu@"$IP" \
    'cd cost/serving/examples && ../../venv/bin/python cost_demo.py --save ~/cost/raw/cost-demo.json \
     | tee ~/cost/raw/cost-demo-run.txt'

The demo ran three times on this box. The first run's table was wider than the terminal, so its lines wrapped. I narrowed the table, copied the new file to the box and ran the step again. The second run's last sentence still wrapped, so I shortened it and did that once more. The recording above is the third run, and it is the one stored in results/cost-demo-run.txt. The first two runs' numbers were not kept, because each run wrote over the last. The two copies were not recorded. This is the command I used both times:

# not recorded: copy the fixed demo to the box, before running step 11 again
scp -q -i $KEY -o LogLevel=ERROR -o UserKnownHostsFile=$STATE/known_hosts \
  examples/cost_demo.py ubuntu@$IP:/home/ubuntu/cost/serving/examples/

Step 12: bring the results home first. A whole box of results was lost once in this chapter, because it was deleted before anyone copied them. So the raw files come home before anything is deleted. Step 12 also ran three times, once after each demo run; the recording is the last.

A real terminal recording under the strip with 12 lit, headed 08:53 on the c6g, with the address hidden. Over ssh, the box records its information after the schedule and prints the size of its raw folder, 11M; then rsync copies the box's cost/raw folder to my laptop; ls and wc count 250 files; du prints 10M for the copy.

The 250 files came to 10 MB. These are the commands:

  ssh -i ~/lab-data/serving/keys/ai-research-course-cost.pem "${SSHO[@]}" ubuntu@"$IP" \
    'cd cost && bash cost_box_info.sh after > raw/box-info-after.txt 2>&1; du -sh raw'
  rsync -a -e "ssh -i $KEY ${SSHO[*]}" ubuntu@"$IP":cost/raw/ "$STATE/raw/"
  ls "$STATE/raw/" | wc -l
  du -sh "$STATE/raw/"

Step 13: delete the box, the key and the door rule, and check. A box you forget keeps costing money. The script deletes the box, waits until AWS says "terminated", then deletes the key pair and the security group, and asks AWS again.

A real terminal recording under the strip with 13 lit, headed 08:53 on the c6g, with the instance and group ids hidden. terminate-instances prints shutting-down; wait instance-terminated; delete-key-pair prints true; delete-security-group prints true. Then the checks: describe-instances prints terminated; the count of key pairs named ai-research-course-cost is 0; the count of security groups of that name is 0; the count of volumes tagged Lesson=cost is 0; the count of course instances not terminated is 0. Last, the key file is removed.

Every count printed 0 and the box printed terminated. Those checks printed 0 after each of the other four boxes, as their cost files record. This recording printed more lines than its window held, so its first line had scrolled away. I drew the picture again from that recording file with a taller window. These are the teardown commands:

  aws ec2 terminate-instances --instance-ids "$ID" --query "TerminatingInstances[0].CurrentState.Name" --output text
  aws ec2 wait instance-terminated --instance-ids "$ID"
  aws ec2 delete-key-pair --key-name ai-research-course-cost --query Return
  until aws ec2 delete-security-group --group-id "$SG" --query Return 2>/dev/null; do sleep 10; done
  date -u +%Y-%m-%dT%H:%M:%S+00:00 > "$STATE/terminated"
  aws ec2 describe-instances --instance-ids "$ID" --query "Reservations[0].Instances[0].State.Name" --output text
  aws ec2 describe-key-pairs --filters "Name=key-name,Values=ai-research-course-cost" --query "length(KeyPairs)"
  aws ec2 describe-security-groups --filters "Name=group-name,Values=ai-research-course-cost" \
    --query "length(SecurityGroups)"
  aws ec2 describe-volumes --filters "Name=tag:Lesson,Values=cost" --query "length(Volumes)"
  aws ec2 describe-instances --filters "Name=tag:Project,Values=ai-research-course" \
    "Name=instance-state-name,Values=pending,running,stopping,stopped,shutting-down" \
    --query "length(Reservations[].Instances[])"
  rm -f ~/lab-data/serving/keys/ai-research-course-cost.pem

The security group can only be deleted once the box is fully gone, so the script retries that command every ten seconds until AWS accepts it.

For the c7g.medium batched, a million answers cost $0.0037 at full work. At 30%, a round number I picked as an example, not a rule, they cost $0.0123. At the shop's 25%, they cost $0.0146.

Here is the c6g.medium both ways, per day. Always on, it costs $0.816 a day. Once a night, as built, it costs $0.00137, or $0.00057 with a ready-made copy.

Per answer, count what the shop actually asked for. Lesson 1 replayed 7,403 real visits over 123 days, about 60 a day. Per 1,000 answers asked for, always on costs $13.56, and the nightly job $0.023 as built, or $0.0094 with a ready-made copy. These comparisons are arithmetic on measured prices and times, not new measurements.

For this shop, both costs are tiny. Lesson 1 also found that most nightly scores are never read, and that a score can be up to a day old. Still, the gap is large: a quiet service pays for its idle hours, and a batch job pays only for its minutes.

  • "G8. The nightly job for 5,722 customers takes under 5 s of wall time on every box, so a job-only box's nightly bill is set by boot plus the 60-second minimum, not by the scoring." Wall time is time on a clock on the wall, from start to end. Partly right. The first half held: the middle runs took 0.89 seconds on the c7a.medium up to 3.26 on the a1.medium.

    The second half was wrong for the boxes as I built them: from plain Ubuntu, setup added about two minutes, so the night ran about 145 seconds. It would hold only with a ready-made copy of the box, which I did not build.

  • "G9. At none's max rate, the tester uses at least 25% of the core's busy time on every box." Busy time is time the core was not idle. Wrong: 4 to 9% of the core's time, and under a tenth of its busy time.

  • "G10. On every box and arm, the 5 rounds' max rates differ by at most 10% (max / min <= 1.10)." That is, the highest round divided by the lowest. Wrong. Five of the ten box-and-arm pairs stayed within 10%, and five did not. The widest was the a1.medium batched, from 749 to 897 a second, 1.20 times. Part of this is the staircase of rates itself: a limit that moves by two steps moves by about 13%.

  • "G11. Launch to SSH-ready under 60 s for every box." Right: 28 to 40 seconds.

  • "G12. Shop traffic (lesson 1's data): averaged over the 24 hours of a day, invoices per hour are between 0.20 and 0.40 of the busiest hour of the day's average; so a box sized for the busy hour is mostly idle." Right: 0.25.

  • You saw the demo running on the c6g.medium in step 11. On that one core, the model scored 1,064 rows a second one at a time, 32,545 rows a second 32 at a time, and 339,309 rows a second all at once. The scores were identical all three ways. That core served only 506 requests a second one by one in the lab. A real request also has to be read, checked, looked up in the database and answered, and the tester shared the core.

    Here is the demo in VS Code on my laptop, run with python cost_demo.py in the examples folder with venv-sv active. My laptop's name shows in the prompt; only the AWS recordings went through the filter.

    A real screenshot of the VS Code terminal on my laptop, in the venv-sv environment, after python cost_demo.py. The first line says the laptop has 10 cores and was busy with other programs: its 1-minute load was 4.49. The chapter's model scored test AP 0.5450 on 26,851 test rows. The price is the example price, $0.1000 an hour, because I did not pass --price. The table has three rows. One at a time: 2,000 rows in 0.3579 seconds, 5,588 rows a second, $0.00000497 per 1,000 at 100% busy and $0.00001657 at 30% busy. 32 at a time: 6,400 rows in 0.0431 seconds, 148,593 rows a second, $0.00000019 and $0.00000062. All at once: 26,851 rows in 0.0326 seconds, 824,654 rows a second, $0.00000003 and $0.00000011. Then: scores that differ between the three ways, to the bit: 0. Last, a note that this times the model alone, so a real service does fewer a second and costs more per 1,000.

    This run is from my laptop, not a rented box. It is a different computer: it has 10 cores, not 1, and other programs were busy on it at that moment. The price is a made-up example, $0.10 an hour, not what my laptop costs. So do not set these numbers beside the rented boxes' numbers one by one.

    Compare inside the run instead. Scoring 32 at a time gave about 27 times as many rows a second as one at a time, 148,593 against 5,588, so the cost per 1,000 at a fixed price fell by about that much. On the c6g.medium the demo gave about 31 times. The scores were identical all three ways. And remember the last line: this is the model alone. The lab's service also reads each request, looks up the customer and sends the answer, so it does far fewer a second.

    cost_files/cost_schedule.py
    bat_files/bat_service.py
    service.py

    cost_files/cost_runbox.sh and cost_files/cost_chain.sh drove the four unrecorded boxes. cost_report.py does not import any of the above; it reads the stored raw files with its own code and recomputes every number.