Serving And Inference

Batch or Online Scoring: A Month-Old Score Cost Nothing I Could Measure Here

0 of 31 complete

0%

Contents

Back|Serving And InferenceBatch or Online Scoring: A Month-Old Score Cost Nothing I Could Measure Here
1/31
72 min left
Prerequisites
What a Feature Is: A Better Model or a Better Feature?requiredFeature Freshness: Two Weeks Stale Cost Nothing I Could MeasurerequiredBatch vs Streaming Data for ML: When Fresh Beats Cheaprequired
Related Topics
Model Signatures: The Right Numbers in the Wrong Shape, and What a Schema Check CatchesPackaging, Registry and VersioningWrong Labels: How Many Can a Model Survive, and Can You Find Them?Data Engineering for MLRebalance, or Just Move the Threshold? Measured on Rare ClassesData Engineering for MLLeakage Before the Split: How Pure Noise Scored 93% AccuracyData Engineering for MLWhat a Feature Is: A Better Model or a Better Feature?Features and Feature Stores
1 of 31

A Card for Every Guest, or One at the Door?

Let me start with a party.

You are hosting a big party, and you want every guest to get a small card with their name and a short note on it. You can do this in two ways.

A flat illustration of a woman at a wooden table gluing small white cards, with a neat stack of finished cards beside her and many more laid out on the table. Below the picture: cards made the night before are ready at once, but most guests never come, and a card cannot know what happened after it was made; a card made at the door is always up to date, and only made for someone who came.

The first way is the one in the picture. The night before, you sit down and write a card for every person on your list. When a guest walks in, you hand over their card at once. It is fast at the door. But you wrote many cards for people who never came. And if a guest told you some news this morning, the card does not know it, because you wrote it last night.

The second way is to write each card at the door, when the guest arrives. Every card is up to date, and you never write a card nobody uses. But each guest waits a moment while you write, and you must be at the door the whole evening.

A machine learning model that gives each customer a number faces the same choice. In this lesson I measure both ways on a real shop.

Where This Lesson Starts

This is the first lesson of a new chapter about serving a model: running a trained model so that other programs can ask it questions and get answers.

The model is the one from the features chapter. If it is new to you, please read what a feature is first. In short: a real online shop in the UK sold gifts, mostly to other shops, from 2009 to 2011. For each customer, the model looks at six numbers made from their past orders. Two examples are the days since their last order and how many orders they made. Then it guesses how likely they are to buy again in the next 30 days.

The survey lesson on model serving already explains, in words, the two main ways to serve a model. In batch serving, a job scores many rows on a schedule. In online serving, each request is scored as it comes. It says the main question is whether someone is waiting for the answer. I will not repeat it. Here I measure a different part of the choice: how much an old stored answer costs, how much work a batch wastes, and who it misses.

Two earlier lessons measured something close. Feature freshness found that features up to 14 days old cost nothing measurable on this same shop. And batch vs streaming data found that fresh counts cut the error almost in half for bike rentals in the next hour. This lesson sits between them.

The Words You Need First

Please read this slide slowly if any word is new. Every slide after it uses these words.

A hand-drawn grid of ten cards, two per row. Score: a number from 0 to 1 the model gives one customer, how likely they buy again within 30 days. Batch scoring: a job that scores every customer on a schedule and stores the scores for later. Online scoring: working out one customer's features and score at the moment the score is needed. Request: here, one real visit, a customer's invoice at one moment, when a score is needed. Schedule: how often the batch job runs, here daily, weekly or monthly, always at midnight. Stale: old, the age of a stored score when it is used, in hours since the job ran. New customer: someone first seen after the last batch run, so the batch has no score for them yet. Fallback: what is served when no score exists, here the model's score on an empty row. Seed: a number that fixes the random part of training; 20 seeds give 20 slightly different models. Bootstrap: redrawing the customers 1,000 times to see how far a result moves by luck alone. Below: AP, features and the top fifth mean what they meant in the features chapter.

A score is the model's answer for one customer: a number between 0 and 1. A higher score means the model thinks the customer is more likely to buy again within 30 days.

Batch scoring is like writing all the cards the night before. A job wakes up on a schedule, for example every night at midnight. It gives a score to every customer the shop knows, and stores the scores in a table. Later, when the shop needs a score, it looks it up. Online scoring is like writing the card at the door. When the shop needs a score, it works out the customer's six numbers right then, from everything that has happened so far, and asks the model.

A request is one moment when the shop needs a score. Here it is a real visit: one invoice from a customer. Most invoices are orders, but returns count too: 1,158 of the 7,403 visits (15.6%, about 1 in 6) were return-only invoices, and 22 mixed returns and purchases. A stored score is stale when it is old. I measure how stale it is in hours, from the moment the job ran to the moment the score was used.

AP, average precision, is the score from the features chapter. It goes from 0 to 1 and says how well the scores put the customers who did buy again above the ones who did not.

Two Ways to Get One Customer's Score

Here are the two ways side by side, for one real visit from the lab: customer 12362, who placed an order on 26 October 2011 at 13:47.

A sequence diagram with four lifelines: nightly job, score table, shop app and model. Step 1, score all, from the nightly job to the score table. Step 2, read, from the shop app to the score table. Step 3, 13.8 h old, back to the shop app. Step 4, ask now, from the shop app to the model. Step 5, score now, back to the shop app. Below: step 1 runs at 00:00, step 2 at the visit, 13:47; steps 1 to 3 are batch scoring, the score was made at midnight and stored; steps 4 and 5 are online scoring; for this visit the stored score was 0.294 and the one made at the visit 0.245, seed 0. Caption: same model, same customer, only the moment the inputs were read differs.

In the batch way, a job ran at midnight. It worked out the six numbers for this customer as they were at 00:00, asked the model, and stored the answer. At 13:47, when the visit came, the shop read that stored answer. It was 13.8 hours old.

In the online way, the shop works out the six numbers at 13:47, from every order before that moment, and asks the model right then.

Both ways use the same model, trained in the same way. The only difference is when the inputs were read. So the whole question of this lesson is this: does reading the inputs earlier make the answer worse, and by how much?

For this one visit, the two scores were 0.294 and 0.245. One visit cannot tell you which way is better. For that, I need thousands of visits.

What the Sources Say

This lesson leans on three facts about tools and data. I downloaded each source at a fixed version and found every quote in it word for word. They are in results/bo-factcheck.json, which the lab's fact-check mode wrote after the main run.

Three rows. scikit-learn 1.9.1, an empty row still gets a score, with the scikit-learn logo: if no missing values were encountered for a given feature during training, then samples with missing values are mapped to whichever child has the most samples. Feast, a feature store, a batch command fills the online store: is the materialize_incremental command, which fetches the latest values for all entities in the batch source and ingests these values into the online store. UCI Online Retail II, the shop's real invoices: non-store online retail between 01/12/2009 and 09/12/2011, licence CC BY 4.0. Caption: 3 quotes, every one found word for word in the fetched file.

scikit-learn is the Python library the model comes from. Its documentation says what happens when a value is missing at prediction time, and the model never saw a missing value in training. The row goes to "whichever child has the most samples". So even a row with all six numbers missing gets a score. I use that as the fallback for customers the batch has no score for.

Feast is a , a tool from the features chapter. Its documentation describes a command, materialize_incremental, that "fetches the latest values for all entities in the batch source and ingests these values into the online store." In plain words: a batch job copies fresh numbers into the fast store that serving reads. That is batch work feeding online serving, and it shows that real systems often mix the two.

The data is the UCI Online Retail II dataset, real invoices from a UK shop, shared under the CC BY 4.0 licence.

How the Lab Was Built

I wrote the lab's design into the docstring of scripts/labs/serving/batch_or_online.py before it first ran. Before that, I had read the features chapter's code and the names of the data's columns. I had not counted or scored any visit in the test months.

A page in four labelled zones, titled one model, real visits, four ways to score them. The model: lesson 1's model from the features chapter, 20 seeds, test AP 0.5450 for seed 0; trained on rows made on the 1st of each month; nothing tuned here. The requests: 7,403 real visits, July to October 2011, by 2,771 customers; label, the customer buys again within 30 days after the visit, 48.2% did. Four ways to score: online, at the visit; or a batch job at 00:00, run daily, weekly from Friday 1 July, or monthly, and its stored score served. Measured: how old each served score was; AP per month on the visits every way could score, 20 seeds, a customer bootstrap; scores made per day, never used, and customers with no score yet. Caption: guess, written first: daily costs nothing measurable, monthly costs clearly; the monthly half was wrong.

The model. The lab trains the features chapter's lesson 1 model 20 times, with seeds 0 to 19, on the same training rows. It stops unless every seed gives exactly the test AP that lesson stored; seed 0 gives 0.5450. Nothing is tuned or chosen in this lab.

The requests are real. Each request is a real moment from the shop's data: one customer's invoice at one time, between 1 July and 31 October 2011. Several invoice lines with the same time count as one request. That gave 7,403 requests from 2,771 customers.

The label. For a visit at time t, the label is 1 if the same customer makes another purchase after t and within 30 days. I left out the visit's own order, because counting it would make every label 1. This is the chapter's question, "will they buy in the next 30 days?", asked at the moment of the visit instead of on the 1st of the month.

Four ways to score each visit. Online: the six numbers worked out at the visit. Daily: a job at midnight every day, and the visit gets the score from the last midnight. Weekly: the same every 7 days from Friday 1 July. Monthly: the same on the 1st of each month. Every way uses the same trained models.

Making the Comparison Fair

Before the results, two choices in the design matter a lot, so I want to explain them.

First, the same visits for every way. Some visits come from customers a batch has never scored.

A daily batch at midnight has no score for someone whose first order was at 09:00 today and who comes back at 15:00. If I let each way skip the visits it cannot score, the four ways would be graded on different visits, which is not fair. So the main comparison uses only the 6,613 visits that all three batches could score. I call it the common set. A second comparison uses every visit with any history, 6,787, and there a batch serves its fallback score for the visits it has no score for.

Second, how sure the result is. Training has a random part, so I trained 20 models with 20 seeds and report the mean. And which customers happened to visit is luck too. So the lab redraws the customers 1,000 times, keeping all of a customer's visits together in each month, and works out the 20-seed mean every time. This is a bootstrap. If the middle 95% of the redrawn differences crosses zero, I say the batch "cannot be told apart" from online. It does not mean the two are equal. It means this data cannot show a difference.

AP is worked out for each month, July to October, and then averaged, as the features chapter did for each month's cutoff.

The lab also checks its own feature code. At every one of 124 midnights, its six numbers for every customer were exactly equal to what the features chapter's own function gives. For 300 visits drawn at random, the visit-time numbers were exactly equal too.

The Lab's Report, Running

This is a real recording of the report script, bo_report.py, on the laptop where the lab ran.

A terminal recording of bo_report.py in ten sections. The model: test AP seed 0 0.5450, 20 of 20 seeds equal lesson 1. The visits: 7,403 visits, 2,771 customers, label rate 0.482, same rows as the lab. Features by lesson 1's build_tables: 123 nights in one call, 2,419 customers' visits one call each, common set 6,613. Age of the served score: daily median 12.9 hours, weekly 108.2, monthly 394.3; visits after one the score never saw 826, 1,294 and 2,530. AP, mean of 20 seeds: online 0.8308; daily 0.8315, -0.0018 to +0.0029; weekly 0.8326, -0.0013 to +0.0052; monthly 0.8324, -0.0027 to +0.0063; the history set the same way. After the results: rank 0.9906, 0.9842, 0.9657; frequency alone 0.8152. Compute: daily 5,353 a day, 99.1% never used; weekly 761, 94.2%; monthly 172, 80.7%; online 60 a day. The worked example, customer 12362. After the review, a model trained on visit-time rows: recency under 1 day, 1st-of-month training rows 1.7%, visit training rows 18.5%, online test rows 13.9%; common set, daily 0.8345 against online 0.8338, -0.0012 to +0.0026, below 4 of 20; weekly 0.8333, -0.0029 to +0.0019, below 11; monthly 0.8332, -0.0037 to +0.0031, below 16; the history set the same way; on the missed visits, monthly +0.0002 to +0.0077, below 0 of 20; return-only invoices 1,158, mixed 22. Demo, playground and fact-check equal the lab. Last line: all 494 checks agree with the stored lab.

The report does not trust the lab. It never imports the lab's code. It builds every feature with the features chapter's own function, build_tables, which the lab did not use for its main pass. For the batches, it calls that function once for all 123 midnights. For online, it calls it once per customer, with that customer's visit times as the moments to measure at.

It rebuilds the visits and their labels with different pandas code, and trains the 20 models again. It scores every way with scikit-learn's AP, and redoes every bootstrap with an AP function written inside the report. It retrains the visit-time model of the review slide from its own rows too. Then it checks the student demo's stored output, the playground and the fact-check. All 494 checks agreed with the stored lab.

When the Visits Came

Before any score, here are the requests themselves. Each one is a moment when the shop needed an answer.

A calendar of the four test months with one row per week, from Friday to Thursday, and one cell per day holding that day's number of visits, for example 53, none, 27, 46, 74, 72 and 87 in the week of 1 July. Thick cells mark the 1st of each month. Every Saturday cell says none. Below: each row is one weekly batch period, from Friday 00:00; a thick cell is the 1st of a month, when the monthly batch runs; 19 days had no visit, 18 Saturdays and 1 other day; the busiest day had 176. Caption: 7,403 visits; each one is a moment the shop needed a score.

The data has no orders on any Saturday in these months. On the days with orders there were between 27 and 176 visits. The busiest day was Thursday 6 October.

Each row of the calendar is one week of the weekly batch. The job runs at the start of the row, on Friday at midnight. Every visit in that row is served its score. A monthly batch runs on the four thick cells, and its scores serve a whole month.

This picture explains the first cost of online scoring too. The service must be ready on any day at any hour, even if only 27 people come. A batch can run once at night and be switched off. I do not measure money or speed in this lesson, but keep that in mind.

How Old Was the Score When It Was Used?

A stored score gets older every hour after the job runs. Here is how old it was at each visit.

A hand-drawn bar chart of visits by how many hours had passed since 00:00 for the daily batch: 6 hours 21, 7 hours 7, 8 hours 137, 9 hours 483, 10 hours 831, 11 hours 782, 12 hours 1,136, 13 hours 906, 14 hours 794, 15 hours 705, 16 hours 402, 17 hours 244, 18 hours 90, 19 hours 61, 20 hours 14. Below: on the 6,613 visits every way could score, median 12.9 hours, 90% under 16.3, never more than 20.6. Caption: no visit came before 6.2 hours after midnight, so a nightly score was never fresh.

For the daily batch, the age of the score is simply the time of day of the visit. Most orders came between 9 in the morning and 4 in the afternoon. So the typical nightly score was 12.9 hours old when it was used, and it was never fresher than 6.2 hours.

A dot chart of the age of the served score in days for the three batches. Daily: median 0.5 days, 90% under 0.7. Weekly: median 4.5 days, 90% under 6.6. Monthly: median 16.4 days, 90% under 27.5. Below: daily, median 0.5 days, 90% under 0.7, oldest 0.9; weekly, median 4.5 days, 90% under 6.6, oldest 6.9; monthly, median 16.4 days, 90% under 27.5, oldest 30.7. Caption: online scores are made at the visit, age zero.

The weekly and monthly scores were much older. Half of the monthly scores were more than 16 days old when a visit used them, and the oldest was 30.7 days. If old scores cost anything, the monthly batch is where I expected to see it.

What a Stored Score Missed

An old score is only a problem if something happened since it was made. So the lab counted the visits where the same customer had already placed another order between the batch run and this visit.

Three panels. Daily batch: 12.5%; 826 visits; 66.5% of them bought again, against 48.3% of the rest. Weekly batch: 19.6%; 1,294 visits; 67.9% of them bought again, against 46.3% of the rest. Monthly batch: 38.3%; 2,530 visits; 68.9% of them bought again, against 39.2% of the rest. Below: here a visit is missed when the same customer had another invoice between the batch run and this visit; those customers were busier, and more of them bought again. Caption: the monthly score had missed something for more than a third of the visits.

For the daily batch, 826 visits, 12.5%, came from a customer who had another invoice earlier the same day. The stored score did not know about that earlier invoice. For the monthly batch it was 2,530 visits, 38.3%: more than a third of all the visits came after an invoice the monthly score never saw.

And these were not random visits. The customers who came back quickly were busy buyers, and 68.9% of them bought again within 30 days, against 39.2% of the other visits. So the information the stored score missed was exactly about the customers most likely to buy.

This made me confident before the run that the monthly batch would lose AP. The next slide shows that it did not, at least not by an amount this data can see.

The Headline: No Schedule Could Be Told Apart

Here is the main result. It is AP on the 6,613 visits all four ways could score, as the mean of 20 models.

An interval plot against a zero line. Daily, 0.8315: -0.0018 to +0.0029. Weekly, 0.8326: -0.0013 to +0.0052. Monthly, 0.8324: -0.0027 to +0.0063. Each row has a dot for the 20-seed mean difference and a bar for the 95% interval, on an axis from -0.004 to +0.008 labelled batch AP minus online AP, 95% interval. Below: online scored 0.8308; the dot is each batch's 20-seed mean minus online's; the bar is the 95% interval from 1,000 redraws of the customers; every bar crosses zero. Caption: the model here was trained on 1st-of-month rows; a model trained on visits gave the same verdict.

Online scored 0.8308. The daily batch scored 0.8315, the weekly 0.8326 and the monthly 0.8324.

All three batch numbers were a little higher than online, and none of the differences could be told apart from zero. For the monthly batch, the middle 95% of the bootstrap went from -0.0027 to +0.0063. So the data is consistent with the monthly score being a little worse than online, a little better, or the same.

I want to be careful with the words here. This is not "batch is better". The differences are smaller than the luck in the data. And it is not "staleness never matters". It is: on this shop, for this 30-day question, scores up to 30.7 days old cost nothing I could measure.

These numbers come from the features chapter's model, which learned from rows made on the 1st of each month. An independent review pointed out that this tilts the race toward batch, and a later slide measures how much. In short: a second model, trained on visit-time rows, gave the same verdict. With it, no schedule could be told apart from online either.

My guess before the run was that the monthly batch would cost clearly. That guess was wrong, with both models.

Twenty Models, Four Ways Each

The headline uses the mean of 20 models. Here is each model on its own.

A dot chart with four columns, online, daily, weekly and monthly, and 20 dots in each, one per training seed, on an axis of AP from 0.824 to 0.838. Below: online 0.8257 to 0.8342; daily 0.8266 to 0.8351; weekly 0.8285 to 0.8357; monthly 0.8294 to 0.8349; seeds below online: daily 3, weekly 2, monthly 4 of 20. Caption: the spread from training luck is wider than the gaps between the four ways.

The 20 online models scored from 0.8257 to 0.8342, a spread of about 0.009. The gaps between the four ways were about 0.002 or less. So the choice of seed moved AP more than the choice between batch and online did.

Within each seed, the batch scored below online for only a few of the 20 seeds: 3 for daily, 2 for weekly and 4 for monthly. So the stored scores had a small edge in most seeds, 16 to 18 of 20. The bootstrap on the last slide says that edge is not bigger than customer luck, but a lean in one direction that steady needed a reason.

The review found one. The model learned from rows made at midnight on the 1st of a month. In those training rows, only 1.7% of customers had their last invoice less than a day before, and the median was 61.5 days. In the online test rows, 13.9% did, with a median of 23.0 days. So online asked the model about rows it had rarely seen, while a batch row, also made at midnight, looked more like its training rows. A later slide tests this with a second model.

With New Customers In

The common set left out visits from customers a batch had never scored. Here are all 6,787 visits with any history, where each batch serves its fallback score when it has no stored score.

An interval plot against a zero line for every visit with any history, 6,787. Daily, 33 new: -0.0020 to +0.0030. Weekly, 66 new: -0.0016 to +0.0050. Monthly, 173 new: -0.0037 to +0.0055. Axis from -0.004 to +0.008, batch AP minus online AP, 95% interval. Below: for a visit it had no score for, the batch served the model's score on an empty row, 0.1424 for seed 0; online scored 0.8263 here. Caption: at most 173 of 6,787 visits were new, too few to move the average much.

The fallback is the model's score for an empty row, a row with all six numbers missing. For seed 0 it was 0.1424, so a new customer was put fairly low in the ranking. Online, by contrast, could score them from the orders they had already placed.

Even so, every interval still crossed zero. Online scored 0.8263 on this set, and the monthly batch 0.8272. The reason is simple arithmetic: at most 173 of the 6,787 visits were new to the monthly batch, about 2.5%. A worse score on 2.5% of the visits cannot move a mean of all of them very far.

In a business where most visitors are new, this would be very different, and the next slides about new customers come back to it.

One Visit, Four Scores

Numbers over thousands of visits can hide what happens to one customer. So the lab kept one worked example, chosen by a rule I wrote before the run. The rule: in October, among the visits where the monthly score had missed an earlier order, take the customer with the smallest id.

A timeline for customer 12362 from late September to early November, with invoices on 09/28, 10/11, 10/26, 10/28 and 11/04; the visit on 10/26 is the filled dot. Dashed marks show the monthly run, the weekly run and the daily run before the visit. A table of four ways. Online: made 10/26 13:47, age 0.0 hours, recency 14.97 days, score 0.245. Daily: made 10/26 00:00, 13.8 hours, 14.39, 0.294. Weekly: made 10/21 00:00, 133.8 hours, 9.39, 0.292. Monthly: made 10/01 00:00, 613.8 hours, 2.50, 0.181. Below: the monthly score was made on 1 October, before the visit of 11 October, so it still counted 6 purchases, not 7; the customer did buy again two days later; online and daily differ only in the two ages, recency 14.97 against 14.39 days and tenure 694.15 against 693.58, yet scored 0.245 and 0.294, a tree model scores in steps. Caption: seed 0 scores; one customer shows how; it shows nothing about how often.

Customer 12362 ordered on 28 September, 11 October and 26 October. On 26 October at 13:47 they came again.

The monthly score was made on 1 October. At that moment the last order was 2.5 days old and they had 6 purchases. It knew nothing about the order of 11 October. So it saw a customer who had just bought, which is a different picture from the online view: last order 15 days ago, 7 purchases.

The four scores went from 0.181 to 0.294. Notice the online and the daily scores. Their inputs differ only in the two ages: recency, 14.97 days against 14.39, and tenure, 694.15 days against 693.58. Yet the scores were 0.245 and 0.294. The model is a set of trees, and a tree answers in steps. A small change that crosses one of its split points can move the score a lot. A change that crosses none moves it not at all.

The customer did buy again, two days later. On this one visit the daily score was "more right" than online. That is one visit, and it says nothing about which way is better in general.

Why Did the Old Scores Cost Nothing?

This slide and the next two answer questions I asked after seeing the results. I wrote them into the lab's docstring before that part ran, and the lab labels them as post-results work.

A bar chart of visits that are in a batch's top fifth but not in online's, per month, seed 0: daily 29, weekly 52, monthly 104. Below: daily, 29 of 1,321 top-fifth places, rank agreement 0.991; weekly, 52 of 1,321, rank agreement 0.984; monthly, 104 of 1,321, rank agreement 0.966; rank agreement is Spearman's, 1 means the same order. Caption: the older the score, the more it moved, but most visits kept their place.

First question: were the stale scores really different from the fresh ones? Rank agreement says how similar two orderings are: 1 means the exact same order. Between online and daily it was 0.991, and between online and monthly 0.966. The top fifth is the 20% of visits each month with the highest scores, the ones a shop would act on first. Of 1,321 top-fifth places, the monthly batch put 104 different visits there than online did. So the old scores did move things, but most visits kept their place.

An interval plot for only the visits each stored score never saw. Daily, 826 visits: -0.0032 to +0.0158. Weekly, 1,294 visits: -0.0012 to +0.0153. Monthly, 2,530 visits: -0.0008 to +0.0115. Axis from -0.004 to +0.016, batch AP minus online AP, 95% interval. Below: 20-seed means, daily 0.9144 against online 0.9082; weekly 0.9312 against 0.9242; monthly 0.9204 against 0.9151. Caption: batch was ahead in 20 of 20 seeds for every schedule here, a lean the next figures explain.

Second question: what about only the visits where the stored score had missed an invoice? If staleness hurt anywhere, it should hurt there. It did not. On the 2,530 such visits, the monthly batch scored 0.9204 and online 0.9151, and the interval, -0.0008 to +0.0115, crossed zero.

But look at the seeds. On these visits, the batch was below online in 0 of 20 seeds, for daily, weekly and monthly alike. Every single model preferred the stale score. This is the lean the review found, at its strongest, and it is why I trained the second model two slides on.

A Visit Says Most of It

The third post-results question: why is AP so high here, 0.83, when the features chapter's test AP was 0.5450? Not because the model got better. AP depends on how many rows are positive: a random order scores about the share of buyers.

An isometric drawing of two blocks of different heights. Visits: 48.2% buy again. 1st of month: 19.7% buy again. Below: block height is the share who bought again within 30 days, among visits, and among every customer on the 1st of a test month in lesson 1; a random order scores about that share as AP, so 0.83 here is an easier task, not a better model; past purchases alone, no model, gave AP 0.8152; the model online 0.8308. Caption: most of the signal is the visit itself, which every way gets for free.

On the 1st of a month, the features chapter scored every customer the shop had ever seen, and only 19.7% of them bought in the next 30 days. Here, every row is a customer who is ordering right now. Among them, 48.2% bought again within 30 days. A customer who is placing an order is, almost by definition, an active customer.

Then I tried a rule with no model at all: rank the visits by how many orders the customer had placed before. That alone gave AP 0.8152. The model online gave 0.8308. So most of the ranking comes from two facts every way already has: this person is visiting, and how often they bought before. The number of past orders changes slowly, so a month-old copy of it is nearly as good as a fresh one.

So 0.83 here and 0.5450 there measure different tasks. A random order would score about 0.48 here and about 0.20 there. The visit task is easier, and the model is the same.

This is my best explanation of why staleness cost little here. It fits the numbers, but I have not proved it.

A Second Model, Trained on Visits

This slide answers the independent review. Its design was written into the lab's docstring before it ran, and the lab labels it as post-review work.

A bar chart of the share of rows whose last invoice was less than a day old. Trained on 1st of month: 1.7%. Trained at visits: 18.5%. Online test rows: 13.9%. Below: trained on 1st of month, 1.7% under 1 day, median 61.5 days; trained at visits, 18.5%, median 15.0 days; online test rows, 13.9%, median 23.0 days. Caption: online rows looked unlike the rows the first model learned from; a batch row, made at midnight, looked more like them.

The first model learned from 44,521 rows made at midnight on the 1st of each month. Only 1.7% of them had a last invoice under a day old. So I built a second training set the same way as the test visits. It holds every invoice in the training months, March 2010 to February 2011, with the six numbers worked out from everything strictly before it. Its label is "buys again within 30 days after it". That gave 19,404 rows.

In them, 18.5% had a last invoice under a day old, close to the online test rows' 13.9%. I trained 20 models on these rows and served the same test visits the same four ways.

An interval plot against a zero line for the model trained on 19,404 visit-time rows, 20 seeds. Daily, 0.8345: -0.0012 to +0.0026. Weekly, 0.8333: -0.0029 to +0.0019. Monthly, 0.8332: -0.0037 to +0.0031. Axis from -0.004 to +0.008, batch AP minus online AP, 95% interval. Below: same 6,613 visits; online scored 0.8338; batch below online in 4, 11, 16 of 20 seeds, daily, weekly, monthly; with the first model it was 3, 2, 4. Caption: the lean toward batch went away; the verdict did not change.

The lean went away. With the visit-trained model, the monthly batch was below online in 16 of 20 seeds, where the first model had it below in only 4. Weekly was below in 11. So the review was right: part of the old edge for stored scores came from the training rows.

The verdict did not change. Every interval still crossed zero: monthly -0.0037 to +0.0031, weekly -0.0029 to +0.0019, daily -0.0012 to +0.0026. With either model, no schedule could be told apart from online on this shop.

I also checked the visits each stored score missed. With the visit-trained model, the daily and weekly intervals crossed zero. The monthly one did not: the monthly score was ahead, +0.0002 to +0.0077, in all 20 seeds. That is one interval out of nine I computed with this model, and it barely clears zero. I do not know why, and I do not read it as "stale is better". With nine tries, one this close to zero can come from luck.

Most Stored Scores Were Never Read

Staleness was not the cost. Waste was. A batch scores every customer the shop has ever seen, every time it runs. Online scores only the customers who show up.

A hand-drawn set of three grids of 100 small cells each, for daily, weekly and monthly, with some cells marked as read: daily about 1 read, weekly about 6 read, monthly about 19 read. Below: daily, 658,422 scores made, 5,918 read by a visit, 99.1% never read; weekly, 90,589 scores made, 5,223 read, 94.2% never read; monthly, 21,129 scores made, 4,084 read, 80.7% never read; a score counts as read if that customer visited before the next run. Caption: the more often a batch runs, the more of its work nobody reads.

I counted a stored score as read if that customer made at least one visit before the next run replaced it. The daily batch made 658,422 scores over 123 nights, and only 5,918 of them were ever read: 99.1% were never used. The monthly batch wasted less, 80.7%, because a month is long enough for more customers to come back.

A bar chart of scores made per day: daily 5,353, weekly 761, monthly 172, online 60. Below: daily 5,353, weekly 761, monthly 172, online 60; online makes one score per visit; a batch scores every customer seen so far. Caption: a nightly batch made 89 times as many scores as online.

Per day, the nightly batch made 5,353 scores and online made 60, one per visit. These are counts, not times. I did not measure how long a score takes on this laptop. The laptop was busy with other work, and the chapter's rule is that timings come only from a quiet machine. For this small model a few thousand scores a night may cost very little. For a large model, or millions of customers, 99% waste is real money.

Who the Batch Had No Score For

The other cost of a batch is the customers it has not met yet.

Three panels. Daily batch: 33 visits by 25 customers it had not scored yet. Weekly batch: 66 visits by 51 customers. Monthly batch: 173 visits by 127 customers. Below: online could score all of them from the history they already had; 616 more visits were a customer's very first, and no way had any history for those. Caption: a first visit has no history for anyone, every way needs a rule for it.

A new customer here is someone whose first invoice came after the last batch run, and who then came back before the next one. The daily batch missed 33 such visits, by 25 customers: people with two invoices on their first day. The monthly batch missed 173 visits, by 127 customers: anyone whose first invoice came earlier in the month. Online scored all of them from the invoices they already had.

There is a bigger group that no way could help: 616 visits were a customer's very first invoice. Nobody had any history for them, online included. A shop that wants a score at a first visit needs a rule for it, such as the fallback, whichever way it serves.

In this shop, new customers were a small share of the visits. In a shop that grows fast, or a website where most visitors are new, the batch would have no score for a large share of requests. Then the choice is no longer close.

What Online Scoring Asks of You

The lab made online scoring look easy, and I should be honest about why. In the lab, online scoring had a perfect record of every past order, ready at the exact second of each visit. A real shop does not get that for free.

The numbers must be ready at request time. To score a visit at 13:47, the service needs the customer's six numbers as they are at 13:47. Someone has to keep a store of those numbers up to date as orders arrive. In the features chapter, late events and watermarks showed that events can arrive late, so the "fresh" numbers may not be quite fresh.

Two copies of the feature code must agree. A batch job usually computes features one way, over a whole table. An online service computes them another way, one customer at a time. In online and offline consistency, two correct versions of the same features still disagreed in the last digits on 17,670 of 26,851 test rows. Here, the lab avoided that by using one pass for both, and checked it against the chapter's own function.

The service must be up all day. A batch job can fail at 02:00 and run again at 03:00, and nobody notices. An online service that fails at 13:47 fails a real visit.

None of these costs showed up in my AP numbers, because the lab did not have to pay them. The rest of this chapter measures some of them on a quiet machine. It asks where the time in one request goes, what the slowest requests look like, and what a request should get when a part of the system is slow.

A Slow Question and a Fast One

This result should not surprise you if you read the two earlier lessons together. They measured the two ends, and this lesson sits between them.

Four rows. Features chapter, lesson 3, the same shop, the same 30-day question: features 14 days old, -0.0026 to +0.0048 AP; 30 days old, -0.0088 to +0.0009, lower in all 20 seeds. This lesson, 1st-of-month model, the same shop, scored at each visit: a monthly score, up to 30.7 days old, -0.0027 to +0.0063 AP, lower in 4 of 20 seeds. This lesson, visit-trained model, the same visits: the monthly score, -0.0037 to +0.0031 AP, lower in 16 of 20 seeds. Data engineering, lesson 1, bike rentals in the next hour: counts from last night, off by 56.4 rides an hour; fresh every hour, 30.5; every 2 hours, 33.8. Caption: every interval crosses zero for the shop, yet the leans differ by model; the bike task is another task; its numbers are not this shop's.

Feature freshness served this shop's model features up to 14 days old and could not measure any cost. At 30 days it was closer: all 20 seeds scored below fresh, on average by 0.0037, though the interval, -0.0088 to +0.0009, still crossed zero. That lean is the opposite of this lesson's first model, which put the monthly score ahead in 16 of 20 seeds.

The review explains the difference. In the freshness lesson, the model was trained and tested on 1st-of-month rows, so stale and fresh rows were both judged by a model that knew their shape. Here, the first model had never learned what a visit-time row looks like. Trained on visit rows, the monthly score fell below online in 16 of 20 seeds, the same lean as the freshness lesson, and again too small to separate from luck.

Batch vs streaming data asked a fast question: how many bikes will be rented in the next hour? There, the count from last night was off by 56.4 rides an hour on average, and the count made fresh every hour by 30.5. Freshness almost halved the error.

So the honest answer to "batch or online?" depends on how fast your answer changes. For a fraud check on a card payment, a news feed, or anything that changes within hours, an old score can be badly wrong. I have not measured those tasks here, so I give no numbers for them.

My Guesses Before the Run, Checked

I wrote four guesses into the lab before it ran. Here they are against the results.

  1. "daily costs nothing the bootstrap can tell from zero on the common set; weekly costs little; monthly costs clearly (its interval below zero)." Right for daily. Wrong for weekly and monthly: neither could be told apart from online. With the first model both were a little ahead on average; with the visit-trained model both were a little behind.

  2. "most batch predictions are never used: more than 90% for daily." Right: 99.1%.

  3. "online makes far fewer predictions per day than daily batch." Right: 60 against 5,353.

  4. "new customers: few under daily (only a second visit on the first day), many under monthly." Right in direction: 33 visits against 173. But "many" was too strong. Even 173 was only about 2.5% of the visits with history.

The one I got most wrong is the one I was most sure of. I had looked at the share of visits the monthly score missed, 38.3%, and decided it must cost AP. It did not, because the missed orders changed the ranking less than I thought, as the post-results slides show. This is why the lab measures before the lesson claims.

Try It Yourself

The full lab trains 20 models and redraws the customers thousands of times. I wrote a smaller demo that does the heart of it with one model.

A page in four labelled zones, headed bo_demo.py, designed before it ran. The model: lesson 1's model, seed 0; it stops unless test AP is 0.5450. One pass over the events: six running numbers per customer; copied for everyone at each midnight, the batch, read for one customer at each visit, online. For each way: age of the served score, AP on the visits all four could score, scores made per day, never used, and visits it had no score for. Left out on purpose: the 20 seeds and the bootstrap, those stay in the lab. Caption: it printed online 0.8335, daily 0.8351, weekly 0.8357, monthly 0.8327, seed 0.

I wrote the demo's design into its docstring after the lab had run and before the demo first ran. It trains the model with seed 0 and makes one pass over the shop's orders in time order. For each customer it keeps six running numbers. At every midnight it copies them for every customer seen so far: that copy is the batch. At every visit it reads that one customer's numbers as they are right then: that is online. Weekly and monthly batches reuse the midnight copies, because they also run at midnight.

A real screenshot of VS Code with bo_demo.py open at the top of the file, showing its docstring: what it needs, how to run it, and the design written before it first ran.

Before you run this lab. You need Python 3 with pandas, pyarrow and scikit-learn: pip install pandas pyarrow scikit-learn openpyxl. First run python fetch_data.py from the scripts/labs/features folder. It downloads the shop data once, about 46 MB, and writes one cleaned file. Then, inside the examples folder of scripts/labs/serving, run python bo_demo.py. It needs no GPU. I ran it with scikit-learn 1.9.1 and Python 3.13 on a Mac. Another version of scikit-learn may give slightly different decimals. The pattern should hold: no batch far from online, and most batch scores never read.

Pick a Schedule, Then Try Your Own Shop

This box holds the real results from the lab. It needs nothing but Python, so it runs in your browser. It has no model inside; it looks up what the lab measured.

Press Run. Then change SCHEDULE to "daily" or "weekly" and run it again. At the bottom, the box also does some plain arithmetic with numbers you choose. Put in your own number of customers, visitors per day and days between batch runs. It shows at least how many batch scores would never be read. That part is arithmetic, not a measurement.

The report script writes this box from the lab's stored results, runs it for all three schedules, and checks what it prints against the lab. For the arithmetic at the bottom, "at least" is honest: if the same customers visit on several days, even fewer scores are read.

The Lab's Code, Piece by Piece

The lab is one file, scripts/labs/serving/batch_or_online.py. Here is what each part does.

build_requests finds every distinct customer and order time between 1 July and 31 October 2011, and works out each label: does the same customer buy again after this moment and within 30 days?

snapshots is the one pass over the events in time order. It keeps six running numbers per customer and copies them at every batch run and at every visit. kahan_add adds money up exactly the way pandas adds a group, so the numbers match the features chapter to the last digit; features_from_state turns the running numbers into the six features with the same date arithmetic.

main trains the 20 models and checks each against the stored test AP, checks the features at 124 midnights and 300 random visits against the features chapter's own function, then serves every visit four ways. served_run finds the last batch run before each visit. ap_months works out AP per month, and bootstrap redraws customers within each month 1,000 times and keeps the 20-seed mean difference each time.

visit holds the post-review arm: it builds visit-time training rows from March 2010 to February 2011, trains 20 models on them, and serves the same test visits four ways again. holds the three questions asked after the results: rank agreement, the missed visits only, and the no-model rule. downloads each source at a fixed version and searches for every quote. only counts visits per day for the calendar figure; it was also added after the run.

How I Would Decide, From This Lab

Here is the order of questions I would ask before choosing, using what this lab and the two earlier ones measured.

A flowchart. You need a score per customer leads to a diamond: does the answer change within hours of an event? Yes leads to: score online, or run the batch often. No leads to a diamond: are many requests from customers the batch never saw? Yes leads to: online, or batch plus an online fallback. No leads to: a batch is enough; measure how stale it may be. Below: here the answer was a 30-day question, so even a score 30.7 days old could not be told apart from a fresh one; 173 visits had no monthly score at all. Caption: measure the staleness cost on your own task before you pay for online serving.

The chart asks the slow-or-fast question first, because on this shop it decided everything: a 30-day answer made even old scores good enough. Here are the same steps in words, each with the number behind it.

  1. Ask how fast the answer changes. If one event can change the right answer within hours, a stored score goes bad fast. Bike rentals in the next hour were like that. This shop's 30-day question was not.

  2. Count who arrives without a stored score. Here it was at most 173 of 6,787 visits for the monthly batch. If it is a large share for you, a batch alone will serve many guesses.

  3. Replay real requests against stored scores of each age. That is what this lab did. It is cheap, it needs only your logs and your model, and it tells you the cost of staleness on your task instead of a guess.

  4. Count the waste. If 99% of batch scores are never read and each score is expensive, online may save money even when staleness costs nothing.

  5. Remember the mix. Many real systems do both: a batch fills a store, as Feast's materialize_incremental does, and an online path covers the new customers.

When Each Way Helps, and When It Does Not

Use a batch when the answer changes slowly. Here, on a 30-day question, scores up to 30.7 days old could not be told apart from fresh ones. A batch is simple, can run when the machines are quiet, and does not need a service that is ready all day.

Do not use a batch when the answer changes within hours. The bike lesson measured that a count from last night nearly doubled the error against a fresh one. A stored score would carry the same problem.

Use online when many requests come from customers the batch has never seen. A batch can only serve a fallback for them. Here that was a small share; on a fast-growing site it may be most of the traffic.

Use online when most stored scores would never be read and each score is costly. The nightly batch here wasted 99.1% of its scores. For a small model that may not matter; for a large one it does.

Do not choose online just because it sounds more modern. It needs a service that works every hour, features that can be computed fast at request time, and the same feature code in training and serving. On this shop, with either model, all of that would buy no measurable gain in ranking.

Do not trust one seed or one customer. Seed 0 alone put the monthly batch below online; 20 seeds put it slightly above; the bootstrap says neither is a real difference.

What This Lab Cannot Tell You

Two columns titled shows and cannot show. Shows: one shop, one 30-day question, four months of real visits; one small tree model, 20 seeds for each training set; two models, trained on the 1st of each month, and on visits. Cannot show: fast questions, fraud, feeds, anything that changes within the hour; speed or cost in money, and models other than these two; a real shop's traffic, here a visit is an invoice.

One shop, one slow question. Everything here is one UK gift wholesaler and one 30-day question. I make no claim about fast questions, and I give no numbers for them.

A visit is an invoice. Real websites have visits that end without an invoice. This data only has invoices, orders and returns, so "a request" here is an invoice; about 1 in 6 was a return. A shop that scores every page view would see many more requests, and many more customers with no history.

Two models, not every model. The first model was trained on 1st-of-month rows, which leaned the race toward batch. The second, trained on visit-time rows, removed that lean. Other models, or more features that change by the hour, could behave differently.

The bootstrap is a little optimistic. It redraws customers within each month, so a customer's visits in different months are treated as separate. They are not, so the true intervals may be slightly wider than shown.

No timings and no money. Nothing was timed on this laptop, by the chapter's rule. The next lessons measure time on a quiet machine.

Labelled additions. The three post-results questions, the visits-per-day count and the fact-check mode were added after the main run. The visit-trained model was added after an independent review. The lab labels each one.

What to Do on Monday

A hand-drawn grid of six cards, titled five habits. 1, replay real requests: take real request times from your logs and serve them stored scores of each age. 2, score the same rows: compare batch and online on the requests both can score, then add the rest. 3, use seeds and a bootstrap: a gap smaller than the spread between seeds is not a finding. 4, count the waste: how many stored scores does a request ever read? 5, count the new ones: who arrives with no stored score, and what do you serve them? The reason: here a score 30.7 days old cost nothing measurable, and 99.1% of nightly scores were never read. Caption: the answer depends on how fast your question moves, so measure it.

If someone on your team asks for real-time scores on Monday, do not argue from feelings. Take a week of real request times from your logs. For each request, work out what a nightly, weekly or monthly batch would have served, and what an online score would have been. Compare them on the same requests, with several seeds and a bootstrap. Count how many stored scores were ever read, and how many requests had no stored score.

That replay takes your existing model and your existing logs, and it answers the question for your task instead of for mine.

A closing card titled staleness, waste and new customers. Staleness: a monthly score, -0.0027 to +0.0063 AP against online; on a 30-day question, nothing measurable. Waste: a nightly batch made 5,353 scores a day; 99.1% were never read; online made 60. New customers: 33 visits had no nightly score; 173 had no monthly one.

The one idea to keep: "batch or online" is three questions, not one. How much does an old score cost? Here, on a slow question, nothing I could measure. How much batch work is wasted? Here, 99.1% of the nightly scores. Who arrives with no score? Here, a small share. Measure all three on your own task before you choose.

Knowledge Check

Knowledge Check

4 questions - Score 80% to pass

Q1

On the 6,613 visits every way could score, what did the monthly batch's stored score cost against online?

Q2

Of the scores the nightly batch made over 123 nights, what share was ever read by a visit?

Q3

Why did the 173 visits with no monthly score barely change the result on all 6,787 visits?

Q4

What does this lesson, read with the bike lesson, say about when freshness matters?

"""Score every customer in a nightly batch, or each visit as it comes? Measure both on real visits.

Lesson 1 of 'Serving and Inference Basics'. It needs Python 3 with pandas, pyarrow and scikit-learn, and the
shop data from the features chapter: run scripts/labs/features/fetch_data.py once first. Then, inside this
folder:
    python bo_demo.py            # print the results
    python bo_demo.py out.json   # and save them
It times nothing and writes nothing to disk except out.json if you ask for it.

Design, written 2026-10-02 after the lab (batch_or_online.py) had run and before this file first ran:
  1. Train lesson 1's model (seed 0; it must score test AP 0.5450 on lesson 1's test months).
  2. A request is a real visit: one customer's invoice at one moment, July to October 2011. Its label is 1 if
     that customer buys again (a new purchase invoice, after this moment) within 30 days.
  3. One pass over the events in time order keeps six running numbers per customer. At 00:00 each day it copies
     them for every customer seen so far: that is the nightly batch. At each visit it reads that customer's
     numbers as they are at that moment: that is online.
  4. A weekly batch is every 7th night from 1 July, and a monthly batch is the 1st of each month; both reuse the
     nightly copies, because they run at midnight too.
  5. For each way: how stale the served score was, its AP on the visits every way could score (seed 0 only; the
     lab uses 20 seeds and a bootstrap), how many predictions it made per day, how many were never used, and how
     many visits came from customers it had not scored yet.

Author: Roni Das
Created: 2026-10-02
"""
import json
import sys
from pathlib import Path

import numpy as np
import pandas as pd
from sklearn.metrics import average_precision_score

HERE = Path(__file__).resolve().parent
sys.path.insert(0, str(HERE.parents[1] / "features"))
import task  # noqa: E402
from what_a_feature_is import HAND_COLS, hgb, joined  # noqa: E402

START, END = pd.Timestamp("2011-07-01"), pd.Timestamp("2011-11-01")
NIGHTS = pd.date_range(START, "2011-10-31", freq="D")
SCHEDULES = {"daily": NIGHTS, "weekly": NIGHTS[::7], "monthly": NIGHTS[NIGHTS.day == 1]}

# 1. the model
ev = task.load_events()
lab_tr, _, lab_te = task.splits(ev)
tr = joined(ev, lab_tr, task.TRAIN_CUTOFFS)
te = joined(ev, lab_te, task.TEST_CUTOFFS)
model = hgb(0).fit(tr[HAND_COLS].to_numpy(float), tr["label"])
p_te = model.predict_proba(te[HAND_COLS].to_numpy(float))[:, 1]
ap_te = np.mean([average_precision_score(te["label"][te["cutoff"] == c], p_te[te["cutoff"] == c])
                 for c in task.TEST_CUTOFFS])
assert round(ap_te, 4) == 0.5450, ap_te
print(f"lesson 1's model: test AP {ap_te:.4f}")

# 2. the visits and their labels
w = ev[(ev["ts"] >= START) & (ev["ts"] < END)]
req = w[["customer_id", "ts"]].drop_duplicates().sort_values(["ts", "customer_id"]).reset_index(drop=True)
buys = ev[~ev["is_return"]]
buy_times = {c: g.to_numpy() for c, g in buys.groupby("customer_id")["ts"]}
labels = []
for c, t in zip(req["customer_id"], req["ts"]):
    a = buy_times.get(c, np.array([], dtype="datetime64[us]"))
    later = a[(a > np.datetime64(t)) & (a < np.datetime64(t + pd.Timedelta(days=30)))]
    labels.append(int(len(later) > 0))
y = np.array(labels)
month = req["ts"].dt.month.to_numpy()
print(f"{len(req):,} visits by {req['customer_id'].nunique():,} customers, July to October 2011; "
      f"{y.mean():.1%} bought again within 30 days")

# 3. one pass over the events: running numbers per customer
state = {}   # customer -> [first, last, lines, return lines, money, compensation, invoices, products]


def add(row):
    s = state.setdefault(row.customer_id, [row.ts, row.ts, 0, 0, 0.0, 0.0, set(), set()])
    s[1] = row.ts
    s[2] += 1
    if row.is_return:
        s[3] += 1
    else:
        s[6].add(row.invoice)
        s[7].add(row.stock_code)
    yk = row.amount - s[5]          # money is added up the way pandas adds a group (Kahan's method),
    tk = s[4] + yk                  # so every number matches lesson 1 to the last digit
    s[5] = (tk - s[4]) - yk
    s[4] = tk


def six(s, t):
    """The six features of lesson 1, as of time t."""
    return [(pd.Timestamp(t).as_unit("ns") - s[1]).total_seconds() / 86400, len(s[6]), s[4], s[3] / s[2],
            (pd.Timestamp(t).as_unit("ns") - s[0]).total_seconds() / 86400, len(s[7])]


nightly = {}       # night -> {customer: six features}
online = [None] * len(req)
events = ev[["customer_id", "ts", "invoice", "stock_code", "is_return", "amount"]].itertuples(index=False)
row = next(events)
moments = sorted([(t, 0, i) for i, t in enumerate(NIGHTS)] + [(t, 1, i) for i, t in enumerate(req["ts"])])
for t, kind, i in moments:
    while row is not None and row.ts < t:
        add(row)
        row = next(events, None)
    if kind == 0:
        nightly[t] = {c: six(s, t) for c, s in state.items()}
    elif req.at[i, "customer_id"] in state:
        online[i] = six(state[req.at[i, "customer_id"]], t)

# 4 and 5. every way of serving
first_seen = ev.groupby("customer_id")["ts"].min()
served, rows = {}, {"online": online}
for name, nights in SCHEDULES.items():
    k = np.searchsorted(nights.values, req["ts"].values, side="right") - 1
    served[name] = nights[k]
    rows[name] = [nightly[n].get(c) for n, c in zip(served[name], req["customer_id"])]
known = np.array([all(rows[n][i] is not None for n in SCHEDULES) for i in range(len(req))])
out = {"test_ap": ap_te, "visits": len(req), "common": int(known.sum()), "ways": {}}
print(f"{known.sum():,} visits from customers every batch had scored: the fair comparison\n")
print(f"{'way':8s} {'stale, median':>14s} {'AP, seed 0':>11s} {'scored per day':>15s} "
      f"{'never used':>11s} {'not scored yet':>15s}")
for name in ["online", *SCHEDULES]:
    x = np.array([r if r is not None else [np.nan] * 6 for r in rows[name]], float)
    p = model.predict_proba(x)[:, 1]
    ap = np.mean([average_precision_score(y[known & (month == m)], p[known & (month == m)]) for m in (7, 8, 9, 10)])
    if name == "online":
        stale, per_day, unused, new = 0.0, len(req) / len(NIGHTS), 0.0, 0
    else:
        stale = float(np.median((req["ts"] - served[name]).dt.total_seconds()[known] / 3600))
        nights = SCHEDULES[name]
        ends = list(nights[1:]) + [nights[-1] + (nights[1] - nights[0]) if name != "monthly" else END]
        made = used = days = 0
        for n, e in zip(nights, ends):
            if e > END:
                continue
            visitors = set(req["customer_id"][(req["ts"] >= n) & (req["ts"] < e)])
            made += len(nightly[n])
            used += len(visitors & set(nightly[n]))
            days += (e - n).days
        per_day, unused = made / days, 1 - used / made
        new = int(sum(rows[name][i] is None and first_seen[c] < t
                      for i, (c, t) in enumerate(zip(req["customer_id"], req["ts"]))))
    out["ways"][name] = {"ap_seed0": ap, "stale_median_hours": stale, "per_day": per_day,
                         "share_unused": unused, "new_visits": new}
    print(f"{name:8s} {stale:11.1f} h {ap:11.4f} {per_day:15,.0f} {unused:11.1%} {new:15,}")
first_visits = int((req["ts"].values == first_seen[req["customer_id"]].values).sum())
print(f"\n{first_visits:,} visits were a customer's very first: no way had any history for them.")
out["first_visits"] = first_visits
if len(sys.argv) > 1:
    json.dump(out, open(sys.argv[1], "w"), indent=1)

This is a real run in VS Code's terminal, inside the examples folder.

A real screenshot of VS Code's terminal after running python bo_demo.py. It prints lesson 1's model, test AP 0.5450; 7,403 visits by 2,771 customers, July to October 2011, 48.2% bought again within 30 days; 6,613 visits from customers every batch had scored, the fair comparison. Then a table with columns way, stale median, AP seed 0, scored per day, never used and not scored yet: online 0.0 hours, 0.8335, 60, 0.0%, 0; daily 12.9 hours, 0.8351, 5,353, 99.1%, 33; weekly 108.2 hours, 0.8357, 761, 94.2%, 66; monthly 394.3 hours, 0.8327, 172, 80.7%, 173. Then: 616 visits were a customer's very first, no way had any history for them.

When I ran it, every number equalled the lab's seed-0 numbers to the last digit: the four APs, the ages, the counts. Notice that for seed 0 alone, the monthly batch scored a little below online, 0.8327 against 0.8335, while over 20 seeds it was a little above. That is why one seed is never enough. The report script checks the demo's stored run against the lab, and its exact output is in results/bo-demo-run.txt.

after
factcheck
--days