Why Production Breaks

A Production Readiness Check: Everything This Chapter Broke, in the Order You Would Check It

0 of 23 complete

0%

Contents

Back|Why Production BreaksA Production Readiness Check: Everything This Chapter Broke, in the Order You Would Check It
1/23
65 min left
Prerequisites
When the Data Answers Back: A Feedback Loop, Simulated on Real Demandrequired
Related Topics
Monitoring and Drift Detection: Catching a Model That Fails Without an ErrorCore ConceptsHandling Imbalanced and Messy Data: Why Your 99% Accuracy Is a LieData Engineering for MLDrift Alarms Against Real Harm: When the Alarm Rings, Is the Model Worse?Data Engineering for MLA Fine-Tune on New Cases: Messages Shaped Unlike Its Training, and a Gap No Model Can FixFine-Tuning: The Narrow DecisionComparing Models Fairly: Same Test, Different ResultsHow Models Generate
1 of 23

A List by the Door

Think of a man who travels often for work. On the inside of his front door there is a short list, written by hand: passport, charger, the key to the office, a printed ticket, medicine. It is not a list anyone gave him. He wrote each line after a trip went wrong. The passport line went up after he reached the airport without it. The charger line went up after three days with a dead phone in another city. The medicine line went up after a long night in a hotel room.

Look at what makes the list useful. It is short. It is in the order he needs it, so he can run through it with his hand on the door handle. And every line has a story behind it, so when he is in a hurry and tempted to skip one, he remembers why it is there.

An illustration of a man in glasses and a light blue shirt, standing with a brown bag over his shoulder, beside text. Headed a list by the door, written from past trips, titled every item is there because something once went wrong. Beside him: one model, bike rentals per hour, six things that broke it in this chapter. Tested on hours picked at random: off by 26.09. On the future: 44.07. Hours shifted by 5 at serving: 215.68, and no input check fired more than on a clean day. A loop scored on its own records: 2.6. On the riders who came: 90.7.

A list like that also has limits, and it helps to say them out loud. It only covers the trips he has already taken. It says nothing about a visa for a country he has never been to. A traveller who trusts the list completely will one day be caught by something that is not on it.

This lesson builds that kind of list for a program that learns from examples and is about to be used for real. Every line comes from something that went wrong earlier in this chapter, measured on real data. The last part of the lesson is about the list's limits: the things this chapter never tested.

Where This Lesson Starts

This is the last lesson of the chapter. Over six lessons I took one program, which guesses how many bikes a city rents in an hour, and broke it in six different ways, one lesson at a time, always on the same public data: two years of hourly bike rentals from Washington, D.C.

Lesson 1 showed that the usual way of testing the program gave it a much kinder score than the future did. Lesson 2 let the world change under it: many more people rode in the second year. Lesson 3 computed one of its inputs a little differently at the moment it was used. Lesson 4 tested simple checks that could notice a broken input before the real answers came in. Lesson 5 asked whether watching how the inputs move could stand in for knowing how wrong the program is. Lesson 6 let the program's own decisions shape the data it learned from next.

Each lesson ended with advice. This lesson puts that advice in one place, in the order a team would actually use it, and next to each piece of advice it puts the number that earned it. Nothing new is measured here. Every number is read back from a file an earlier lesson stored, and checked again against the raw results of that lesson before it is shown.

Eight Words for This Lesson

A hand-drawn list headed eight words for this lesson, titled checking a model before, and after, it goes live. Readiness check: a short list a team runs before a model goes live, and again while it runs. Item: one line of that list: a question, and what to measure to answer it. Evidence: the stored result, from an earlier lesson, that says why the item is there. Layer: one kind of check, cheapest first: inputs, then the shape of the day, then labels. Break: one thing wrong in serving, such as the hour shifted by 5. Label: the real answer for one case, here the bikes really rented in an hour. False alarm: a check firing on a day when nothing was broken. Coverage: which layer showed which break, read from the stored results. Beneath: a check you never ran on a known-good day is a guess about its false alarms.

A model is a program that learned from examples to guess a number; here, the bikes rented in one hour. It is served when it is used for real, and the code that prepares its inputs at that moment is the serving code. A readiness check is a short list a team runs before a model goes live and keeps running while it is used. Each line is an item: a question, plus what to measure to answer it. The evidence for an item is the stored result, from an earlier lesson, that says why it is on the list.

A break is one thing that goes wrong at serving time, such as the hour of the day being sent 5 hours late. A layer is one kind of check that could notice a break, and the layers come in order of cost: checks on the inputs, then a check on the shape of a whole day, then the real answers. The real answer for one case is its label. A false alarm is a check firing when nothing was broken. Coverage, in this lesson, means which layer showed which break.

Two scores from the earlier lessons come back on almost every slide. The MAE (mean absolute error) is the average size of a miss, in rentals an hour. The signed error is the same average with the minus sign kept, so below zero means the guesses lean too low.

Six Items, in the Order a Team Runs Them

A checklist is only useful if it is in the order you meet the problems. Here is this chapter's list in that order.

A flowchart headed the readiness check, in the order a team would run it, titled before launch, at serving, and once the answers come back. Six boxes in a column, joined by arrows: 1 score on the future, not a shuffle; 2 score the newest labelled month; read its sign; 3 one code path for inputs; pin units and clocks; 4 cheap input checks in front; know what they miss; 5 drift numbers are prompts; labels are the alarm; 6 watch for loops; offer a margin; log the offer. Beneath: items 1 and 3 are done before launch. Item 4, the drift numbers of item 5 and item 6's ran-out share need no labels. Item 2, item 5's label sample and item 6's true error need the real answers.

Item 1 comes before anything else, because it decides whether the number you ship with means anything: test the model on the future, not on a shuffle. Item 2 is the habit you keep once it is live: when a stretch of real answers arrives, score the model on it and read which way it leans. Item 3 is about the code that feeds the model: training and serving must compute each input the same way. It is a property of the code, settled before launch, not a check that runs every day. Item 4 puts cheap checks in front of the model, on every batch of inputs, before any answer can arrive. Item 5 is about the time between a guess and its real answer: numbers that describe how the inputs moved are a reason to look, not an alarm, and a small random sample of real answers is the better alarm. Item 6 asks whether the model's own decisions limit what can be written down, because then its records can hide its error.

A two-column list headed six items, each backed by a stored result, titled the checklist, and the number behind each item. Each row gives an item's title and its evidence with the lesson it came from, from item 1, score on the future, not a shuffle: random split 26.09, forward split 44.07 (x1.69) (lesson 1), to item 6, watch for loops; offer a margin; log the offer: loop 1.0 90.7, control 43.6; on its own records 2.6 (lesson 6). The rows hold the same lines the student script prints, shown on the try-it slide. Beneath: every number is loaded from the stored results by wrap_report.py.

The list also says who does what. Item 1 belongs to whoever builds the model, and it is done before anyone else sees a number. Item 3 belongs to whoever writes the serving code, and it is easiest to do while that code is being written, not after. Items 2, 4 and 5 belong to whoever looks after the model once it is live: they are routines, run every batch or every month, and they only work if someone reads what they print. Item 6 belongs to whoever decides what the model's answer is used for, because only they know whether the answer limits what gets written down. A team that gives every item to one person usually finds that the items nobody owns are the ones that stop being run. That last point is my experience, not something this chapter measured.

What This Lesson Ran, and What It Did Not

This lesson trains no model and runs no model. I wrote one report for it, wrap_report.py, and it only reads files that the six earlier labs stored. For each item it reads that lesson's own result file, and where it can, it checks it against the raw rows the lesson kept.

Many of the checks are real recomputations, not a second copy of the same number; the rest compare a lab's own file with that lesson's report file. For lesson 1 the report computes the forward error again from all 3,476 stored guesses. For lesson 2 it computes the error over all 8,734 hours of 2012 from every stored guess, for each of the three policies (one of which never retrains). For lesson 3 it computes each condition's error and the share of guesses that moved from the stored guesses. For lesson 4 it compares every share in the table of checks with lesson 4's own report. For lesson 5 it computes all eight rank correlations again from the monthly numbers, with its own code. For lesson 6 it computes the loop's error on its own records, and the share of hours when every bike went, from the hour-by-hour file.

In all, the report makes 123 of these comparisons before it prints anything, and it stops if any one of them fails. They all agree.

It also follows every correction the lessons made after their reviews. The break that shifts every hour by 5 is called that, "hours shifted by 5", and not UTC, because lesson 3's follow-up found that a fixed +5 matches real UTC only for the last part of the test; with the correct offset the error was 199.49. Where a lesson's number came from a follow-up designed after its main results, this lesson says so too.

What came after all the results, and is mine alone: the order of the items, the wording of each item, which numbers stand for each one, and the coverage table later in the lesson. The report writes all of them from the stored files, but the choices behind them were made by reading the results.

Item 1: Score on the Future, Not a Shuffle

The first question on the list is about the number you ship with. Lesson 1 found that the usual recipe, a random split, made the model look much better than it would ever be in use.

A bar chart headed item 1: boosted trees, MAE in rentals an hour, lesson 1, titled a shuffle said 26.09; the future said 44.07. Five bars on a scale up to 50, one per way of holding back the 20% test: random about 26, whole days about 28, whole weeks about 28, whole months about 35, forward about 44. Beneath: random 26.09, whole days 28.14, whole weeks 28.41, whole months 34.95, forward 44.07. The gap, 17.97: whole days explain 2.04, whole months 8.86, and 9.11 is the future itself. Rolling folds: 36.6 to 81.8.

On hours picked at random, the boosted trees missed by 26.09 rentals an hour. That is the mean of five random splits, each shuffled from a different seed (a starting number for the shuffle, so the same seed always gives the same split); the five ran from 25.60 to 26.74. On the last 20% of the hours in time, the future the model would really meet, they missed by 44.07: 1.69 times as much, a gap of 17.97.

A follow-up in lesson 1, designed after the main results and prompted by a review, measured where that gap came from. It held back whole days, weeks or months at random instead of single hours. Whole days gave 28.14, only 2.04 of the 17.97 points: the neighbouring hours in training explain little. Whole weeks gave 28.41. Whole months gave 34.95, 8.86 points, about half the gap: the random test had seen other hours of each test month while it learned. The last 9.11 points, from 34.95 to 44.07, went with the future itself, which was also busier than the past.

The forward score also depends on which stretch of the future you test. Five rolling folds, each trained on everything before its own block, scored between 36.6 and 81.8.

So the item is: hold the test back by time, roll it forward a few times, and report the spread. The question to ask of your own model: was the test held back by time, the way the model will really be used? In lesson 1's code the change was one line, slicing the rows in time order instead of shuffling them. If the rows have a time and the model will only meet later rows, a random split describes a job the model will never do.

Item 2: Score the Newest Labelled Month, and Read the Sign

Once the model is live, the next item is a habit. Lesson 2 trained the model on 2011 and used it through all of 2012, while the city's riders grew a lot.

A line chart headed item 2: MAE each month of 2012, lesson 2, titled the frozen model's newest month said 'too low' from the start. Three lines over January to December on a scale from -120 to 120 rentals an hour, with a dashed line at 0 labelled no lean. The frozen model's MAE starts near 69 in January, rises to about 111 in March, stays between about 78 and 110 through October and falls to about 62 in December. The frozen model's signed error, dashed, mirrors it below zero, from about -68 in January down to about -110 in March and up to about -54 in December. The retrained monthly model's MAE starts at the same point in January, drops to about 29 in February and stays between about 35 and 57. Beneath: read on 1 February: January's MAE 68.9, signed -67.8, 53% of the month's average. Over 2012: frozen 89.5, monthly 43.7. Frozen signed error below zero in all 12 months.

Lesson 2 compared three policies, rules for when to retrain. The model that was never retrained, the frozen one, missed 2012 by 89.5 rentals an hour, with a signed error of -85.0: almost all of its error was guessing too low, on 85.5% of the hours. Retrained every month on everything so far, the model missed by 43.7. Retrained every month on only the last 12 months, a window, it missed by 47.9. And the frozen model's problem was visible on the first day it could be: on 1 February, January's real counts gave an MAE of 68.9, which was 53% of the month's average, and a signed error of -67.8, so the misses were almost all too low.

Three panels headed item 2, lesson 2's follow-up: MAE from February to December 2012, titled retraining helped only through an input that could show the change. Retrained monthly, with the year input: 41.38, signed -8.1. Retrained monthly, year made constant: 88.89, signed -84.0. Never retrained: 91.41, the same with or without the year: it never used it. Beneath: designed after the main results, prompted by a review of lesson 2. One year, one run per setting.

The follow-up is the part of this item people skip. Lesson 2's review asked why retraining helped, and a follow-up, designed after the main results, made the year column constant so that no tree could use it. From February to December, retraining monthly then scored 88.89, against 91.41 for the frozen model, and it kept the lean (signed -84.0). With the year input it had scored 41.38. So retraining caught up here only because one input could say "this row is from the new, busier year".

Item 3: One Code Path for Inputs; Pin Units and Clocks

The third item is about the code that feeds the model. Lesson 3 kept the model and the test hours fixed and changed only how one input was computed at serving time, one input at a time.

A bar chart headed item 3: one serving input computed another way, lesson 3, titled from 44.07 to 215.68, depending on one input. Seven bars on a scale up to 250, with a dashed line at 44.07 labelled as trained: as trained about 44, Fahrenheit about 72, humidity in percent about 63, hour +5 about 216, weekday about 49, label about 44, wind 0 about 44. Beneath: as trained 44.07, Fahr. 71.84, hum. % 63.32, hour +5 215.68, weekday 48.86, label 43.67, wind 0 43.69. Hour +5 is the hours shifted by 5; with the correct UTC offset, 199.49.

With the inputs computed as in training, the model missed by 44.07. With every hour shifted by 5, it missed by 215.68, almost five times as much. UTC is the world's reference time; Washington was 4 hours behind it until 4 November 2012 and 5 hours after, and with that correct offset the error was 199.49. Temperatures sent in Fahrenheit gave 71.84, and it was the only break that produced impossible answers: 54 guesses below zero bikes. Humidity sent as a percentage gave 63.32 and pushed the guesses much lower (signed -42.5). The weekday counted from Monday instead of Sunday gave 48.86. In none of these cases did the model complain. It answered every time.

An isometric drawing of six blocks, heights to scale, headed item 3: share of guesses that moved by more than 10 bikes, titled two breaks moved guesses and left the average where it was. From left to right: hours shifted by 5, 98.0%, MAE 215.68; Fahrenheit, 79.2%, MAE 71.84; humidity in %, 63.7%, MAE 63.32; weekday from Monday, 40.4%, MAE 48.86; label capitalised, 22.0%, MAE 43.67; wind filled with 0, 7.8%, MAE 43.69. Beneath: as trained: MAE 44.07. The capitalised label moved 22.0% of guesses and the missing wind 7.8%; their averages stayed at 43.67 and 43.69.

Two of the six breaks are the reason this item also says the average error cannot be the alarm. A weather label spelled with a capital letter moved 22.0% of the guesses by more than 10 bikes, and a missing wind speed filled with 0 moved 7.8%, yet their average errors were 43.67 and 43.69, next to the 44.07 as trained. Lesson 3 found why for the label: every hour was treated as clear, rainy hours got worse and misty hours better, and the two nearly cancelled.

The item: compute each input with one shared function, used in training and in serving; write every unit, time zone and allowed word into a schema, a written description of each input; and never judge a serving change by the average error alone. Lesson 3 measured the damage only. It did not test any fix, so this half of the item is common practice, not something these numbers prove.

Item 4: Cheap Input Checks in Front, and Know What They Miss

The fourth item is about the time before the real answers arrive. Lesson 4 served the model 145 test days, broke one input at a time in seven ways (lesson 3's six plus a stuck temperature sensor), and ran six simple checks on every day.

A two-column table headed item 4: share of the 145 test days each check fired, lesson 4, titled four cheap checks caught the break each was built for. Left, the break; right, the checks that fired more than on clean days. Fahrenheit: range 1.00. Humidity as %: range 1.00. Hours shifted by 5: none: every check as on clean days. Weekday from Monday: none: every check as on clean days. Label capitalised: unseen 1.00. Wind filled with 0: zeros 1.00. Temperature stuck: stuck 1.00. Beneath, left: clean days: zeros 0.08 (12 of 145), psi 1.00, output 0.01. Beneath, right: psi fired on 145 of 145 clean days, broken or not.

Four cheap checks each caught the break they were built for, on every day: the range check (a number outside what training ever had) caught Fahrenheit and humidity as a percentage; the unseen check (a word training never had) caught the capitalised label; the stuck check caught the stuck temperature; and the zeros check caught the wind filled with 0. Together they caught 5 of the 7 breaks, and their only false alarms were 12 calm days out of 145 for the zeros check. The worst break, the hours shifted by 5, and the weekday counted from Monday got past all four: their values stayed inside the normal range, so no check that looks at one column at a time could see them.

Three panels headed item 4, lesson 4's two psi follow-ups; panels: random training hours, 2,000 draws each, titled at day size, psi fires from batch size alone; a bigger batch still needs a like-for-like reference. 24 hours: 82.55% of draws with temperature psi over 0.25. 168 hours: 0.05%. 720 hours: 0.00%. Beneath: clean days scored at least 3.74, against a median of 0.82 for 24 random hours. Against all training, clean weeks fired 20 of 20 and months 5 of 5; against the same month of 2011, 2 of 5 months.

The fifth check failed in a way worth remembering. The psi (population stability index: one number for how far the mix of values in a column moved), with its usual line of 0.25, fired on all 145 clean days. Lesson 4 ran two follow-ups, both designed after the main results. The first made the batches bigger and changed the reference. Against all the training hours, clean weeks still fired on 20 of 20 and clean months on 5 of 5. Against the same month of 2011, clean months fired on 2 of 5, because the weather really did differ from a year before.

Item 5: Drift Numbers Are Prompts; Labels Are the Alarm

The fifth item is about what to watch while the real answers are late. Lesson 5 computed four measures of how the inputs and outputs moved, month by month through 2012, and ranked each against the model's real error.

A bar chart headed item 5: rank correlation of each drift measure with the monthly MAE, lesson 5, titled no drift measure moved in step with the error; best 0.58. Four pairs of bars, frozen model then monthly model, on a scale from -1.0 to 1.0 labelled rank correlation with MAE, growing up or down from a dashed line at 0 labelled no relation: psi_year about -0.34 and 0.05; psi_same about 0.41 and 0.42; weather_mix about 0.39 and 0.06; pred_shift about 0.58 and 0.19. Beneath: frozen: -0.34, +0.41, +0.39, +0.58. Monthly: +0.05, +0.42, +0.06, +0.19. Below 0: months with more drift tended to have less error. 12 months: a pattern, not a proof.

A rank correlation (the kind called Spearman's) says whether two lists rise and fall in the same order: 1 is the same order, 0 no same-order relation, and a value below 0 means they tend to move in opposite orders. The best of the four measures, for the frozen model, reached 0.58. The usual psi, against the whole training year, reached -0.34: its months with more drift tended to be months with less error. None moved in step with the error. Lesson 5 found why: over the year, the weather of 2012 looked much like 2011's, while the riders grew by 1.41 to 2.53 times, month by month. The error came from the riders, which no weather input carries.

An editorial page in five labelled zones, headed item 5: the frozen model and its labels, lesson 5, titled the drift numbers said something differs; the labels said which way. psi against the whole year: 1.34 to 4.94 in the 12 months of 2012, and 1.53 to 5.10 in the 2011 months the model learned from: it measures the season. psi against the same month: over 0.25 in 9 of 12 months with the bands cut from all of 2011, 7 of 12 cut lesson 4's way; February, the retrained model's best month (MAE 28.7), crossed it. January's labels: MAE 68.9, 53% of the month's average; signed -67.8: too low. The first day of each month: scored on its labels alone, the frozen model was below zero in 11 of 12 months. 24 labelled hours a month: drawn from across each month, 1,000 times: too low in every draw, every month. All 24 exist only at the month's end. Beneath: both label readings were chosen after the results; a daily labelling routine was not tested.

The drift numbers were also noisy as alarms. The whole-year psi sat between 1.34 and 4.94 every month of 2012, and between 1.53 and 5.10 in the 2011 months the model had learned from: it was measuring the season. The same-month psi crossed 0.25 in 9 of 12 months, including February, which was the retrained model's best month of the year. And that count depends on how the bands are cut: cut from each month of 2011 itself, the way lesson 4 did, it crossed the line in 7 of 12 months instead of 9. A line whose verdict moves with a choice like that is a weak alarm.

Item 6: Watch for Loops; Offer a Margin; Log the Offer

The last item is about a model whose answer changes what can happen next. Lesson 6 was a simulation: the hourly counts were real and played the demand, but the rule that placed bikes from the forecast was invented. Riders could only rent bikes that were there, and each month's model learned from what was rented.

A bar chart headed item 6: a simulation on real demand, mean monthly MAE over 2012, lesson 6, titled on its own records 2.6; on the riders who came 90.7. Four pairs of bars on a scale up to 100, against true demand then against its own records: control about 44 and 44; loop 1.0 about 91 and 3; loop 1.2 about 60 and 31; loop 1.5 about 46 and 40. Beneath: true demand: control 43.6, loop 1.0 90.7, loop 1.2 60.0, loop 1.5 46.4. Loop 1.0 turned away 769,320 riders; every bike went in 86.1% of its hours. The supply rule is invented.

With no margin (loop 1.0), the model missed the riders who really came by 90.7 bikes an hour, against 43.6 for a control that learned from every rider, and it turned away 769,320 riders over the year. Scored on its own records, the only error an operator could compute, it missed by 2.6. Every record was true; the riders who found an empty rack were never written down. Bikes placed at 20% above the forecast brought the error to 60.0, and at 50% to 46.4, at a cost of 820,304 idle bike-hours. That last number is a unit of this simulation, which counts a bike again every hour it stands unused.

One number was visible to the operator at once, with no labels and no knowledge of demand, and it pointed straight at the problem: in 86.1% of loop 1.0's hours, every bike placed was rented. Only the error against true demand needed the real counts.

The item: ask whether the model's answer limits what can be recorded; if it does, log what was offered as well as what was taken, count how often the offer ran out, keep a margin, and never judge the model only on its own records. Lesson 6 tested margins only. Making a small share of decisions without the model is common advice with the same aim, and it was not tested here. The same shape appears in recommendations, fraud rules, credit and stock ordering, but those were general knowledge in lesson 6, not measured.

Where Each Item Sits in a Day of Service

A list is easier to use when you know where each line happens. This sequence places the items in one day of a model in service.

A sequence diagram with four columns: serving code, input checks, the model, the team. Step 1, serving code sends the input checks the day's inputs (item 4). Step 2, the input checks send the team pass, or an alarm. Step 3, serving code sends the model the same inputs. Step 4, the model sends the team guesses, always. Step 5, the team: a label sample (item 5). Step 6, the team: newest month, sign (item 2). Step 7, the team: offered vs taken (item 6). Headed one day in service, and where each item sits, titled input checks before the answer; labels and records after. Beneath: items 1 and 3 come before any of this: a score measured on the future, and serving code that matches training. Step 4 never fails; the model always answers.

The serving code builds the inputs, and the cheap checks read them before anything else happens. The model always answers, broken input or not; that was the point of lesson 4's title. Then the team's work begins: a sample of labels as they can be had, the newest labelled month when it is complete, and, for a model that decides something, a record of what was offered next to what was taken. Items 1 and 3 do not appear in the day at all, because they are settled before launch.

A hand-drawn sketch headed sketched: three moments, three kinds of item, titled some items run once, some on every batch, some after the answers. Three boxes joined by arrows: before launch, items 1, 3; every batch, no labels, items 4, 5, 6; after labels, items 2, 5, 6. Under them: 26.09 vs 44.07; 5 of 7 breaks caught; 2.6 vs 90.7. Below: 5 and 6 span both: drift and ran-out now, labels later. Beneath: left: random split against forward split. Middle: lesson 4's cheap checks; loop 1.0's ran-out share (86.1%) needs no labels. Right: loop 1.0 on its own records against the riders who came.

The sketch groups the same items by when they run. Items 1 and 3 run once, before launch, and again whenever the model or its serving code is rebuilt. Item 4 runs on every batch, before any label exists, which is why its checks have to be cheap and why their false alarms matter so much. Items 5 and 6 sit on both sides. Drift numbers and the share of hours when every bike went can be read at once, with no labels. The label sample of item 5 and the true-demand error of item 6 need the real answers, and so does item 2. A team that only has the items on the right finds out about a break when the damage is already done; a team that only has the checks in the middle sees many breaks early and misses the ones that stay in range.

No Single Layer Showed Every Break

After all the results were in, I put lesson 4's breaks next to each layer that could have shown them. This table is derived from the stored files; nothing in it was measured for this lesson.

A two-column table headed after all results: which layer showed each break of lesson 4, titled no single layer showed every break. Left, the break and the input check that fired; right, the change in MAE from 44.07 and whether the shape check fired. Fahrenheit: range; +27.77; shape no. Humidity as %: range; +19.26; shape no. Hours shifted by 5: none; +171.61; shape yes. Weekday from Monday: none; +4.79; shape no. Label capitalised: unseen; -0.40; shape no. Wind filled with 0: zeros; -0.38; shape no. Temperature stuck: stuck; not stored; shape no. Beneath, left: label and wind: the input check was the only layer that showed them. Beneath, right: the shape check was designed after lesson 4 knew the break.

Read it row by row. The unit breaks were caught by the range check and also raised the real error, by 27.77 and 19.26 rentals an hour: two layers saw them. The hours shifted by 5 passed every input check, raised the real error by 171.61, and only the shape check, built after the fact, saw it before the answers came. The weekday counted from Monday passed every check, including the shape check, and showed up only in the real error, as a rise of 4.79. The capitalised label and the missing wind speed were the opposite: the real error did not rise at all (it moved by -0.40 and -0.38), and only the input checks saw them. The stuck temperature was caught by its check, and its cost is unknown, because lesson 4 did not store its guesses.

So the answer to "which layer should we build?" is: more than one. Here, the input checks alone would have missed the most costly break, the hours shifted by 5, and the weekday break; the real error alone would have missed the two that moved hundreds of guesses (22.0% and 7.8% of 3,476) without moving the average. That is my reading of seven rows from one run, not a law. What it does show is that each layer saw something the others could not.

The Bottom Line: What This Chapter Did Not Test

A checklist built from six lessons is only as wide as those lessons. Here is what they covered, and what they did not.

A two-column table headed the bottom line, titled what this chapter tested, and what it did not. Tested here: one public dataset: 17,379 hours of bike rentals; one city, Washington, D.C., over two years; boosted trees at default settings (a linear model once, in lesson 1); one number to guess, per hour: a regression; breaks made on purpose, one at a time; a feedback loop, as a simulation with an invented rule. Not tested: other domains: loans, fraud, text, images; other cities, other years, longer spans; other model types, tuned models, large language models; classification: choosing a class, not a number; real incidents, several breaks at once; a real operator's rule and riders who react. Beneath, left: mostly one run per setting, on this data. Beneath, right: read the checklist as a place to start.

One dataset, one city, two years. Every number in this chapter comes from 17,379 hours of bike rentals in Washington, D.C., in 2011 and 2012. The kind of drift here was mostly one kind: many more riders, with weather much like the year before. Data where the inputs themselves change a lot, or where the world turns around rather than growing, could rank these items differently.

One model family. Every lesson used scikit-learn's boosted trees at default settings; lesson 1 added a linear model once, as a check. Some findings depend on how trees work. A tree cannot extend what it learned past the range it saw, which is why Fahrenheit acted like the hottest training hour. A linear model would react differently, and a large language model very differently. None was tested here.

Regression only. The model guessed a number. Nothing in this chapter tested classification, where a model chooses a class (fraud or not, spam or not). Its errors are counted differently, and some items, such as reading the sign of the error, change shape there.

Breaks made on purpose, and a simulated loop. Every break was applied by me, one at a time, to every hour. Real incidents start in the middle of a day, affect some requests and not others, and arrive several at once. The loop was a simulation with an invented rule, and its riders never reacted to an empty rack.

Mostly one run per setting. Each lab ran its design once. Only lesson 1's random split used five seeds; lesson 2's follow-up cut the monthly policy's rows with three seeds, and lesson 5 drew its label samples 1,000 times. The designs were written before the runs, and every follow-up was designed after its lesson's results. So the checklist is a floor to build on: each item is here because something broke on this data. What it cannot tell you is what will break on yours.

Try It Yourself

This script prints the checklist with each item's measured evidence, or one item with the question to ask of your own model. It uses only Python's standard library, reads one small file, and runs no model.

A real screenshot of VS Code with readiness_demo.py open, showing the docstring that describes the script and how to run it, the imports, the two line-width settings, and the functions that find the file, format one item and print the checklist. The rest of the file is in the code box on this slide. Beneath: copy it from the box on the slide.

Before you run this lab. This script needs only Python 3; it installs nothing and calls no model. It reads rc-checklist.json, printed in full in the second box below: save it next to the script. The chapter's own labs used scikit-learn and pandas (pip install scikit-learn pandas), and each earlier lesson says how to run them.

"""A production readiness check, with the evidence this chapter measured for each item.

Lesson 7 of 'Why Production Breaks', made small. It reads rc-checklist.json (copy it from
the lesson and save it next to this file) and needs only Python's standard library.
No model runs:
    python readiness_demo.py        # the whole checklist, in the order a team would run it
    python readiness_demo.py 4      # one item, with the question to ask of your own model

Author: Roni Das
Created: 2026-09-29
"""
import json
import sys
from pathlib import Path

HERE = Path(__file__).resolve().parent
WIDTH = 79
HEAD = 72


def find_file() -> Path:
    for p in (HERE / "rc-checklist.json", HERE.parent / "results" / "rc-checklist.json"):
        if p.exists():
            return p
    sys.exit("save rc-checklist.json next to this file first")


def show(item: dict, ask: bool) -> list[str]:
    head = f"{item['n']} {item['title']}"
    tag = f"lesson {item['lesson']}"
    lines = [head + tag.rjust(HEAD - len(head))]
    if ask:
        lines.append(f"  ask: {item['ask']}")
    lines += [f"  {e}" for e in item["evidence"]]
    return lines


def main() -> None:
    data = json.load(open(find_file()))
    items = data["items"]
    pick = sys.argv[1:2]
    if pick:
        n = int(pick[0])
        if not 1 <= n <= len(items):
            sys.exit(f"choose an item from 1 to {len(items)}")
        items = [items[n - 1]]
    out = [data["title"], data["scope"], ""]
    for item in items:
        out += show(item, ask=bool(pick))
    out += ["", data["footer"]]
    too_long = [line for line in out if len(line) >= WIDTH + 1]
    if too_long:
        sys.exit(f"a line is longer than {WIDTH} characters: {too_long[0]}")
    print("\n".join(out))


if __name__ == "__main__":
    main()

The Lab Report

A real terminal recording headed python wrap_report.py, titled every table in this lesson, from the stored files. It opens with 123 cross-checks against the labs' own files: all agree, then prints one numbered section per item, 1 to 6, each headed by the item's title, with each lesson's table: the splits and where the gap went, the three policies over 2012, the seven serving inputs, the share of the 145 days each check fired with psi on bigger batches and the shape check, the eight rank correlations with psi_same under both band cuts and the label readings, and the four loop policies. Section 7, after all results, shows which layer showed each break. Beneath: the lab's own report. It trains no model.

The report lives in scripts/labs/prodbreaks/wrap_report.py. It reads each lesson's lab files from scripts/labs/prodbreaks/results/: split.json and split_blocked.json for lesson 1, drift.json and drift_followup.json for lesson 2, skew.json and skew_followup.json for lesson 3, checks.json, checks_followup.json and checks_psi_size.json for lesson 4, watch.json for lesson 5, and loop.json and loop_detail.json for lesson 6. It checks each one against that lesson's own report file ( to ), recomputing from the stored guesses and hourly rows wherever the lab kept them. If any of the 123 comparisons fails, it stops before printing.

Run the Checklist Yourself

This box has no model in it. It holds the six items with their evidence, lesson 4's table of how often each check fired under each break, lesson 3's error and moved share for each serving input, and the coverage table. It runs in your browser.

As it is, the box prints the checklist with its evidence, then the coverage table: for each of lesson 4's breaks, which input check fired, whether the shape check fired, and the change in the error from 44.07 (for example +171.61 for the hours shifted by 5 and -0.40 for the capitalised label). The report's box mode checks every printed evidence line and every coverage row against its own numbers.

Try item(4) to see one item with the question to ask of your own model. Try caught('weekday_shift') to confirm that no check fired more often than on clean days, and caught('wind_missing') to see the zeros check. Then describe your own set-up: gaps(forward_test=True, same_code=True) lists the items you did not name, each with what this chapter measured for it. The names it knows are forward_test, newest_month, same_code, input_checks, early_labels and loop_log.

The Code, Part by Part

Finding the file. find_file looks for rc-checklist.json next to the script, and if it is not there, in the lab's results folder. If neither exists, it stops with a message rather than guessing.

One item. show builds the lines for one item: its number and title, with the lesson it came from lined up on the right at column 72, then its evidence lines, indented by two spaces. When you ask for one item, it also prints the question to ask of your own model.

The whole list. main reads the file, keeps either all six items or the one you named on the command line, and adds the title, the scope and the closing line. Before it prints, it checks that no line is 80 characters or longer, so the output fits an ordinary terminal without wrapping; if one were, it would stop and say which.

Where the numbers come from. Nothing in the script computes a result. The evidence lines are written by wrap_report.py, which builds each one from the stored results after its 123 checks, so the demo can only print what the labs measured. In the report, each item has its own function (item1 to item6) that loads one lesson's files, recomputes what it can, and returns the numbers; checklist turns them into the six items, using one list of titles that the report's own headings use too, and coverage builds the table of which layer showed which break.

How to Run the Check on Your Own Model

A hand-sketched column of six boxes joined by arrows, headed running the check on your own model, titled one item at a time, in this order. 1, find the line that made the test set. 2, find when labels arrive; score the newest. 3, put training and serving code side by side. 4, run each check on known-good days first. 5, price a small random label sample. 6, ask: does the answer limit the record? Beneath: advice drawn from one dataset and one model family: a place to start, not a law.

Find the line that made the test set. If the rows have a time and it shuffles them, change it to a cut in time, run the score again, and roll it forward over a few later stretches. Write down the spread. That is the number you ship with.

Find out when the real answers arrive, and score the newest stretch. Write down the MAE, the signed error and the MAE as a share of the average, next to the launch numbers. Then ask what input could show the change you most expect.

Put the training code and the serving code side by side. Run the training code on a hundred real serving records and compare the inputs value by value. Look hardest at time stamps, units, day numbering, words and missing values.

Run every check on days you know were fine before anyone gets an alarm. Range, unseen words, stuck values and counts of zeros are a few lines each. Count how often each fires on clean data; a check that fires every day gets ignored.

Price a small random sample of labels. Find out what one label costs, and how many a month you can afford. Here 24 a month showed the frozen model's lean in every draw, and the first day alone showed it in 11 of 12 months.

Ask whether the model's answer limits what gets recorded. If it does, start logging what was offered, not only what was taken.

When to Use This Checklist, and When Not To

A two-column table headed grounded in this chapter's numbers, titled when to run this checklist, and when it will not be enough. Run it when: rows have a time and the model meets later ones: 26.09 vs 44.07; serving code is not the training code: up to 215.68; labels arrive late: 24 labelled hours showed the frozen model's lean in every draw; the model's answer sets what is recorded: 2.6 vs 90.7. Not enough when: the job is not a number per hour, such as a class or a text; several things break at once, or only some requests; labels never arrive at all; you have not run each check on known-good data. Beneath, left: each item needs someone who owns it. Beneath, right: the checklist is a floor, not a guarantee.

Use it when your model guesses a number from rows that arrive in time. Demand, traffic, prices, sensor readings. That is the kind of job this chapter measured, and every item has a number from it: a shuffle said 26.09 where the future said 44.07.

Use it when serving code is written separately from training code. A different team, a different language, a data provider. One line in that code moved the error to 215.68 here.

Use it when labels arrive late. Items 4 and 5 are about that gap: cheap checks on every batch, drift numbers read as prompts, and a small random sample of labels as soon as you can afford one.

Use it when the model's answer decides something. Stock, supply, a ranked list, an approval. Item 6 is the one that stops a model being judged on records its own decisions shaped.

Do not treat it as enough when the job is different. Classification, text, images or ranking were not tested here; the items still make sense, but the numbers behind them do not transfer, and each needs its own measure of error.

Do not treat it as enough when things break in messier ways. Several breaks at once, a break in only some requests, or labels that never arrive at all were not tested. And a check you have never run on known-good data is not yet a check: you do not know its false alarms.

Do not use it as a score. Passing all six items does not mean a model is safe. It means it does not fail in the six ways this chapter found.

What These Numbers Can and Cannot Tell You

A two-column page headed read before you trust these numbers, titled what these runs are, and what they are not. They are: one dataset, one city, two years; boosted trees, default settings; mostly one run per setting; designs fixed before each run; this lesson's order: chosen after; nothing new measured here. They are not: not a rule for other data; not tuned; not other models; seeds: 5 in lesson 1 only; follow-ups came after results; not a tested ranking of items; not real incidents.

Nothing new was measured. Every number in this lesson was stored by an earlier lesson and read back here, after 123 checks against the labs' own files. So every limit of those lessons applies here too, and I repeat the main ones.

One dataset, one model family, mostly one run per setting. Bike rentals in one city over two years, boosted trees at default settings. Each lab ran its design once; only lesson 1's random split used five seeds, lesson 2's follow-up used three seeds for its cut rows, and lesson 5 drew 1,000 label samples. None of the numbers is a rule for other data.

Designs before, follow-ups after. Each lab's design was written before it ran. Several of the numbers here come from follow-ups designed after the main results, often prompted by a review: the whole-month split in lesson 1, the constant year in lesson 2, the correct UTC offset in lesson 3, the bigger psi batches, the shape check and the psi batch sizes in lesson 4, and the label readings in lesson 5. I have said so where each appears.

The order and the coverage are mine. The order of the items, their wording and the coverage table were chosen after every result was known. They are a way of reading the chapter, not a test of which item matters most.

Guesses stay guesses. Where a lesson labelled a reason as a guess (why the frozen model guessed 2011's level, why the window did worse, why loop 1.0 stayed stuck), this lesson does not promote it to a finding.

What to Do Next

A hand-drawn list headed before your model goes live, titled six questions. Item 1?: was the test held back by time, the way the model will be used? Item 2?: can any input show the change you expect, as the year did here? Item 3?: does one function compute each input for training and serving? Item 4?: how often does each check fire on days you know were fine? Item 5?: can someone label a small random sample of cases each month? Item 6?: does the model's answer limit what can be recorded next? Beneath: here: 44.07 on the future, 215.68 with the hours shifted by 5, 90.7 for the loop.

Take one model that you or your team is about to ship, or already runs, and ask the six questions on the card, one at a time. You do not need to answer all of them this week. Start with the first: find the line of code that made the test set. If the rows have a time and it shuffles them, you will learn the number your users are going to see before they see it.

Keep the answers in a file, the way this chapter kept its results. The next time someone asks whether the model is ready, you can point at evidence rather than a feeling. And when something breaks that is not on the list, as it will, measure it and add a line. That is how this list was written.

This is the end of the chapter. The next chapter planned for the course steps back from a single model to the whole life of one: how a model moves from an idea to data, training, release and the watching this chapter measured.

A closing card headed to keep, titled every item on the list comes from a measured break. In large type: 44.07, 215.68, 90.7. Beneath: the future against a shuffle's 26.09; the hours shifted by 5 against 44.07; a loop that looked like 2.6 on its own records. Then: score on the future, check inputs before the answers, label a random sample, and log what you offered. Last: one dataset and one model family: a floor to build on, not a proof.

Knowledge Check

Knowledge Check

4 questions - Score 80% to pass

Q1

A team scored its model on test rows picked at random and got a small error. Which checklist item says what to do first?

Q2

In lesson 3 the capitalised weather label moved 22.0% of the guesses, but the average error stayed at 43.67. Which layer showed that break?

Q3

In lesson 2's follow-up, with the year input made constant, retraining monthly scored 88.89 against 91.41 for never retraining. What does item 2 take from that?

Q4

Which of these did this chapter NOT test, so the checklist cannot speak for it?

I chose this order after every lesson's results were known, and it is a choice, not a finding. The lessons were run in a different order, one break at a time; nothing here tested whether one item matters more than another. The order follows the life of a model: tested, shipped, fed, checked, watched, and finally allowed to act.

That gives the item two halves. Score the newest stretch of real answers as soon as it arrives, and read the sign as well as the size: a signed error almost as large as the MAE means the guesses have fallen behind the real counts, all in one direction. And before you count on retraining, ask what input lets the model tell the new world from the old one. If there is none, fresh rows may teach it very little.

The second follow-up, prompted by a review, drew random batches from the training hours themselves, where nothing had changed. At 24 hours the temperature psi crossed the line in 82.55% of the draws; at a week, 0.05%; at a month, never. So at one day, psi fires from batch size alone. But batch size does not explain how large the clean days scored: at least 3.74, against a median of 0.82 for 24 random training hours. One summer day against a year and a half of every season is not a like-for-like comparison. A bigger batch helps only with a reference from the same kind of period. The output check, on the day's average guess, fired once on a clean day and caught nothing.

The first follow-up also tried a check designed knowing which break it had to catch. It compared the day's guesses with a typical day, ordered by the time each request really arrived, as logged by the server itself: a clock the broken serving code never touched. It needs all 24 hours of a day, so it ran on the 140 of the 145 test days that were complete. It caught the shifted hours on all 140, fired on no clean test day, and fired on 1.2% of the training days, as its line was set to do. Ordered by the hour the serving code sent, the shifted days looked normal. Because the check was built after the fact, it shows that such a check can exist, not that it would find breaks nobody has seen yet.

The item: put range, unseen, stuck and zero-count checks in front of the model, run each on days you know were fine before you trust it, and keep one clock that the serving code does not compute. Then remember what passed all of them.

What did say which way the model was wrong was the labels. After the results, lesson 5 read them two ways. The first day of each month, scored on its own 24 hours, put the frozen model below zero in 11 of 12 months; that is the earliest evidence the lab has. And 24 hours drawn at random from across each month, 1,000 times, put the frozen model too low in every draw, in every month. Those 24 hours come from the whole month, so all of them exist only at its end; a team would have to collect them as a small sample through the month, and the lab did not test that routine.

The item: treat a drift number as a reason to look at what moved, compare inputs with the same season rather than the whole year, and pay for a small random sample of labels. Let the real error on labelled data, even a few cases, decide when someone acts.

This is the file it reads. wrap_report.py writes it from the six lessons' stored results, after checking them.

{
 "title": "Production readiness check: Why Production Breaks",
 "scope": "UCI Bike Sharing, hourly, 17,379 hours; boosted trees; MAE in rentals an hour",
 "items": [
  {
   "n": 1,
   "lesson": 1,
   "title": "Score on the future, not a shuffle",
   "ask": "Was the test held back by time, the way the model will be used?",
   "evidence": [
    "random split 26.09, forward split 44.07 (x1.69)",
    "whole months held out 34.95; rolling folds 36.6 to 81.8"
   ]
  },
  {
   "n": 2,
   "lesson": 2,
   "title": "Score the newest labelled month; read its sign",
   "ask": "Can any input show the change you expect, as the year did here?",
   "evidence": [
    "2012: frozen 89.5 (signed -85.0), monthly 43.7",
    "Feb to Dec, year made constant: monthly 88.89, frozen 91.41"
   ]
  },
  {
   "n": 3,
   "lesson": 3,
   "title": "One code path for inputs; pin units and clocks",
   "ask": "Does one function compute each input for training and serving?",
   "evidence": [
    "as trained 44.07; hours shifted by 5 215.68; correct UTC 199.49",
    "Fahrenheit 71.84; label 43.67 with 22.0% of guesses moved"
   ]
  },
  {
   "n": 4,
   "lesson": 4,
   "title": "Cheap input checks in front; know what they miss",
   "ask": "How often does each check fire on days you know were fine?",
   "evidence": [
    "range, unseen, stuck, zeros caught 5 of 7 breaks, every day",
    "false alarms: 12 of 145 days; hours shifted by 5 passed all four",
    "psi on one day fired on 145 of 145 clean days"
   ]
  },
  {
   "n": 5,
   "lesson": 5,
   "title": "Drift numbers are prompts; labels are the alarm",
   "ask": "Can someone label a small random sample of cases each month?",
   "evidence": [
    "frozen model: best rank correlation with its error 0.58",
    "24 labelled hours a month: frozen too low in all 1,000 draws"
   ]
  },
  {
   "n": 6,
   "lesson": 6,
   "title": "Watch for loops; offer a margin; log the offer",
   "ask": "Does the model's answer limit what can be recorded next?",
   "evidence": [
    "loop 1.0 90.7, control 43.6; on its own records 2.6",
    "margin 1.2: 60.0; 1.5: 46.4; 1.0 ran out in 86.1% of hours"
   ]
  }
 ],
 "footer": "One city, one model family, two years. A simulation in item 6."
}

This is a real run in VS Code's terminal (python readiness_demo.py).

A real screenshot of VS Code's terminal after running python readiness_demo.py. It prints the title and the scope, UCI Bike Sharing, hourly, 17,379 hours, then the six items, each with its lesson number on the right and two or three evidence lines: 1 score on the future, not a shuffle, beginning random split 26.09, forward split 44.07 (x1.69); 3 one code path for inputs; pin units and clocks, with hours shifted by 5 215.68; 5 drift numbers are prompts; labels are the alarm; 6 watch for loops; offer a margin; log the offer, beginning loop 1.0 90.7, control 43.6; on its own records 2.6. It closes with: one city, one model family, two years. A simulation in item 6.

When I ran it, the output was exactly the run the report stored in rc-demo-run.txt. The report's demo mode runs the script again, checks that its output is the stored run character for character, checks that no printed line is 80 characters or longer, and checks that every evidence line is the one the report builds from the stored results. Try python readiness_demo.py 4 to see item 4 alone, with the question to ask of your own model.

sp-report.json
lp-report.json

Its json mode writes every number to results/rc-report.json, which the figures read, and the small rc-checklist.json for the demo. The figure script checks the numbers against the labs' files again before it draws anything. The demo mode checks the student script's stored run, and the box mode writes the playground below and checks that it prints the report's numbers.

What was fixed before each lesson's run is in that lesson's lab file. What was chosen for this lesson, after every result was known: the order of the items, their wording, which numbers stand for each, and the coverage table in section 7.

Four brand cards headed the tools, with their logos, titled what ran where. scikit-learn: every model in lessons 1 to 6, the data download. pandas: the table of hours in every lab. SciPy: the rank correlation in lesson 5. Python: this lesson: the report, the demo, the box.