Ml Lifecycle

The Lifecycle as One Loop: Nine Steps, Nine Measured Failures, and What Saw Each One

0 of 28 complete

0%

Contents

Back|Ml LifecycleThe Lifecycle as One Loop: Nine Steps, Nine Measured Failures, and What Saw Each One
1/28
80 min left
Prerequisites
When the Answers Arrive Late: What a Label Delay Costs a Model, and How Late You See Its Mistakesrequired
Related Topics
A Production Readiness Check: Everything This Chapter Broke, in the Order You Would Check ItWhy Production BreaksTomorrow Is Different: A Model Frozen on 2011, Scored Through 2012Why Production BreaksTrain/Serve Skew: One Input Computed Two WaysWhy Production BreaksFailing Silently: Which Checks Catch a Broken Input Before the Answers ArriveWhy Production BreaksWatching Inputs Before the Answers Arrive: Drift Measures Against the Real ErrorWhy Production Breaks
1 of 28

The Map With the Pins

Imagine a small team that ships parcels along one regular route, through many ports and cities around the world. Every week the same kind of parcel leaves the warehouse, passes the same places, and a report comes back. At the end of a long season the team leader pins a large world map to the wall and asks everyone to stick a pin wherever something went wrong: the port that closed in a storm, the customs office that wanted a form nobody had filled in, the depot whose clock ran an hour behind, the address that no longer existed.

An illustration of four people in an office: a bearded man in a shirt with a lanyard points at pins on a large world map on the wall, while a woman holding a folder, a woman seated at a desk and a man seen from behind listen. Headed the map with the pins, titled one loop, nine steps, and a measured failure at each. Beneath: one model on electricity prices, followed through nine lessons, from its baseline to its late answers. Persistence 0.8484 beat every model. Twenty seeds gave 0.7179 to 0.7906. A day-old input: 0.8193 to 0.7489, and the answers looked normal. Last: my reading: one failure was seen by no check at the time, and two more only by checks chosen after the results.

The pins are useful because each one is a real event, with a date and a story, not a guess about what might go wrong. But the leader also notices something about the map. Some pins were found by a check the team already ran, such as the morning look at the weather reports. Other pins were found by nobody at the time. The depot's slow clock, for example, was only noticed weeks later, when the complaints came in.

So the leader does two things. She walks the route once on the map, in the order a parcel travels it, and at each pin she asks: what is this stop for, what went wrong here, what would have told us sooner, and what do we still not know about it? Then she writes one page of checks that the next season's team can keep by the door.

This lesson does the same for a program that learns from examples. The chapter followed one such program through nine lessons, and each lesson put a pin in the map. Here I walk the whole route once, in order, and then write the one page.

Where This Lesson Starts

This is the last lesson of the chapter. Over nine lessons I followed one kind of program, trained on the same public record of electricity prices, from the first question anyone should ask about it to the problems that only appear months after it is in use. Each lesson took one step of that life and measured, on the data, what goes wrong at that step.

A flowchart headed the chapter as one loop, titled nine steps, in the order a team meets them. Nine boxes in a column joined by arrows: 1 baseline: the bar to beat; 2 reproducible training; 3 selection of settings; 4 the promotion gate; 5 canary and shadow; 6 when to retrain; 7 stale pieces in the pipeline; 8 lineage and rollback; 9 answers that arrive late. A dotted arrow runs from box 9 back up to box 1, labelled fresh labels feed the baseline. Beneath: each step is one lesson of this chapter. Persistence, the best baseline, needs the previous label, so step 9 decides what step 1 can use.

The nine steps, in the order a team meets them, are the nine lessons: start with the dumbest model, same code, different model, picking the best of many, the promotion gate, canary and shadow, when to retrain, stale pieces in a pipeline, lineage and rollback and when the answers arrive late.

They form a loop, not a line. Once a program is in use, new answers come back, it is trained again, judged again and put back into use, and the same nine questions come round again. The last step also feeds the first in a very direct way on this data, which a later slide shows.

The course's lesson on training pipelines and orchestration explains in words how software runs such a loop. The chapter before this one ended with its own checklist, a production readiness check, built on bike rentals. This lesson does not repeat either. It reads this chapter's nine results back as one route.

Eight Words for This Lesson

A hand-drawn list headed eight words for this lesson, titled reading a whole chapter back as one loop. Loop: the steps a model goes through again and again; the last feeds the first. Step: one stage of the loop, and one lesson of this chapter. Measured failure: what went wrong in a lesson's lab, with its stored number. Check: a test a team runs that could see a failure. Silent failure: one whose answers look normal from outside. Follow-up: a run designed after the main results, to answer what they raised. Persistence: the rule that repeats the previous half-hour's label. Point: one hundredth of accuracy. Beneath: every number here was stored by an earlier lesson.

A model is a program that learned from examples. In this chapter it answers one question for each half-hour of the electricity market: is the price UP or DOWN compared with its average over the last 24 hours? That right answer is the half-hour's label. Accuracy is the share of half-hours a model got right, and a point is one hundredth of accuracy, so 0.75 to 0.78 is 3 points.

The loop is the set of steps a model goes through again and again, with the last step feeding the first. A step is one stage of the loop, and in this chapter each step is one lesson. A measured failure is what went wrong in that lesson's lab, with the number the lab stored. A check is a test a team can run that might see a failure. A failure is silent when the model's answers look normal from outside, so nothing makes anyone look. Whether some check saw a failure in time is a separate question, and a later slide states the rule I used for it.

A follow-up is a run I designed after seeing a lab's main results, to answer a question those results raised. I wrote each follow-up's design down before it ran, but it was still chosen knowing the results, and I say so wherever one appears. And persistence is the laziest rule of lesson 1: say whatever the previous half-hour's label was.

How I Read the Chapter Back

This lesson runs no new experiment. Every number in it was stored by one of the nine earlier labs, in files under scripts/labs/lifecycle/results/. I wrote one report for this lesson, wrap_report.py, that reads those files and prints them step by step.

A two-column table headed nine steps, read back from their files, titled what each step is for, and what went wrong there. Left, the step and what it is for; right, the measured failure. 1 baseline: the bar a model must beat; persistence 0.8484 beat every model. 2 training: get the same model again; 20 seeds: 0.7179 to 0.7906. 3 selection: pick settings; best of 200 on validation: 0.6639 on test, rank 173 of 200. 4 gate: replace the live model, or keep it; equal models: higher-wins promoted 0.43 of 200-row windows. 5 canary and shadow: live rows; a 4.09-point regression: a 20% canary saw it in 12 of 20. 6 retrain: stay fresh; the trigger's bar sat at 0.405: 11,328 half-hours, no retrain. 7 pipeline pieces: caches, saved files; a day-old input: 0.8193 to 0.7489, answers looked normal. 8 lineage: get last month's model back; seed not written down: 0.920 to 0.970 agreement, 0 exact. 9 late answers: learn and know; twelve weeks late: weekly 0.6954, first model kept 0.7077. Beneath, left: what the step is for. Beneath, right: from that lesson's stored file.

Reading numbers back is where copying mistakes creep in, so the report checks two things before it prints anything, and stops on the first mismatch. First, where a lab's own file and its lesson's report file both hold a number, the two must be equal; there are 43 such comparisons. Second, for every lesson, the report trains the lab's models again with the same data, split, settings and seeds, and at least one headline number must come back exactly, to the last digit the computer stores. There are 30 of these refit checks. For example, it trains the boosted trees with seed 0 and gets lesson 1's 0.7527 again; it plays lesson 4's coin-swap test on 20 pairs of models and gets every one of its rates again; it serves lesson 6's monthly and triggered policies again; and it retrains lesson 9's weekly model with labels twelve weeks late and gets 0.6954. All 73 checks agreed.

For each step I then answer four questions, always in the same order. What is the step for? What went wrong there, measured? What saw it, or did nothing see it? And what could the lesson not show? The first three come from the lessons. The fourth comes from each lesson's own limits slide, and I have kept the corrections those lessons made after their reviews.

What is mine, and chosen after every result was known: the order of the steps, which numbers stand for each step, the wording of each check, and a later table of which check saw which failure. Read those as my way of reading the chapter, not as a measurement.

Step 1, the Baseline: Beat the Laziest Rule First

What it is for. A trained model's score has no scale of its own. An accuracy of 0.75 is good if the laziest rule gets 0.55, and a loss if the laziest rule gets 0.85. So the first step is to score rules that learn nothing, called baselines, on the same test, and keep the best of them as the bar.

A bar chart headed step 1, lesson 1: forward test, 9,063 half-hours, titled the rule that learns nothing beat every model. Six bars on a scale from 0.50 to 0.95, with two dashed lines, persistence near 0.85 and trees, shuffled test near 0.88: persistence about 0.85, trees plus lag about 0.82, trees about 0.75, same period about 0.67, logistic about 0.65, majority about 0.55. Beneath: 0.8484, 0.8193, 0.7527, 0.6741, 0.6491, 0.5488. The same trees on a shuffled test: 0.8830. The axis starts at 0.50.

What went wrong, measured. On a forward test, the last 20% of the half-hours in time, which is the future from the model's point of view, persistence scored 0.8484. No trained model beat it. Boosted trees with the previous label as an extra input scored 0.8193, the plain trees 0.7527, logistic regression 0.6491, and the rule that always says the most common label 0.5488. The trees were not bad in general; they were worse than doing nothing clever. A follow-up, designed after the results, gave the models only the previous half-hour's market, as a strict forecast would; then the trees with the lag scored 0.8493 against 0.8484, 8 rows ahead out of 9,063, which I read as a tie. Twenty seeds of those trees, in the benchmark's framing, ran from 0.7884 to 0.8406, and none of the 20 beat persistence.

What saw it. Scoring persistence on the same forward test. Nothing else would have: the same trees tested on a shuffle of the rows scored 0.8830, and a team that tested that way would have shipped a model that loses to a one-line rule. The same lesson found where persistence fails: on exactly the 1,374 test rows where the label changed, and nowhere else.

What the lesson could not show. Whether a tuned model, or one given more lags, would beat persistence; the lab used default settings and ran each model once. And the label is sticky partly by its own definition, a price against a slow 24-hour average, so the result belongs to this data.

Step 2, Reproducible Training: Can You Get the Same Model Twice?

What it is for. A team needs to rebuild a model to debug it, to show an auditor, to roll back to it, and to compare two versions fairly. All four need a training run that gives the same model when it is run again.

A dot chart headed step 2, lesson 2: the same trees, 20 seeds, forward test, titled same code, same data: twenty different models. Twenty dots, one per seed from 0 to 19, on a scale from 0.70 to 0.86, scattered between about 0.72 and 0.79, with the highest near seeds 2 and 5 and the lowest at seed 4. An unlabelled dashed line runs through the dots near 0.75, and a dashed line near 0.85, labelled persistence, sits above every dot. Beneath: 0.7179 to 0.7906, 20 models. The dashed line through the dots: early stopping off, 0.7520 for every seed. The axis starts at 0.70.

What went wrong, measured. The same boosted trees, the same code and the same data, trained 20 times with only the seed changed (the starting number for the program's random choices), gave 20 different models. On the future they scored from 0.7179 to 0.7906, a spread of 7.3 points, and two of them disagreed on up to 19.7% of the test rows. The cause was a default setting nobody wrote: above 10,000 rows, scikit-learn's boosted trees switch on early stopping, which sets aside a random tenth of the training rows, chosen by the seed. Every run built all 100 trees, so nothing ever stopped early; the seed only changed which tenth was left out. With early stopping switched off, every seed gave one model, 0.7520.

What saw it. Running more than one seed. One run shows one number and hides the spread completely. After the results, lesson 2's report rebuilt each seed's hidden slice and showed that, in all 20 of 20 seeds, the default model gives the same answer on every test row as a model trained without early stopping on the 90% that was kept. Running the same seed 20 times gave one model, so the computer itself added no noise.

What the lesson could not show. Other kinds of model, such as neural networks, have other sources of randomness that this lab did not touch, and it never checked another library version or another machine. It also showed why one run is not enough to compare two models: 70 of the 190 pairs of seeds were more than 2 points apart with nothing changed at all.

Step 3, Selection: The Best of Many Is Partly the Wrong Months

What it is for. Before training, someone chooses settings, such as how fast the trees learn and how big they grow. Tuning tries many settings and keeps the one that scores best on a validation set, rows held back only for choosing. A separate test set is looked at once, at the end, to find out what the choice is worth.

A line chart headed step 3, lesson 3: the winner of K random settings, mean of 1,000 draws, titled the more settings I tried, the worse the winner did on the test. Two lines over K = 1, 5, 20, 50 and 200 on a scale from 0.65 to 0.73, with a dashed line near 0.692 labelled median test. Validation, solid, rises from about 0.703 to about 0.718. Test, dashed, starts just below the median line and falls to about 0.664. Beneath: K=1: 0.703 and 0.690. K=200: 0.718 and 0.664. Luck alone, measured inside the validation months: about 21% of the growth in the gap.

What went wrong, measured. I drew 200 random settings, kept the best on validation, and repeated the game 1,000 times for several sizes of search. Picking one setting at random gave 0.703 on validation and 0.690 on test. The best of 200 gave 0.718 and 0.664: its test score came 173rd of the 200. Trying more raised the score I could see and lowered the one I could not. The gap between them grew from 1.3 points with no choosing to 5.4 points at 200.

Here is the correction lesson 3 made after its review, and it matters. Most of that growth was not luck. A fair control, which split the validation months' own days into two halves, found that ordinary selection luck accounted for 0.86 of the 4.06 points of growth, about a fifth. The rest came from choosing on months that were unlike the months the model then served: the validation months, November 1997 to the first morning of June 1998, rewarded different settings from the test months, June to December. Lesson 3 also kept a limit on that control: neighbouring days are alike, so two halves made of mixed days are more alike than two truly independent samples, and the control may understate luck.

What saw it. Only the untouched test set. While choosing, nothing warned: the rank correlation between validation and test scores over all 200 settings was 0.03, almost no relation. The best setting on the test itself scored 0.750, and persistence still scored 0.848 on the same rows.

What the lesson could not show. Smarter search methods, a winner retrained on training and validation together, and data that does not drift, where luck might be most of the story.

Step 4, the Promotion Gate: When Is New Really Better?

What it is for. A pipeline that retrains a model must decide whether the new one, the challenger, replaces the live one, the champion. The rule that decides is the promotion gate, and it is only as good as what it demands.

Three panels headed step 4, lesson 4: two models equal by construction, 200-row windows, titled higher-wins promoted equal models; a paired test mostly did not. Naive, higher wins: 0.431 of windows promoted. Margin, 1 point: 0.298 of windows promoted. McNemar: 0.018 of windows promoted. Beneath: at 1,000 rows: 0.511, 0.094, 0.022. A real 6.7-point gain: every rule caught it in every 1,000-row window.

What went wrong, measured. I first planned pairs of seeds as models that are equally good. Lesson 4 found that they were not: on the whole test, the second seed of a pair was from 3.2 points worse to 7.2 points better, and 18 of the 20 pairs clearly differed. That is lesson 2 again, showing up at the gate: a retrain with a new seed is a different model. So a follow-up built models equal by construction, by tossing a coin on every row to decide which model got which outcome. On those, the rule "promote if the new score is higher" promoted the challenger in 0.43 of 200-row windows and 0.51 of 1,000-row windows: a coin flip at every size. A fixed 1-point margin promoted it in 0.30 of 200-row windows and 0.09 of 1,000-row windows, so its safety depended on the window. And checking again and again made things worse: over 45 checks as rows arrived, McNemar's test promoted 6 of the 20 equal challengers at some check, against 2 with one check at the end.

What saw it. The coin swap itself, which showed that higher-wins and a fixed margin promote equal models far too often. McNemar's test, which counts only the rows where exactly one of the two models was right and asks whether that split is too uneven for coin flips, promoted in 0.018 and 0.022 of windows. Lesson 4 is careful about what that shows, and so am I: under the swap, McNemar holding near its line is what the maths requires, because the swap builds exactly the coin-flip assumption the test rests on. The new evidence is about the other two rules. McNemar still caught a real 6.7-point gain in every 1,000-row window, though only in 0.578 of 200-row windows.

What the lesson could not show. The coin swap also broke the link between neighbouring half-hours, so it cannot say whether real windows of alike rows fool the test. And there was no pair with a small real gain, so how many rows a 1-point gain needs is unknown here.

Step 5, Canary and Shadow: Seeing a Worse Model on Live Rows

What it is for. A model that passed the gate on stored rows can still be worse on today's rows. A canary sends a share of live rows to the new model; shadow runs the new model silently on every row next to the old one. Both look for a regression, a new version that is worse.

A bar chart headed step 5, lesson 5's follow-up: injected regressions, 20 starts, titled shadow saw both regressions in every start. Two groups of four bars, canary 5%, canary 20%, canary 50% and shadow, at 1.75 points and 4.09 points, on a scale of starts with an alarm, of 20, from 0 to 20. At 1.75 points: 2, 4, 7 and 20. At 4.09 points: 6, 12, 15 and 20. Beneath: injected errors only ever point one way: the easiest case for shadow.

What went wrong, measured. First, my own design went wrong twice, and lesson 5 kept both on the record. The planned "small regression", the trees without two inputs, was in fact 0.61 points better than the champion. And the planned equal challenger was the champion itself, which can never disagree with itself, so shadow could not have raised any alarm on it. A follow-up, designed after the results, replaced both. On a known regression of 4.09 points, a canary sending 20% of the rows raised the alarm in only 12 of 20 starting points in 2,000 half-hours; at 5% in 6. On 1.75 points, the 20% canary caught 4 of 20. And checking the canary after every half-hour, instead of once at the end, raised its false alarms on equal models from 1.1% to 9.8% of runs.

What saw it. Shadow, because it compares both models on the same rows: it caught both injected regressions in all 20 starts, the 4.09-point one after a median of 241.5 half-hours. On the real regression of the main run, the logistic model, shadow raised the alarm first in every start where a 20% canary also did, 15 of 15, and alone in 3 more.

What the lesson could not show. The injected errors only ever turned right answers wrong, which is the easiest case for a paired test; the real regression came in runs, with a correlation of 0.57 between neighbouring rows against 0.01 for the injected errors. There was no real traffic, one row per half-hour, and nothing about how users react to new answers, which only a canary can see.

Step 6, the Retrain Policy: A Bar Anchored to a Bad Week

What it is for. A model learns from the past and then answers a future that drifts away from it. A retrain policy decides when to train it again: on a schedule, or when a measurement says it got worse, which is called a trigger.

A hand-drawn sketch headed sketched: step 6, lesson 6, the trigger's last model, real numbers, titled one bad first week set a bar nothing later went under. Three boxes in a row joined by arrows: retrain, 1998-04-15; first week 0.455; bar 0.405. Two boxes below: 230 checks, lowest week 0.452; no retrain for 11,328 half-hours. Beneath: the bar was the model's own first week minus 5 points. Replayed from the lab's serving loop, after the results.

What went wrong, measured. Over 22,656 served half-hours, never retraining scored 0.6610, retraining every 28 days 0.7516 with 16 retrains, and every week 0.7636 with 67. The trigger, which retrained when the last week fell 5 points below the model's own first week, scored only 0.6812 with 8 retrains. The report replayed its log and found why: the last model it trained, on 15 April 1998, had a first week of 0.455, so its bar was 0.405. In 230 daily checks its weekly accuracy never went below 0.452, so it served 11,328 half-hours, half the run, without a retrain. The rule was anchored to one bad week.

A bar chart headed step 6, lesson 6 and its follow-up: points gained over never retraining, titled most of retraining's gain came through the date input. Two pairs of bars, every 28 days and every week, on a scale from 0 to 12 points over never: with the date about 9 and about 10; date set to 0 about 2 and about 3. Beneath: with the date: +9.06 and +10.26. Date set to 0: +1.94 and +3.19. Persistence: 0.8622, above every policy.

The second finding came from a follow-up and changed how I read the first. Most of what retraining bought came through one input, date. Every served half-hour had a date larger than any the first model had learned from, and the frozen model gave the same answer on 100.0% of rows when its dates were replaced by its last training date. With the date input set to 0, monthly and weekly retraining gained only 1.94 and 3.19 points over never, instead of 9.06 and 10.26. And persistence, 0.8622 on the same rows, stayed above every policy.

What saw it. Nothing running at the time. The trigger checked every day and saw nothing wrong by its own rule; only a replay of its log after the run showed the anchored bar. A fixed floor, retraining whenever a week fell below 0.65, 0.70 or 0.75, avoided the anchoring in a follow-up (0.7578, 0.7518 and 0.7656), but I chose those levels after seeing the results.

Step 7, Stale Pieces: A Cache and an Old File

What it is for. Between the data and the model sit other parts of a pipeline: an input stored in a cache and reused instead of computed again, or a saved file such as a scaler, which rescales the inputs before the model reads them. They must stay in step with the model.

Two panels headed step 7, lesson 7: two stale pieces, the same test rows, titled the quiet one looked normal; the loud one did not. An input cached for a day: 0.7489; fresh 0.8193; said UP 0.541 against 0.484. An old scaler: 0.4512; fresh 0.6491; said UP on every row. Beneath: stuck check: the cache on 188 of 188 days. Range on the model's inputs: the scaler on 188 of 188. A 1% recompute: 33 of 79 sampled rows for the cache, 79 of 79 for the scaler.

What went wrong, measured. The previous half-hour's label, cached once a day instead of computed each half-hour, cut the trees plus lag from 0.8193 to 0.7489, a loss of 7.04 points. Lesson 7's review corrected how I first read that: 0.7489 is a tie with the trees that never had the lag at all, 0.7527 (McNemar p 0.353), not worse than them. The day-old lag simply threw away everything the lag had added. And from outside it looked normal: the share of answers that were UP moved only from 0.484 to 0.541. The second stale piece was loud. The logistic model served through a scaler left over from the first 10% of the data scored 0.4512 and said UP on every row, because the old scaler turned the growing date into a number more than 200 spreads above anything the model had seen.

What saw it. A different check for each piece, both chosen after the results. A stuck check, which fires when an input that should move holds one value all day, caught the cache on 188 of 188 test days, and fired on 1 day with the fresh input. A range check on the numbers the model actually received caught the old scaler on all 188 days, while the same check on the raw inputs could not see it at all. As first written that range check had no margin, and lesson 7 found it needs one: with the date left out and a margin of 0.1 spreads it fired on 0.037 of days, and with any sensible margin only the date tripped it, every day. The daily share of UP answers missed the old scaler completely, because on training days the logistic model had already said DOWN on every row on 110 days and UP on every row on 2. Recomputing 1% of served inputs from scratch caught both: 33 of 79 sampled rows for the cache, 79 of 79 for the scaler.

What the lesson could not show. Both faults were built on purpose, one input and one file, so this is a demonstration of what each check can see, not a fair trial of the checks.

Step 8, Lineage and Rollback: Getting Last Month's Model Back

What it is for. On a bad morning, a team goes back to the model that worked. That needs either the saved model file or a complete record of how it was made, its lineage, so it can be rebuilt exactly.

An isometric drawing of five blocks, heights to scale above 0.70, headed step 8, lesson 8: agreement with last month's answers, titled two roads gave the model back; three did not. From left to right: full record 1.000 and saved file 1.000, the two tallest; seed guessed, median 0.957; labels backfilled 0.958; rows shifted 0.739, a small cube. Beneath: seed guessed: 0 of 19 exact, the data fingerprint quiet, the answers fingerprint fired on 19 of 19. Rows shifted scored 0.9182 against the original's 0.7067.

What went wrong, measured. Lesson 8 trained one model on rows 0 to 30,000 with seed 7 and tried five ways back. The full record and the saved file both gave the same answer on all 6,000 rows it had served. The other three did not. With the seed not written down, 19 guessed seeds agreed with it on 0.920 to 0.970 of rows, and none was exact. With 282 old labels corrected upstream, a rebuild agreed on 0.958. With the rows written as "the last 30,000" and read a month later, it agreed on only 0.739, and it scored 0.9182 on the served rows against the real model's 0.7067, because it had trained on those very rows. A wrong rebuild looked better than the truth.

What saw it. Two fingerprints in the record. A hash is a short code computed from exact numbers that changes if one number changes. The hash of the training rows fired for the corrected labels and the shifted rows. It stayed quiet for the missing seed, because the rows were right. Lesson 8's correction is the important one here: the missing seed is silent only to a record that describes the inputs. The hash of the model's answers on the stored month fired on all 19 guessed seeds.

What the lesson could not show. A change of library version, which it did not test. And the "month" was simulated: every rebuild ran seconds later, on the same machine, in the same program.

Step 9, Label Delay: Learning Late and Knowing Late

What it is for. A model can only learn from answers that have arrived, and a team can only score it on them. The label delay is how long the right answer takes to arrive. Here it decides both learning and knowing.

A bar chart headed step 9, lesson 9: weekly retraining against the same run's first model, kept, titled retraining every week clearly helped only when labels came within a day. Five pairs of bars on a scale from 0.60 to 0.80, at half hour, a day, a week, 4 weeks and 12 weeks: retrained weekly and first model kept. The weekly bar is taller at the first four delays, by a wide margin at half an hour and a day and a narrow one at a week and 4 weeks; at 12 weeks the first-model bar is taller. Beneath: weekly: 0.7636, 0.7563, 0.7299, 0.7315, 0.6954. First model kept: 0.6610, 0.6667, 0.7101, 0.7152, 0.7077. Weekly minus first, week bootstrap 95%: +6.87 to +13.78; +5.59 to +12.34; -1.12 to +4.98; -1.15 to +4.28; -4.26 to +1.47 points. The axis starts at 0.60.

What went wrong, measured. Lesson 9 retrained the model every week while the labels arrived half an hour, a day, a week, four weeks or twelve weeks late. Accuracy went 0.7636, 0.7563, 0.7299, 0.7315 and 0.6954. Four weeks scoring above one week looks backwards, but the two differed by 36 rows out of 22,656 (McNemar p 0.59), and a bootstrap over whole weeks gave an interval of -2.38 to +2.72 points, so the order is noise. A day's cost, 0.73 points, also could not be told apart from chance.

Lesson 9's review changed the comparison with never retraining, and I use the corrected one. The 0.6610 "never" line from lesson 6 was a model trained with fresh labels, which a team with late labels could never have had. The fair comparison for each delay is that run's own first model, kept forever: 0.6610, 0.6667, 0.7101, 0.7152 and 0.7077 for the five delays. Weekly retraining clearly helped only when the labels arrived within a day: +10.26 and +8.96 points, with week-bootstrap intervals of +6.87 to +13.78 and +5.59 to +12.34, both far above zero. At a week and at four weeks the small gains, +1.98 and +1.62, cannot be told apart from chance (-1.12 to +4.98 and -1.15 to +4.28). Twelve weeks late it scored lower than its first model kept, 0.6954 against 0.7077, and that gap also cannot be told apart from chance (-4.26 to +1.47). Lesson 9 added this bootstrap after its review.

What saw it. Nothing at the time. With labels twelve weeks late, a dashboard of weekly accuracy had nothing to show for the first 90 days, and then showed a fine week on 65 days when the real week was bad, and a bad week on 94 days when the real one was fine. Those two counts rest on levels, 0.65 and 0.70, that were chosen after lesson 6's results. A signal that needs no labels, how far the share of UP answers sat from training's, moved partly with weekly accuracy (a rank correlation of -0.60); the model's own confidence did not (-0.08).

The Loop Closes at Step 1

The last step feeds the first, and on this data it does so in a very plain way. The best rule of the whole chapter, persistence, uses exactly one thing: the previous half-hour's label. It scored 0.8484 on lesson 1's test and 0.8622 on the served half of lessons 6 and 9, above every trained model and every retraining policy.

A sequence diagram with four columns: pipeline, model, checks, team. Headed one turn of the loop, and where each step's check sits, titled the checks run at different moments, by different people. Step 1, the pipeline to the model: train; write the record (2, 8). Step 2, the model to the checks: score against persistence (1). Step 3, the checks to the team: test once, after choosing (3). Step 4, the team to the model: gate on disagreeing rows (4). Step 5, the model to the checks: shadow on live rows (5). Step 6, the checks to the team: stuck, range, recompute (7). Step 7, a dashed arrow from the team back to the pipeline: labels in: retrain (6, 9). Beneath: numbers in brackets are the steps of the loop. The order of this turn is my reading of the chapter, chosen after the results.

That rule only exists because the labels come back in half an hour. With a delay of a day, persistence would have to repeat yesterday's label at the same half-hour, a weaker rule that lesson 1 measured at 0.6741. With a delay of twelve weeks it would be almost meaningless. So the label delay of step 9 decides which baselines step 1 is allowed to use. A team that measures its baseline with labels it will not have in service is setting a bar it cannot reach, which is a framing point lesson 1 made at its start: write down what is known when each guess is made.

The sequence above places each step's check in one turn of the loop, as I read the chapter. The pipeline trains a model and writes its record, the record that steps 2 and 8 need. The model is scored against persistence. The chosen setting is tested once. The gate compares on the rows where the two models disagree. Shadow runs on live rows. The pipeline's pieces are checked for staleness. Then labels arrive, and the loop starts again with a retrain. The order is my reading, chosen after the results; the lessons ran one step at a time.

One more thing ties the turn together. Every check in it runs at a different moment, and usually belongs to a different person: the person who trains, the person who owns the gate, the person on call when something breaks. In my experience, which this chapter did not measure, the checks that nobody owns are the ones that stop being run.

Which Check Saw Which Failure

After all nine results were in, I put each step's measured failure next to what saw it. Nothing in this table was measured for this lesson; it is read from the stored files.

A two-column table headed after all the results: my reading of the stored files, titled which check saw each failure, or that nothing did. Left, the failure; right, what saw it. 1 persistence ahead of every model: persistence on the same forward test; a shuffle said 0.8830. 2 a 7.3-point spread from the seed: more than one seed; one run shows one number. 3 the winner fell on the test: the untouched test; validation's rank correlation was 0.03. 4 equal models promoted: higher-wins failed; McNemar 0.018, as the swap requires. 5 a 4.09-point regression: shadow in 20 of 20; a 20% canary in 12. 6 the trigger's bar stuck low: only checks chosen after: a fixed floor; a replay. 7 a day-old input, normal answers: only a check chosen after: stuck input, 188 of 188 days. 8 a seed nobody wrote down: the answers fingerprint, 19 of 19; the data one stayed quiet. 9 bad weeks, twelve weeks late: nothing at the time: 65 bad days looked fine. Beneath, left: my reading, by the rule stated beside it. Beneath, right: each number is from its lesson's file.

First, the rule I used for "seen", because the count that follows depends on it, and it is my reading, not a measurement. A failure counts as seen at the time if some check in the chapter fired on it clearly using only what the pipeline had at that moment: no replay of a log after the run, and no labels that had not yet arrived. I also note whether that check was written into the lab's design before the run, or chosen by me after the results. This is a different question from whether a failure was silent, which is about whether the answers looked normal. By that rule the nine rows fall into four kinds.

In the first kind, a check from the lab's own design saw the failure clearly: persistence scored on the same forward test (step 1), shadow on the same live rows (step 5) and the fingerprint of the model's answers (step 8). None of these needs anything exotic. Each is a few lines of code or a single extra number.

In the second kind, the check was also in the design, but it costs more than a single run: running more seeds (step 2), or keeping a test set that nobody chose with (step 3). These cost more training, and they are exactly the things that get skipped when a team is in a hurry. Step 4 belongs here too, with a caveat. The coin swap showed that higher-wins and a fixed margin promote equal models far too often; McNemar holding near its line under the swap is what the maths requires, not a test of McNemar on real windows of alike rows.

In the third kind, only a check I chose after the results saw it. For the trigger of step 6, that was a fixed floor, which retrained where the self-referenced bar did not, and a replay of the trigger's log, which showed the bar stuck at 0.405. For the day-old cache of step 7, it was a stuck-input check. I count those two alike, because both checks were written knowing what they had to catch.

The Failures No Check Saw in Time

A failure that no check sees in time is the most expensive kind, because nothing makes anyone look, and it can run for months. By the rule on the previous slide, the chapter had one that nothing saw at the time, and two that only checks I chose after the results saw. A fourth is easy to miss for a different reason.

An editorial page in four labelled zones, headed my reading, by the rule on the previous slide, titled one failure no check saw in time; two seen only by checks chosen after. Step 9, seen by nothing: twelve weeks late, a dashboard had nothing to show for 90 days, then showed a fine week on 65 days when the real week was bad. Step 6, only a check chosen after: the trigger's bar was its own first week, 0.455, minus 5 points. It checked every day and never fired again; a fixed floor and a replay of the log, both after the results, showed it. Step 7, only a check chosen after: a day-old input moved the share of UP answers only from 0.484 to 0.541: silent in lesson 7's sense. A stuck-input check, written after the results, saw it every day; the daily output mix fired on 18% of days, against 5% with a fresh input. Step 2, easy to miss: one training run shows one number. The 7.3-point spread appears only when you run more seeds.

Seen only by a check chosen afterwards: the trigger. It did its job exactly as written. It measured a rolling week every day and compared it with a reference. The fault was in the reference: one bad week, 0.455, set the bar at 0.405. No alarm could ever say "the bar itself is wrong", because the bar was the thing the alarm used. What showed it was a comparison the trigger never made: a fixed floor, which lesson 6 tried in a follow-up with levels chosen after the results, and a replay of the log after the run.

Seen by nothing: the late dashboard. Every number it showed was correct about the week it described. The problem was which week that was: twelve weeks late, the newest complete week was 84 days old. A chart titled "accuracy, last 7 days" that shows a week from three months ago is not wrong about the past; it is wrong about now.

Seen only by a check chosen afterwards: the cached input. This one is also silent in lesson 7's sense: a day-old lag cost 7.04 points while the share of UP answers moved from 0.484 to 0.541, a gap that hides inside the ordinary swings of the season. The output checks fired on 18% of days against 5% with a fresh input, far too often on a healthy model to act on. Only the stuck check saw it every day, and I wrote that check knowing the result.

Easy to miss: the seed. Lesson 2's design ran 20 seeds, so the lab saw it. But one training run prints one number. The 7.3-point spread of step 2 only exists for someone who runs the same code more than once.

Three Threads That Ran Through Every Step

Some findings did not belong to one step. They came back in lesson after lesson, and they are the most useful things I learned from the chapter as a whole.

Three panels headed three threads that ran through the loop, titled the same three things kept coming back. Persistence: 0.8484; lesson 1; 0.8622 on lessons 6 and 9's rows; it needs the previous label. The date input: +9.06 to +1.94; monthly retraining's gain, with and without it; the date 211.89 spreads out under the old scaler. The seed: 0 of 19; guessed seeds gave the exact model back; 18 of 20 seed pairs clearly differed. Beneath: each thread is measured in more than one lesson; the lessons are named in the text beside this figure.

Persistence. A rule that learns nothing beat every trained model in lesson 1 (0.8484), stayed far above the best tuned setting in lesson 3 (0.848 against 0.750), stayed above every retraining policy in lesson 6 (0.8622) and needed labels half an hour old in lesson 9. On data whose labels repeat, the question "does this model beat persistence?" came before every other question, and the answer here was no.

The date input. An input that only says when did damage in two lessons and mattered in a third. In lesson 6 the frozen model treated every served half-hour as the last day of its training. Monthly retraining gained 9.06 points over never retraining with the date and 1.94 with the date set to 0, so about 7.1 of those points are tied to the date; lesson 6 calls its account of why, that retraining moved the model's idea of "latest" forward, its interpretation. In lesson 7 an old scaler turned the same date into a number more than 200 spreads out, and the model said UP on every row. In lesson 9 the date added 1.09 points with fresh labels (0.7636 against 0.7527) and almost nothing twelve weeks late (0.6954 against 0.6962); lesson 9 found that its first, unfair baseline, not the date, had decided how much late retraining appeared to help. Before measuring anything about time, find the inputs that only say when.

The seed. Randomness hidden in a default setting made 20 models from one piece of code in lesson 2, made pairs of seeds genuinely different models at lesson 4's gate (18 of 20 clearly differed), and made a lost seed impossible to recover in lesson 8 (0 exact of 19 guesses). Write it down, and never compare two models on one run each.

What the Chapter Corrected About Itself

Every lesson in this chapter was reviewed before it went live, and several of them changed what they said after a review or a follow-up. The wrap has to agree with what each lesson finally says, so here are the corrections in one place.

A table of seven rows headed what the chapter corrected about itself, in review or a follow-up, titled seven results that changed after a second look. Lesson 3: the winner's fall was mostly a mismatch of months; luck was about a fifth (0.86 of 4.06 points). Lesson 4: the seed pairs were not equal (-3.2 to +7.2 points); a coin swap built a true null. Lesson 5: the planned small regression was 0.61 points better; the equal challenger could not test shadow. Lesson 6: about 7.1 of monthly retraining's 9.06 points were tied to the date input; the why is lesson 6's reading. Lesson 7: a day-old lag tied the trees with no lag (p 0.353). Lesson 8: a missing seed is silent to the data fingerprint, not to the answers fingerprint. Lesson 9: against each run's own first model, retraining clearly helped only within a day; a week and later cannot be told apart from chance. Beneath: each correction came from a review or a follow-up designed after the results.

Lesson 3 first read its growing gap as the winner's curse, selection luck. A fairer control found luck was about a fifth of the growth, 0.86 of 4.06 points, and most of it came from choosing on months unlike the months served. Lesson 4 designed its "null" pairs as equal models, and they were not; a coin swap built a true null. Lesson 5's planned small regression was 0.61 points better, and its equal challenger could not test shadow; a follow-up replaced both. Lesson 6 found that about 7.1 of monthly retraining's 9.06 points were tied to the date input, and labelled its account of why as its interpretation. Lesson 7 first called the day-old cache worse than no lag at all; it was a tie, p 0.353. Lesson 8 first called the missing seed silent; it is silent to the data fingerprint and caught by the answers fingerprint every time. Lesson 9 first compared late retraining with a frozen model that had fresh labels; against each run's own first model, retraining clearly helped only within a day, and from a week on the difference could not be told apart from chance.

I include this slide because it is the pattern, not the exception. Each correction came from asking a result one more question, usually one a reviewer asked. Several of the lessons' clean first readings were wrong in a way the first design could not see, and the fix was always another measurement, never a better sentence. That is a reason to treat every number in this chapter, including the corrected ones, as one careful reading of one dataset.

A One-Page Checklist for Your Own Pipeline

Here is the chapter as one page, one line per step, in the loop's order. Each line is one lesson's advice, and each has a measured failure behind it. The wording is mine.

A hand-sketched column of nine boxes joined by arrows, headed a one-page checklist for your own pipeline, titled nine questions, one per step, in the loop's order. 1, beat the best no-learning rule, on a forward test. 2, train twice; compare answers row by row; run seeds. 3, test once, on rows nobody chose with; count what you tried. 4, gate on the rows where the two disagree; fix the rule first. 5, shadow on the same rows before a canary; fix when to look. 6, score a frozen copy; a fixed floor; find inputs that say when. 7, store each input's age; stuck, range with a margin, 1% recompute. 8, keep the file; rows as numbers; seed; both fingerprints. 9, measure the label delay; date every accuracy you show. Beneath: each line is one lesson's advice. The wording is mine, chosen after the results.

  1. Beat the best no-learning rule, on a forward test. Score majority, persistence and a seasonal rule on the same future rows as the model. Here persistence won, 0.8484 against 0.8193.
  2. Train twice and compare the answers row by row; then run several seeds. Here one piece of code gave 0.7179 to 0.7906.
  3. Test once, on rows nobody chose with, and count what you tried. Report the winner's test score and K. Here the best of 200 was 0.718 on validation and 0.664 on test.
  4. Gate on the rows where the two models disagree, with the rule written down first, and look once. Here higher-wins promoted equal models in 0.43 of windows, McNemar in 0.018.
  5. Shadow on the same live rows before any canary, and decide in advance when to look. Here shadow caught a 4.09-point regression in 20 of 20 starts, a 20% canary in 12.
  6. Score a frozen copy block by block, give any trigger a fixed floor, and find the inputs that only say when. Here a self-referenced trigger sat at 0.405 for half the run.
  7. Store each cached value's age and each file's model version; run a stuck check, a range check with a margin on what the model receives, and a 1% fresh recompute. Here a day-old input cost 7.04 points and looked normal.

Try It Yourself

This script runs four of the loop's cheapest checks on Elec2: the baseline against persistence (step 1), the fingerprint of the training rows and of a model's answers, with one rebuild from the recorded seed and one from a guess (steps 2 and 8), and the output mix, the share of answers that were UP, for the trees and for the logistic model behind a stale scaler (a signal from lessons 7 and 9). It trains five small models and does not need a GPU.

A real screenshot of VS Code with loop_demo.py open, showing the docstring that says what the script is and how to run it, the imports, the lines that load Elec2 and turn the label into 1 for UP and 0 for DOWN, the cut at 80%, the version line, the sha function that fingerprints arrays, and the start of the baseline check that compares the trees with persistence; the rest of it and the lineage and output-mix parts are further down. Beneath: copy it from the box on the slide.

Before you run this lab. You need Python 3 and two libraries: pip install scikit-learn pandas. scikit-learn holds the models, the scaler and the download, and brings NumPy with it; pandas holds the table; hashlib comes with Python. The first run downloads Elec2 from OpenML, a free public website of datasets for machine learning (under 1 MB compressed), so it needs an internet connection once; scikit-learn keeps a copy in a folder in your home directory (scikit_learn_data) for later runs. I ran it with scikit-learn 1.9.1 on a Mac. scikit-learn runs the same way on Windows and Linux, but I have not checked the numbers there, and another version of scikit-learn may give different models and different fingerprints, which is lesson 8's point; that is why the first line printed is the version.

"""The lifecycle as one loop: four of the chapter's cheapest checks.

Lesson 10 of 'The ML & AI Lifecycle', made small. It needs Python 3 with
scikit-learn and pandas (pip install scikit-learn pandas). The first run
downloads Elec2 from OpenML (under 1 MB) and keeps a copy for later runs.
    python loop_demo.py

Author: Roni Das
Created: 2026-09-29
"""
import hashlib

import numpy as np
import sklearn
from sklearn.datasets import fetch_openml
from sklearn.ensemble import HistGradientBoostingClassifier as Trees
from sklearn.linear_model import LogisticRegression
from sklearn.preprocessing import StandardScaler

# 45,312 half-hours in time order. The label: is the NSW price UP (1)
# or DOWN (0) against the average of the last 24 hours?
data = fetch_openml("electricity", version=1, as_frame=True,
                    parser="auto").frame
X = data.drop(columns="class").astype(float)
y = (data["class"].astype(str) == "UP").astype(int).to_numpy()
cut = int(0.8 * len(y))  # learn from the first 80%, test on the rest
print(f"scikit-learn {sklearn.__version__}, {len(y):,} half-hours")


def sha(*arrays):
    # a short fingerprint of the exact numbers in the arrays
    h = hashlib.sha256()
    for a in arrays:
        h.update(np.ascontiguousarray(a).tobytes())
    return h.hexdigest()[:12]


# Check 1 (lesson 1): does the model beat the laziest rule, on the future?
trees = Trees(random_state=0).fit(X[:cut], y[:cut])
said = trees.predict(X[cut:])
model = (said == y[cut:]).mean()
persist = (y[cut - 1:-1] == y[cut:]).mean()  # say what the last one was
print("\n1. baseline, on the last 20% of the rows")
print(f"   trees {model:.4f}   persistence {persist:.4f}")
if model > persist:
    print("   the model wins; go on to the next step")
else:
    print("   persistence wins: ship the rule, or rethink")

# Checks 2 and 3 (lessons 2 and 8): fingerprint the training rows and
# the answers, then rebuild: once with the seed, once with a guess.
rows, month = slice(0, 30000), slice(30000, 36000)
first = Trees(random_state=7).fit(X[rows], y[rows])
served = first.predict(X[month])
data_hash = sha(X[rows].to_numpy(), y[rows])
answers_hash = sha(served)
print("\n2. lineage, rows 0 to 30,000, seed 7")
print(f"   data hash {data_hash}, answers hash {answers_hash}")
for seed, name in ((7, "seed 7 again"), (3, "seed 3, a guess")):
    again = Trees(random_state=seed).fit(X[rows], y[rows])
    answers = again.predict(X[month])
    now = sha(X[rows].to_numpy(), y[rows])  # the rows it learned from
    d = "same" if now == data_hash else "CHANGED"
    a = "same" if sha(answers) == answers_hash else "DIFFER"
    print(f"   {name}: data {d}, answers {a}")
agree = (answers == served).mean()
print(f"   seed 3 agrees on {agree:.4f} of the served rows")

# Check 4 (lessons 7 and 9): the output mix, which needs no labels.
# The logistic model is served through a scaler fitted on the first 10%.
scaler = StandardScaler().fit(X[:cut])
logistic = LogisticRegression(max_iter=1000)
logistic.fit(scaler.transform(X[:cut]), y[:cut])
old = StandardScaler().fit(X[: int(0.1 * len(y))])
stale = logistic.predict(old.transform(X[cut:]))
print("\n3. output mix: share of answers that were UP")
print(f"   training labels {y[:cut].mean():.3f}")
print(f"   trees {said.mean():.3f}")
print(f"   logistic, stale scaler {stale.mean():.3f}")

The Lab Report

A real terminal recording headed python wrap_report.py, titled every number in this lesson, from the nine lessons' files and refits. It opens with: 43 cross-checks between files, 30 against refits: all agree. Then nine numbered sections, one per step: 1, baseline, persistence 0.8484 and the shuffled trees 0.8830; 2, reproducible training, 20 seeds 0.7179 to 0.7906; 3, selection, best of 200 at 0.7175 and 0.6639, rank 173 of 200, luck 0.86 (21%); 4, the promotion gate, the true null at three window sizes; 5, canary and shadow, the injected 4.09-point regression; 6, the retrain policies and the trigger's bar of 0.405; 7, stale pieces, the daily cache and the stale scaler, 1% recompute 33 and 79 of 79; 8, lineage and rollback, seed missing 0.9197 to 0.9702; 9, label delay, a table of weekly accuracy, first model kept, the difference and its week-bootstrap interval for five delays, from 30 min 0.7636, 0.6610, +10.26, +6.87 to +13.78 to 12 weeks 0.6954, 0.7077, -1.22, -4.26 to +1.47, then a line saying only half an hour and a day are clearly above zero. Section 10 lists the refits, one or more per lesson, all equal to the stored numbers. Beneath: the lesson's own report. It stops unless every file agrees and every refit gives the stored number.

The report lives in scripts/labs/lifecycle/wrap_report.py. It reads the nine labs' stored files and their reports' files, from baseline.json for lesson 1 to delay.json, delay_followup.json and delay_fairnever.json for lesson 9, and it checks them in two ways before it prints anything. Where a lab's file and its lesson's report file both hold a number, they must agree: 43 comparisons. And for every lesson, it trains the lab's models again with the lab's own data, split, settings and seeds and stops unless the stored headline comes back exactly: 30 refit checks, including lesson 4's coin-swap null on 20 pairs of models and lesson 9's weekly retrain with labels twelve weeks late, together with its first model kept. It changes nothing in any lab's files. Its tryit mode, added after this lesson's review, makes the try-it slide's edit on a copy of the demo, runs it, checks the printed agreement against its own refit, and stores the run in .

Walk the Loop Yourself

This box has no model in it. It holds the nine steps, each with what it is for, the failure its lesson measured, what saw it and what the lesson could not show, written from the report's numbers. It runs in your browser.

As it is, the box prints the loop: each step with what it is for and its measured failure, and a last line that sends you back to step 1. The report's box mode checks that every printed failure is the one it built from the stored files.

Try step(8) to see one step in full, with what saw the failure and what the lesson could not show. Try unseen() to list, by my reading and the rule on the matrix slide, the step no check saw at the time and the steps only checks chosen after the results saw: it prints steps 6, 7 and 9. Then describe your own pipeline: gaps(1, 4, 8) lists every step you did not name, each with what this chapter measured there.

The Code, Part by Part

Loading. fetch_openml("electricity", version=1) downloads Elec2 once and reads the local copy after that. The label column holds the words UP and DOWN, turned into 1 and 0. cut is 36,249: the first 80% of the rows in time are for learning, the rest are the future.

sha. Feeds the exact bytes of one or more arrays into SHA-256 and keeps the first 12 characters of the fingerprint. np.ascontiguousarray lays the numbers out in one fixed order first, so the same table always gives the same bytes. Twelve characters are enough to compare by eye; a real record keeps all 64.

The baseline. The boosted trees with seed 0 learn from the first 80% and answer the rest. y[cut - 1:-1] == y[cut:] scores persistence: each test row is paired with the label before it. The if prints the verdict, because the check is only useful if someone acts on it.

The fingerprints. The model of lesson 8 learns from rows 0 to 30,000 with seed 7 and answers the next 6,000, the rows it "served". The data hash and the answers hash are the record. The loop rebuilds it twice, with the recorded seed and with a guess, and compares both hashes with the record each time. The data hash is computed from the rows the rebuild learned from; here they are the same rows, so it says "same" both times, which is exactly why it cannot see a lost seed.

The output mix. The logistic model learns behind a scaler fitted on the training rows, then answers through an old scaler fitted on the first 10% of rows, as in lesson 7. The script prints the share of UP answers, which needs no labels, next to the share of UP in the training labels.

How to Walk the Loop on Your Own Pipeline

The checklist is the what. Here is the how, in the order I would add the checks to a pipeline that has none, cheapest and most often useful first.

Start with the two numbers that cost nothing. Score persistence, or whatever your laziest rule is, on the same future rows as your model, and print the share of each answer the model gives next to the share in its training labels. Both need a few lines and no new data. The first tells you whether the model is worth running at all. Trust the second little: in lesson 9 it moved with weekly accuracy only partly, a rank correlation of about -0.60, so it says where to look, not what is wrong. In the demo the trees said UP on 0.512 of the test rows against 0.418 of the training labels, while 0.451 of the test rows really were UP; a gap like that is a reason to look closer, not proof that anything broke.

Then make every run leave a record. Save the model file, the row numbers or dates it learned from, the seed, every setting and the library versions, and two fingerprints: of the training rows and of the model's answers on a fixed set of rows. Keep it next to the model, by version, in a registry you control, because loading a pickle can run hidden code, and do not count on a file loading under another scikit-learn version, which scikit-learn does not support. From then on, a rollback is a load and a comparison of two hashes.

Then put a paired check in front of every replacement. Keep the champion's answers on your evaluation rows, and compare any challenger on the same rows with McNemar's test, with the rule and the number of rows written down before anyone sees a score. If you can afford it, run the challenger in shadow on live rows before any user sees it.

Then look for stale pieces. List every cache and every saved file the pipeline loads. Store the time each value was computed and the model version each file belongs to, and recompute a small sample fresh every day.

Then measure time. Find out how late your labels really arrive, score a frozen copy of the model block by block, and find the inputs that only say when. Only then choose a retraining schedule, and compare it with the first model kept, not with a model that had labels you will never have.

Last, run the expensive checks when a decision depends on them. Several seeds before you believe a gain, a clean test set before you report a tuned score, a replay of any trigger's log before you trust it.

When This Loop Fits, and When It Is Not Enough

A two-column table headed grounded in this chapter, titled when this loop fits, and when it is not enough. It fits when: rows arrive in time order and the right answers come back; a model is retrained and replaced again and again; two models can answer the same rows; a pipeline keeps caches and saved files. It is not enough when: answers never arrive, or only for what users saw; the model's answers change what happens next; the model is a large neural network; several things break at once, in some requests. Beneath, left: all nine steps were measured on Elec2. Beneath, right: none of these was tested here.

Use this loop when rows arrive in time order and the right answers come back. Prices, demand, sensor readings, whether a server was busy. That is the kind of job this chapter measured, and every step has a number from it.

Use it when a model is retrained and replaced again and again. Steps 2, 4, 6 and 8 are about exactly that: a new model is a different model, and each replacement needs a record, a paired gate and a way back.

Use it when two models can answer the same rows. The two strongest checks in the chapter, McNemar's test at the gate and shadow in rollout, both depend on it.

Use it when a pipeline keeps caches and saved files. Step 7 applies to any stored piece that a later step loads.

It is not enough when the answers never arrive, or arrive only for what users saw. Then accuracy cannot be measured at all, shadow cannot be scored, and steps 5 and 9 change shape. This chapter never tested that.

It is not enough when the model's answers change what happens next. A recommendation changes what people click; the electricity price went up or down whatever the model said. The previous chapter's feedback-loop lesson measured that case, on bikes; this one did not.

It is not enough for large neural networks or language models as they stand. Their randomness, their cost of a retrain and their outputs are different, and none was tested here. The questions carry over; the numbers do not.

It is not enough when several things break at once. Every failure here was one fault at a time, often built on purpose.

The Limits of This Chapter

A two-column page headed read before you trust these numbers, titled the limits of this chapter. They are: one dataset, Elec2; scikit-learn only; mostly single runs; designs came first; many follow-ups came after; faults built on purpose; stored rows, not traffic. They are not: not a rate for other data; not other model families; not tested for significance, mostly; not changed after the runs; chosen knowing the results; not bugs found in the wild; not a live service.

One dataset. Every number in this chapter comes from Elec2, 45,312 half-hours of one electricity market in New South Wales, from May 1996 to December 1998. Its label is sticky by its own definition, a price against a slow 24-hour average, and its date input grows with time. Both shaped several results. On other data the ranking of these failures could be very different, and persistence could be weak.

scikit-learn only. Every model was a scikit-learn model: boosted trees, and logistic regression as a second model. Some findings depend on how boosted trees work, such as a date past the training range falling into the last range the trees saw, and on one default setting, early stopping. Neural networks and language models were never tested.

Single runs, mostly. Most labs ran each setting once. The exceptions were deliberate and small: 20 seeds in lesson 2 and in lesson 1's follow-up, 200 settings and 1,000 draws in lesson 3, 20 seed pairs in lesson 4, 20 starting points that overlap in lesson 5, 19 guessed seeds in lesson 8, and bootstraps in lessons 3 and 9. Few gaps in this chapter were tested for significance, and I have called gaps that could not be told apart from chance exactly that.

Several follow-ups designed after the results. Every lab's design was written before it ran. But lessons 1, 3, 4, 5, 6, 7, 8 and 9 each added a follow-up designed after the main results, and several key numbers come from them: the strict forecast, the fair luck control, the true null, the injected regressions, the date set to 0, the cache intervals, the answers fingerprint's count, and the first model kept. Lesson 2's rebuild of each seed's hidden slice, behind its 20 of 20, was also added after the results. Many of the checks in lesson 7, the levels behind lesson 9's hidden bad days, and the table of which check saw what were chosen knowing the results.

No real production traffic. Live traffic was always stored rows replayed in time order, one row per half-hour. The faults were built on purpose. The lineage "month" was simulated in one process on one machine. No user ever saw an answer.

What to Do Next

Take one model your team runs, open the checklist, and answer the nine lines in order. For each one, write down the check you run today, or "nothing". The lines where you write "nothing" are your map with the pins still to come.

Then do the three cheapest things this week. Score persistence, or your own laziest rule, on the same future rows as the model. Print the share of each answer next to the share in the training labels. And add the answers fingerprint to the record of the model you would roll back to, so that the next rebuild, or the next reload, can prove it is the same model. The demo on this lesson's try-it slide runs all three on Elec2.

A closing card headed to keep, titled walk the loop once, with a check at every step. In large type: 9 steps, 9 measured failures, 1 that nothing saw at the time. Beneath: persistence 0.8484 beat every model. A day-old input cost 7.04 points and looked normal. A missing seed passed the data check and failed the answers check 19 times in 19. Then: retraining every week clearly helped only with labels a day old or newer. Twelve weeks late it scored 0.6954 against its own first model kept, 0.7077: lower, and not beyond chance. Last: the count of 1 is my reading, by a stated rule. One dataset, scikit-learn only: a way to check your own loop, not a law.

This is the end of the chapter. Each lesson measured one step, and this one put the steps back into a loop. The next chapter planned for the course moves on from the life of one model to the next part of the curriculum, and the checks from this page go with it.

Knowledge Check

Knowledge Check

4 questions - Score 80% to pass

Q1

In lesson 8, a model was rebuilt with a guessed seed. Which check noticed that it was not last month's model?

Q2

With labels twelve weeks late, how did weekly retraining compare with that same run's first model, never retrained?

Q3

Lesson 6's triggered policy went 11,328 half-hours without a retrain. What did the replay of its log find?

Q4

A day-old cached input cut accuracy from 0.8193 to 0.7489. Which check saw it on every one of the 188 test days?

What the lesson could not show. Triggers on input drift, the cost of a retrain in time, and which level of floor is right.

What the lesson could not show. Real feedback, which arrives unevenly rather than at one fixed delay, a trigger fed with late labels, and the many drift checks it did not try.

In the fourth kind, nothing in the chapter saw it at the time: with labels twelve weeks late, step 9's dashboard showed the wrong weeks for months, and the one label-free signal moved only partly with accuracy. So, by this rule, one failure was seen by nothing, and two more only by checks chosen afterwards. The next slide looks at those three.

One caution about this table. It shows which kind of check can see which kind of failure. It is not a fair trial of the checks: several were chosen knowing the results, and lesson 8's table of which field fires was measured after them.

What the four have in common, as I read them, is that each check looked at the thing itself and not at what it was compared with. The trigger trusted its own reference, the dashboard trusted its own window, the output mix trusted its own training range, and one run trusted itself.

  • Keep the model file; write the rows as numbers, the seed and every setting; store both fingerprints. Load a file only from a registry you control, because loading a pickle can run hidden code, and treat a file saved with another scikit-learn version as unsupported. Here a guessed seed passed the data check and failed the answers check 19 times in 19.
  • Measure how late the labels arrive, compare with the first model kept, and date every accuracy you show. Here weekly retraining clearly beat its first model kept only with labels a day old or newer.
  • You do not need all nine this week. Lines 1 and 8 cost a few lines of code each, and each saw its failure clearly in this chapter. The demo on the next slide runs four cheap checks on the same data: the baseline of line 1, the two fingerprints of lines 2 and 8, and the output mix, a label-free signal from lessons 7 and 9 that is not one of line 7's checks.

    This is a real run in VS Code's terminal (python loop_demo.py).

    A real screenshot of VS Code's terminal after running python loop_demo.py. It prints scikit-learn 1.9.1, 45,312 half-hours; then 1, baseline, on the last 20% of the rows: trees 0.7527, persistence 0.8484; persistence wins: ship the rule, or rethink; then 2, lineage, rows 0 to 30,000, seed 7: data hash 7f391cd43a8a, answers hash b978599166b1; seed 7 again: data same, answers same; seed 3, a guess: data same, answers DIFFER; seed 3 agrees on 0.9575 of the served rows; then 3, output mix: share of answers that were UP: training labels 0.418, trees 0.512, logistic, stale scaler 1.000.

    When I ran it, it printed scikit-learn 1.9.1 and 45,312 half-hours; then trees 0.7527 against persistence 0.8484, and the verdict that persistence wins; then the data hash 7f391cd43a8a and the answers hash b978599166b1, the same first 12 characters as lesson 8's record; the rebuild with seed 7 gave the same data and the same answers, and the guess, seed 3, the same data and different answers, agreeing on 0.9575 of the served rows; and last the output mix: training labels 0.418, trees 0.512, and the logistic model behind the stale scaler 1.000. Every one of those numbers matches the lessons' stored files, and the longest printed line was 52 characters. The report's demo mode checks all of it.

    To see the data fingerprint fire, add one line right after answers_hash = sha(served): y = y.copy(); y[:100] = 1 - y[:100]. It corrects 100 old labels after the record was written, as an upstream fix would. The copy() matters: the labels that to_numpy() returns here are read-only, and changing them in place raises an error. Put it before the record is written and the record is built from the corrected labels too, so the data check still says same. When I ran exactly that edit, after this lesson's review, both rebuilds printed data CHANGED and answers DIFFER, seed 3 agreed on 0.9158 of the served rows, and the training-label share in the last part moved to 0.419, because the corrected labels stay in y. The report's tryit mode makes the same edit on a copy of the script, runs it, checks the agreement against its own refit, and stores the run in results/lo-tryit.json.

    results/lo-tryit.json

    Its json mode writes every number to results/lo-report.json, which the figures read, and the figure script checks the numbers against the lessons' files again before it draws anything. The demo mode checks the student script's stored run, and the box mode writes the playground below and checks it.

    What was fixed before each lesson's run is in that lesson's lab file. What was chosen for this lesson, after every result was known: the order of the steps, which numbers stand for each, the wording of the checks and the checklist, and the table of which check saw which failure.

    Four brand cards headed the tools, with their logos, titled what ran where. scikit-learn: every model in the nine labs. Python: this lesson's report: json, hashlib. NumPy: the rows, the draws, the fingerprints. pandas: the table of half-hours.

    So read the chapter as one careful reading of one dataset: a way to check your own loop, with a number behind each check, and not a set of rates that transfer to your data.