Ml Lifecycle

Stale Pieces in a Pipeline: A Cached Input, an Old Scaler, and What Each One Hides

0 of 23 complete

0%

Contents

Back|Ml LifecycleStale Pieces in a Pipeline: A Cached Input, an Old Scaler, and What Each One Hides
1/23
66 min left
Prerequisites
When to Retrain: A Schedule, a Trigger, and What Each Retrain Buysrequired
Related Topics
Tomorrow Is Different: A Model Frozen on 2011, Scored Through 2012Why Production BreaksTrain/Serve Skew: One Input Computed Two WaysWhy Production BreaksFailing Silently: Which Checks Catch a Broken Input Before the Answers ArriveWhy Production BreaksWatching Inputs Before the Answers Arrive: Drift Measures Against the Real ErrorWhy Production BreaksWhen the Data Answers Back: A Feedback Loop, Simulated on Real DemandWhy Production Breaks
1 of 23

The Board by the Door

Imagine a small school with a board by the front door. Every morning at seven, the caretaker looks out of the window and writes the day's weather on the board: "dry", or "rain". Parents read it when they drop their children off, and teachers read it at lunch to decide whether the children can play outside.

One day the caretaker writes "dry" at seven. At two in the afternoon it starts to rain hard. The board still says "dry". Nothing on the board looks wrong. It looks exactly like every other dry day, in the same handwriting, in the same place. A teacher reads it and sends thirty children out into the rain.

An illustration of a man in a sweater and trousers, standing with his hands at his sides, next to text. Headed the board by the door, titled a stale piece raises no error. Does anything look wrong? Beside him: one model, its inputs built by a pipeline. Two pieces of that pipeline went stale. An input cached once a day: accuracy 0.8193 fell to 0.7489, and it still said UP on 0.541 of rows, near the fresh 0.484. An old scaler, fitted on the first 10% of rows: 0.4512, and UP on every row. Last: the quiet one looked normal from outside. Write down when each input was computed.

Now imagine a second problem. The school's thermometer was marked years ago, in the depth of winter, and whoever marked it wrote "normal" at a very cold point. Ever since, every reading shows "far above normal", in summer and in winter alike. That one is easy to spot: a board that says "far above normal" every single day is clearly broken, and someone will ask about it within a week.

Both problems come from the same thing: a stored piece of information that nobody updated. But the first one is quiet and the second one is loud. This lesson measures both kinds in a program that learns from examples, and asks which one does more harm before anyone notices.

Where This Lesson Starts

This chapter follows one model through its life, on the same electricity data. Start with the dumbest model set the baselines. Same code, different model and picking the best of many looked at training. The promotion gate and canary and shadow looked at how a new model is judged and put live, and when to retrain asked how often to train it again.

All of those lessons treated the model as the thing that goes stale. This one looks at the pieces around it. The course's lesson on training pipelines and orchestration has a slide called "Skip the Work You Already Did: ". It explains in words how a pipeline saves the output of a step and reuses it when nothing changed. I will not repeat it. This lesson measures what happens when a saved piece is reused after something did change.

A flowchart headed inside the serving pipeline, titled where a stale piece sits. A box, a new half-hour arrives, leads to step 1: build the inputs. A cylinder, cache: refreshed once a day, feeds step 1 with a dotted line labelled lag input. Step 1 leads to step 2: rescale the inputs, which a second cylinder, saved scaler file, also feeds with a dotted line. Step 2 leads to step 3: the model says UP or DOWN. Beneath: in the lab the cache fed the trees and the old scaler fed the logistic model: one stale piece in each pipeline.

The previous chapter measured two close relatives. Train/serve skew computed one input two different ways, and failing silently tested which simple checks catch a broken input before the right answers arrive. Here the input is computed the right way, by the right code. It is just old. I reuse that lesson's checks later on, on these new inputs.

Nine Words for This Lesson

A hand-drawn list headed nine words for this lesson, titled the pieces between the data and the model. Pipeline: the chain of steps that turns a new record into an answer. Step: one link of the chain: build an input, rescale, predict. Artifact: a file a step saved, which later steps load: a fitted scaler, a model. Cache: a stored result reused instead of computing it again. Refresh: computing the stored result again: here once a day, or once a week. Stale: a stored piece that no longer matches what a fresh computation gives. Lag input: the previous half-hour's label, given to the model as an input. Silent: a failure whose answers look normal from outside. Loud: a failure whose answers look strange from outside. Beneath: a stale piece raises no error. The model answers either way.

A pipeline is the chain of steps that turns a new record into an answer. Each link is a step: build the inputs, rescale them, ask the model. Many steps save what they made so that later steps, or later runs, can load it. A saved file like that is called an artifact. A fitted scaler is an artifact, and so is the trained model itself.

A cache is a stored result that is reused instead of being computed again. Computing it again is a refresh. A cache saves time and money, and it is safe as long as the stored result is still what a fresh computation would give. When it is not, the stored piece is stale.

This lab caches one input, the lag input: the label of the previous half-hour, which lesson 1 added to the trees to make its best model. The label is the right answer for a row: here, whether the price went UP or DOWN against its average over the last 24 hours. Accuracy is the share of rows a model got right, and a point is one hundredth of accuracy.

Last, two words for how a failure looks from outside. A failure is silent when the model's answers look normal: the same mix of UP and DOWN, nothing strange in any chart of the outputs. It is loud when the answers themselves look strange.

How a Cached Input Goes Stale

The lag input changes every half-hour. A fresh pipeline looks up the previous half-hour's label each time a new row arrives. A cached pipeline looks it up once, stores it, and serves the stored value until the next refresh. Here is a real day from the lab's test, sketched, with the cache filled once a day just before midnight.

A hand-drawn sketch headed sketched: a real day, 1998-06-09, every third hour, titled the cached input held the night's value all day. Six rows of eight boxes, for the clock times 00, 03, 06, 09, 12, 15, 18 and 21. Real label: D, D, D, D, then U, U, U, U. Fresh input: D, D, D, D, U, U, U, U. Cached input: D in all eight, the last four marked as different. Fresh answer: D, D, D, D, U, U, U, U, all right. Cached answer: D, D, D, D, D, D, U, D, with three marked wrong. Beneath the boxes: U = UP, D = DOWN. The cache was filled just before 00:00. Beneath the sketch: the first whole test day where the fresh answers were right at all 8 of these half-hours and the cached ones wrong at 3 or more, a rule chosen after the results; here, 3. The whole day: fresh 0.875, cached 0.729.

On this day the price was DOWN through the night and morning and turned UP around midday. The fresh input followed it: once the price turned, the previous half-hour's label was UP, and the model said UP. The cached input still held the label of the last half-hour of the day before, which was DOWN. It went on saying DOWN all afternoon, and the model, which leans heavily on this input, followed it for most of the afternoon. Over the whole day the fresh input scored 0.875 and the cached one 0.729.

Nothing in the pipeline raised an error. The cached value was a perfectly normal label, 0 or 1, of the same type and in the same place as the fresh one. The only thing wrong with it was its age.

Why would anyone cache an input like this at all? Because computing inputs fresh has a cost. In a real system the previous label might live in another database, owned by another team, and looking it up for every request adds time to every answer and load on that database. A popular pattern is to compute inputs in a batch job, once a night, and store them where the serving code can read them quickly. That design is sensible for inputs that change slowly. The trouble starts when the same pattern is used for an input that changes every half-hour, and nobody measures what the delay costs. That is exactly the choice this lab makes on purpose, so its cost can be counted. I picked this day with a rule I wrote after seeing the results, so it shows the mechanism clearly; it is not a typical day, and the averages come on the next slides.

What the Lab Ran

I wrote the lab's design at the top of its file, stale_lab.py, before it ran. The dated record is in the chapter plan (2026-09-29, "STALE LAB (batch 7) designed before running").

An editorial page in four labelled zones, headed what the lab ran: stale_lab.py, designed before it ran, titled two trained models, five ways to serve them. The data: Elec2, as in lessons 1 to 6: train on the first 36,249 half-hours, test on the last 9,063, 1998-06-01 to 1998-12-06. 45.1% of test rows were UP. Model 1: boosted trees with the lag input, lesson 1's trees plus lag: 0.8193 with fresh inputs. Model 2: logistic regression behind a scaler, lesson 1's logistic: 0.6491 with its own scaler. What changed: only the piece in front of each model at serving: where the lag input came from, and which scaler rescaled the inputs. Nothing was trained again. Beneath: the cache intervals and the scaler one input at a time came after the results.

The data is Elec2 again: 45,312 half-hours of the New South Wales electricity market, May 1996 to December 1998, in time order. The models learn from the first 36,249 rows and are tested on the last 9,063, from 1 June to 6 December 1998, the same split as lesson 1.

The two models are also lesson 1's. The first is the boosted trees with the lag input: many small trees of yes-or-no questions, built one after another. The second is logistic regression, which adds up a weight for each input and turns the sum into UP or DOWN. Logistic regression needs its inputs on a similar scale, so a scaler sits in front of it. A scaler learns each input's average and spread from the training rows, and then standardises every new row: it subtracts the average and divides by the spread, so a value of 0 means "the training average" and 1 means "one spread above it". The fitted scaler is an artifact: a small saved file with two numbers per input.

A table of five rows headed the five cases, fixed before the run, titled one fresh and one or two stale versions of each piece. Lag, fresh: the lag input computed for every row: the label of the half-hour before. Lag, daily cache: the same trained trees; the lag comes from a cache filled once a day with the last label of the day before. Lag, weekly cache: the same, the cache filled once a week (every 336 half-hours). Scaler, fresh: logistic regression; inputs rescaled by a scaler fitted on the training rows. Scaler, stale: the same weights, rescaled by an old scaler fitted on the first 10% of rows. Beneath: reported for each: accuracy, the share of answers that were UP, and for the caches how often the input differed from the fresh one.

The Main Run

Here is what the lab stored in results/stale.json.

A two-column table headed the main run, stale.json: accuracy and share of answers UP, titled both stale pieces cost accuracy; only one made the answers look strange. Left, the case; right, accuracy and said UP. Lag, fresh: 0.8193, 0.484. Lag, daily cache (input differs on 44.6% of rows): 0.7489, 0.541. Lag, weekly cache (differs on 48.0%): 0.7422, 0.528. Scaler, fresh: 0.6491, 0.757. Scaler, stale (first 10% of rows): 0.4512, 1.000. Beneath, left: really UP: 0.451 of test rows. Beneath, right: the trees without any lag scored 0.7527.

The daily cache cost 7.04 points: accuracy fell from 0.8193 to 0.7489. The cached input differed from the fresh one on 44.6% of test rows. The weekly cache cost a little more, 0.7422, with the input different on 48.0% of rows.

Now compare those with lesson 1's trees without any lag input, which scored 0.7527 on the same rows. The daily cache landed about the same as trees that never had the lag: 0.7489 against 0.7527, a 0.39-point gap in one run, too small to call. After the results I compared the two row by row. The cached-lag trees were right where the no-lag trees were wrong on 652 rows, and the reverse on 687; McNemar's test from lesson 4, counting only those rows, gives an exact two-sided p of 0.353, so the gap cannot be told apart from luck. The weekly cache, 0.7422, I did not test. So a day-old lag threw away all of what the lag had added, 7.04 points, and left the trees roughly where they would be without it. One possible reason, which is my reading and not a separate measurement: the trees had learned to trust that input a great deal, because a fresh lag is a very good guide on this data, and a stale one sent them the wrong way as often as it helped.

Persistence, the dashed line at 0.8484 on the chart below, is lesson 1's rule that simply copies the previous half-hour's label. It is a rule, not a model, and no model in this chapter has beaten it on this split.

A bar chart headed stale.json: accuracy on the 9,063 test half-hours, titled the daily cache fell to the level of the trees with no lag. Five bars on a scale from 0 to 1, with two dashed lines, persistence near 0.85 and no lag near 0.75. Trees with the lag: lag fresh about 0.82, daily about 0.75, weekly about 0.74. Logistic: scaler fresh about 0.65, scaler stale about 0.45. Beneath: 0.8193, 0.7489, 0.7422, 0.6491, 0.4512. Lines: persistence 0.8484; trees without the lag 0.7527. Daily cache against no lag: McNemar p 0.353, too close to call. One run each.

What the Outputs Showed

Accuracy needs the right answers, and in many systems those arrive late or never. So the first question a team can ask, before any label arrives, is whether the model's answers look normal. The simplest measure of that is the output mix: the share of answers that were UP.

Two panels headed share of answers that were UP, all test rows; really UP 0.451, titled silent and loud, seen from outside. Daily cache: 0.541; said UP; fresh 0.484. Accuracy fell 7.0 points. Stale scaler: 1.000; said UP on every row; fresh 0.757. Beneath: the cached input left the mix of answers close to normal. The old scaler made the model give one answer only.

With the daily cache, the trees said UP on 0.541 of rows, against 0.484 with the fresh input and 0.451 in reality. That is a shift, but a small one, and the fresh model was itself 0.033 away from reality. Nobody looking at this number without the fresh one beside it would suspect anything. With the stale scaler, the logistic model said UP on every row. Its fresh version already said UP too often, 0.757, but "every row, every day" is a different kind of number.

A line chart headed share of answers UP in each of the 26 whole test weeks; after the results, titled week by week, the cached input moved with the fresh one. Four lines over 26 weeks on a scale from 0 to 1.05. Really UP, dotted, moves between about 0.3 and 0.65. Lag, fresh moves between about 0.14 and 0.85, rising to a peak near week 9 and falling after week 16. Lag, daily cache, dashed, follows the fresh line closely, a little higher in most weeks. Scaler, stale is a flat line at 1.0 in every week. Beneath: daily cache against fresh: at most 0.214 apart in any week. The stale scaler: 1.000 in every week.

Week by week, which I measured after the results, the cached input's line rose and fell with the fresh one. In most weeks it was a little higher, and in the worst week the two were 0.214 apart. But the fresh line itself swings from 0.14 to 0.85 across the season, so a gap of that size hides inside the ordinary movement. The stale scaler's line is flat at 1.000, week after week.

So, from outside, the daily cache was silent and the stale scaler was loud. That is the most important point of this lesson: the silent failure cost 7 points of accuracy and looked normal, and a silent failure can run for months, because nothing makes anyone look. The loud one was far worse per row, but a model that says UP every day for a week is the kind of thing a person notices. I will show on a later slide that "notices" depends on which check you run, because the obvious daily check missed it too.

Where the Damage Sits

A cached input can only change an answer when the cached value differs from the fresh one. I measured this after the results, and it holds exactly: on every row where the two were equal, the model received the same row and gave the same answer.

Two panels headed the daily cache, all test rows; after the results, titled all the damage sits on the rows where the input changed. Input the same, 5,017 rows: 0.824; accuracy, cached and fresh alike: the model saw the same row. Input differs, 4,046 rows: 0.655; accuracy with the cached input; fresh 0.813. Beneath: on the rows that differ, the cached input changed the answer 28.1% of the time.

The cached input equalled the fresh one on 5,017 rows, and both scored 0.824 there. On the other 4,046 rows, the fresh input scored 0.813 and the cached one 0.655. On those rows the cached input changed the model's answer 28.1% of the time. That is the whole 7-point loss, concentrated on under half the rows.

A stale value fits best just after the refresh, so I expected the damage to grow through the day as the cached value got older. I measured it after the results, in blocks of three hours. As in lesson 1, the data does not say whether the first half-hour of a day starts at midnight, so I count it as 00:00; the clock times may be half an hour early.

A bar chart headed the daily cache by time of day, 3-hour blocks; after the results, titled the gap did not grow steadily through the day. Eight pairs of bars, fresh and daily cache, on a scale from 0.6 to 1, for blocks starting at 00:00, 03:00, 06:00, 09:00, 12:00, 15:00, 18:00 and 21:00. Fresh is about 0.78, 0.91, 0.79, 0.85, 0.83, 0.82, 0.84 and 0.76; daily cache about 0.77, 0.77, 0.78, 0.78, 0.72, 0.74, 0.76 and 0.68. Beneath: gap, points: 1.1, 13.6, 1.4, 6.9, 10.8, 7.7, 7.6, 7.3. Input differs: 24%, 55%, 37%, 36%, 45%, 49%, 54%, 58%. Over the 48 half-hours the gap follows the second (correlation 0.59) more than the input's age (0.21).

It did not grow steadily. The first three hours after the refresh lost only 1.1 points, as expected. But the next three hours, 03:00 to 06:00, lost 13.6 points, the most of any block, and 06:00 to 09:00 lost only 1.4. In this data the gap followed how often the cached value differed from the fresh one more closely than its age, and how often it differed follows the shape of the day's prices. From 03:00 to 06:00 the cached value, the label from late in the evening, differed from the fresh one on 55% of rows; from 06:00 to 09:00 on only 37%. Across the 48 half-hours of the day, the gap had a correlation of 0.59 with how often the input differed, and only 0.21 with the input's age. (A correlation of 1 means two numbers rise and fall together; 0 means no link.) Why the evening's label happens to match the morning better than the small hours is a question about this market's daily rhythm that I did not study.

Inside the Stale Scaler

Why did the old scaler make the model say UP on every row? After the results, I looked at the numbers the model actually received: every input, standardised by the old scaler, on every test row, next to the same inputs standardised by the training scaler on the training rows.

A two-column table headed the stale scaler's inputs, as the model received them; after the results, titled one input landed more than 200 spreads from anything the model had seen. Left, input: training range; right, stale, on the test rows. date: -1.30 to 1.95, against 211.89 to 244.53. nswprice: -1.50 to 25.46, against -1.71 to 26.09. nswdemand: -2.58 to 3.22, against -2.56 to 3.30. period: -1.70 to 1.70, against -1.69 to 1.70. vicdemand: -3.76 to 5.37, against -0.39 to 0.49. Beneath, left: standardised: 0 is the training average, 1 is one spread. Beneath, right: the old scaler saw rows up to 1996-08-09; the dates rose after that.

One input stands out. The date column in Elec2 is a number that mostly grows from 0 to 1 over the whole data (it steps back a few times, as lesson 6 noted). The old scaler had only seen the first 10% of the rows, whose dates ran from 0 to about 0.013, so it learned a tiny average and a tiny spread. The test rows' dates, around 0.87 to 1.0, came out as 211.89 to 244.53 spreads above its average. During training the model had never seen this input above 1.95.

Most other inputs landed close to their training range, because prices and demand in 1996 were not so different from 1998. Three inputs, vicprice, vicdemand and transfer, did something odd: in those first rows they held a single value, as the Victorian data in Elec2 begins later. A scaler cannot divide by a spread of zero, so scikit-learn set their spread to 1, and they came out squeezed near 0. Their weights in the model are small, so they did little here.

Two panels headed the date input and the model's score; after the results, titled one input decided every answer. Date, rescaled: 211.89; the smallest test value; training went up to 1.95. Lowest score: 150.9; over all test rows; a score above 0 means UP. Beneath: the date's weight, 0.728, times its average stale value added 160.3 to the score; with the fresh scaler it added 1.19.

Which Check Would Have Caught It

The previous chapter's lesson failing silently built simple checks that need no right answers and run once a day. After the results, I ran the same kinds of check on these inputs and outputs, one batch per whole test day, 188 days in all. None of them was tuned for this lab, and I wrote their exact rules knowing the results, so read this as a demonstration, not a fair trial. A range check fires when any input is below the lowest or above the highest value that input had in training. An unseen value check fires when an input that takes a few fixed values, such as the weekday or the lag, holds a value training never had. A zeros check fires when the day's share of inputs that are exactly 0 is more than 0.30 above the training share, the way a missing value filled with 0 would look.

A table headed the previous chapter's checks on these inputs, share of 188 whole test days; after the results, titled a different check caught each stale piece. Columns: lag fresh, lag daily, lag weekly, scaler fresh, scaler stale. Range, raw inputs: 0.03 in all five. Range, model's inputs: 0.03, 0.03, 0.03, 0.03, 1.00. Unseen value: 0.00 in all five. Stuck all day: 0.01, 1.00, 1.00, 0.00, 0.00. Zeros: 0.03, 0.40, 0.45, 0.00, 0.00. Output mix, day: 0.05, 0.18, 0.14, 0.00, 0.00. Output mix, week: 0.38, 0.50, 0.58, 0.46, 1.00. One answer all week: 0.00, 0.00, 0.00, 0.00, 1.00. Cells at 0.90 or more are outlined: stuck all day for lag daily and lag weekly, and range on the model's inputs, output mix by week and one answer all week for scaler stale. Beneath: weeks: 26. The range rule had no margin; with 0.1 spreads and the date left out it fires on 0.037 of days for scaler stale.

The daily cache was caught by the stuck check on every day. A stuck check fires when an input that normally moves within a day holds one single value all day. A cached lag does exactly that: it holds the night's value until the next refresh. With the fresh input, the lag held one value all day on 1 of 188 test days and 0.7% of training days, so the check is also quiet when nothing is wrong. The output checks were much weaker: the daily output mix fired on 18% of days with the cache against 5% without it, and the weekly one on 50% against 38%. Those checks fire often on the fresh model too, so they cannot tell you which days are broken.

The stale scaler was caught only by checks on what the model received. The raw inputs were untouched: the old scaler sits after them. So range, unseen, stuck and zeros on the raw inputs fired exactly as often as with the fresh scaler. The range check on the standardised inputs, the numbers the model really got, fired on all 188 days. But as I wrote it, the rule had no margin, and that matters: the period input crossed its training range by 0.003 spreads on every day, so even without the date the check would have fired every day. I measured the rule with margins after the review of this lesson. Leaving the date out, a margin of 0.01, 0.1 and 0.5 spreads made it fire on 0.197, 0.037 and 0.005 of days; with the date kept in, on every day at every margin, because the date was over 200 spreads outside. So with any sensible margin, only the date trips it, and it trips it on every day.

How Short Must a Cache Be?

The main run tested a day and a week. A real team would ask a practical question: if we must cache, how often must we refresh? I designed a follow-up after seeing the main results and wrote its design into the lab file's followup mode before it ran. It served the same trees with the lag cached for 1, 2, 6, 12, 24, 48 and 336 half-hours. A cache of 1 half-hour is the fresh input, and 48 and 336 repeat the main run; both had to come back exactly, and they did. Its results are in results/stale_followup.json.

A bar chart headed follow-up, designed after the results: how often the cache is filled, titled even a cache filled every hour cost something. Seven bars on a scale from 0.70 to 0.86, with dashed lines for persistence near 0.85 and no lag near 0.75: fresh about 0.82, 1 hour about 0.81, 3 hours about 0.79, 6 hours about 0.76, 12 hours about 0.77, 1 day about 0.75, 7 days about 0.74. Beneath: 0.8193, 0.8121, 0.7891, 0.7577, 0.7670, 0.7489, 0.7422. Every 6 hours scored below every 12 hours. The axis starts at 0.70.

A cache refreshed every hour scored 0.8121, 0.72 points below fresh. Every 3 hours: 0.7891, 3.02 points below. Every 6 hours: 0.7577, which is already close to the trees with no lag at all, 0.7527. From there on, every cache landed near that level: 0.7670 for 12 hours, 0.7489 for a day and 0.7422 for a week, none tested against it. The order was not perfectly neat: every 12 hours, 0.7670, scored higher than every 6 hours. I did not study why; one possible reason, a guess, is the same daily rhythm as before, since a 12-hour cache is filled at midday as well as midnight.

An isometric drawing of six blocks, heights to scale, headed follow-up: share of test rows where the cached input differs from the fresh one; heights to scale, titled the longer between fills, the more rows get an out-of-date input. From left to right: 8.1%, 1 hour, a short block; 19.9%, 3 hours; 36.1%, 6 hours; 36.5%, 12 hours; 44.6%, 1 day; 48.0%, 7 days, the tallest. Beneath: under each block: how often the cache is filled. The input's average age went from 1.5 to 168.4 half-hours. Designed after the results.

The share of rows with an out-of-date input rose with the interval: 8.1% for one hour, 19.9% for three, 36.1% for six, 36.5% for twelve, 44.6% for a day and 48.0% for a week. The last is close to what two unrelated labels would give, so a week-old lag carries almost no information about now.

The practical reading: for an input that changes every half-hour and that the model relies on, there is no cache interval that is free. Even one hour cost almost a point here. For an input that changes slowly, such as a customer's country, a day-old cache might cost nothing. The number that decides it is how often the cached value would differ from a fresh one, and you can measure that before you choose the interval.

The Stale Scaler One Input at a Time

The same follow-up asked which inputs carried the stale scaler's damage. It used the training scaler for every input except the named ones, which got the old scaler's average and spread.

Two panels headed follow-up, designed after the results: the stale scaler, one input at a time, titled the date did all the damage; the other old centres happened to help. Only the date stale: 0.4512; said UP 1.000: the whole failure. All stale but the date: 0.8030; said UP 0.498; the fresh scaler scored 0.6491. Beneath: a stale piece is not always worse. The old price centre pushed scores toward DOWN (-0.43 against 0.84), against the date's push toward UP.

With only the date stale, the model scored 0.4512 and said UP on every row: the whole failure, from one input. With every input stale except the date, it scored 0.8030. That is 15.39 points better than the logistic model with the right scaler, and I did not expect it.

Here is what I measured, after the results. With the right scaler, the date pushed the average score toward UP by 1.19 and the price by 0.84, and the model said UP on 75.7% of rows when only 45.1% were UP. The old scaler's average price was higher than the training average, so with it the price pushed the average score toward DOWN by 0.43 instead. One possible reason for the jump, and it is my reading rather than a test: that old price centre happened to cancel most of the date's drift toward UP, so the mix of answers came out close to reality.

I include this because it teaches something uncomfortable. A stale piece is not always worse, and a pipeline that mixes old and new pieces can score better by accident. That does not make the old piece right. It makes the model's behaviour depend on a coincidence nobody chose, which can reverse with the next month of data. The fix is still to load the scaler that belongs to the model.

Two Cheap Checks on the Pipeline Itself

The checks on the previous slide look at inputs and outputs. Two cheaper checks look at the pipeline itself, and neither needs the right answers. I measured both after the results.

Two panels headed two checks on the pipeline itself; after the results, titled recompute a sample; write down each input's age. Fresh recompute, 1%: 33 of 79; sampled rows where the daily cache disagreed; the stale scaler: 79 of 79. Input's age: 97.9%; of rows got a lag older than half an hour; average 24.5 half-hours. Beneath: neither needs the right answers. The first sampled disagreement came at half-hour 197 for the cache, 12 for the scaler.

Recompute a small sample fresh. Pick 1% of served rows at random, compute their inputs again from scratch, the slow way, and compare with what the pipeline served. Here each of the 9,063 test rows had a 1% chance of being picked, with seed 0, which gave 79 rows. The daily-cached lag disagreed with a fresh computation on 33 of them, the first at half-hour 197, about four days in. The stale scaler disagreed on all 79, from the first sampled row. Any disagreement at all is worth an alarm, because a fresh recompute and a correct cache should agree on every row.

Write down each input's age. If the pipeline stores, next to every input it serves, the time the value was computed, then staleness is a number you can read. A fresh lag is always one half-hour old. The daily-cached lag was older than that on 97.9% of rows, on average 24.5 half-hours old and at worst 48. A rule such as "refuse a lag older than one half-hour" would have fired at once. The same idea works for artifacts: store, with the saved scaler, which rows it was fitted on. Here the scaler said 9 August 1996, and the model it served had trained up to 1 June 1998.

What I Would Do Instead

Here is what the lab points to, in the order a pipeline meets it.

Record when every input and file was computed. Store the time next to each cached value, and store with each artifact the rows or dates it was fitted on and the version of the model it belongs to. That turns "is this stale?" from a guess into a lookup.

Give every cache an expiry. Decide how old a value may be before the pipeline must compute it again, and make the pipeline refuse older values instead of serving them. Choose the expiry by measuring how often an old value differs from a fresh one: here a cache filled every hour, whose values were on average 1.5 half-hours old, already differed on 8.1% of rows.

Load artifacts by the model's version, not by a file path. The stale scaler here is what happens when a pipeline loads "the scaler" from a fixed place and a retrain replaces the model but not the file. If the model and its scaler are saved together, or the scaler is part of the saved model, they cannot drift apart. In scikit-learn, a Pipeline that holds both does exactly that.

Run input checks on what the model receives. Checks on raw inputs could not see the stale scaler at all. A range check on the standardised inputs, with a small margin, saw it on every day through the date. Add a stuck check for inputs that should move within a batch: it caught the daily cache on every day.

Watch the output mix over a long enough window. A single day's share of UP answers was inside the normal range even for a model that said UP on every row. Over a week, one answer on every row happened in none of 107 training weeks.

Recompute a small sample fresh, and compare. It is the one check here that caught both stale pieces, from the first days, without any labels.

Try It Yourself

This script is the lab made small. It downloads the same data, trains the trees with the lag input and the logistic model with its scaler, and serves each one twice: fresh, and with a stale piece. For the trees the stale piece is the daily cache; for the logistic model it is a scaler fitted on the first 10% of the rows. It prints the accuracy and the share of answers that were UP for each, and how often the cached input differs from the fresh one. It does not need a GPU.

A real screenshot of VS Code with stale_demo.py open, showing the docstring that says what the script is and how to run it, the imports, the lines that load Elec2 and turn the label into 1 for UP and 0 for DOWN, the cut at 80%, the two lines that print the version and the real share of UP, the show function, the trees trained with the previous half-hour's label as an input, and the start of the daily cache; the scaler part is further down. Beneath: copy it from the box on the slide.

Before you run this lab. You need Python 3 and two libraries: pip install scikit-learn pandas. scikit-learn holds the models, the scaler and the download, and brings NumPy with it; pandas holds the table. The first run downloads Elec2 from OpenML (under 1 MB compressed), so it needs an internet connection once; scikit-learn keeps a copy in a folder in your home directory (scikit_learn_data) for later runs. It trains two small models, so it is quick; I did not time it. I ran it with scikit-learn 1.9.1 on a Mac. scikit-learn runs the same way on Windows and Linux, but I have not checked the numbers there, and another version of scikit-learn may give different decimals, which is why the first line printed is the version.

"""Stale pieces in a pipeline: a cached input and an old scaler, on the future.

Lesson 7 of 'The ML & AI Lifecycle', made small. It needs Python 3 with scikit-learn
and pandas (pip install scikit-learn pandas). The first run downloads Elec2 from
OpenML (under 1 MB) and keeps a copy for later runs.
    python stale_demo.py

Author: Roni Das
Created: 2026-09-29
"""
import numpy as np
import sklearn
from sklearn.datasets import fetch_openml
from sklearn.ensemble import HistGradientBoostingClassifier
from sklearn.linear_model import LogisticRegression
from sklearn.preprocessing import StandardScaler

# 45,312 half-hours in time order, 48 a day. The label: is the NSW price UP or
# DOWN against the average of the last 24 hours?
data = fetch_openml("electricity", version=1, as_frame=True, parser="auto").frame
X = data.drop(columns="class").astype(float)
y = (data["class"].astype(str) == "UP").astype(int).to_numpy()
cut = int(0.8 * len(y))  # learn from the first 80%, test on the rest
print(f"scikit-learn {sklearn.__version__}, {len(y) - cut:,} test half-hours")
print(f"really UP: {y[cut:].mean():.3f}")


def show(name, pred):
    acc = (pred == y[cut:]).mean()
    print(f"  {name:<17} accuracy {acc:.4f}  said UP {pred.mean():.3f}")


# 1. A cached input. The trees also read the previous half-hour's label.
prev = np.r_[y[0], y[:-1]]
Xl = X.assign(prev_label=prev)
trees = HistGradientBoostingClassifier(random_state=0).fit(Xl[:cut], y[:cut])
show("fresh lag", trees.predict(Xl[cut:]))

# The same input from a cache refreshed once a day: every half-hour of a
# day gets the label of the last half-hour of the day before.
day_start = (np.arange(len(y)) // 48) * 48
cached = prev[day_start]
show("daily-cached lag", trees.predict(Xl.assign(prev_label=cached)[cut:]))
differs = (cached[cut:] != prev[cut:]).mean()
print(f"the daily-cached input differs on {100 * differs:.1f}% of test rows")

# 2. A stale scaler. Logistic regression reads inputs rescaled to an average
# of 0 and a spread of 1, by a scaler fitted on the training rows.
scaler = StandardScaler().fit(X[:cut])
logistic = LogisticRegression(max_iter=1000)
logistic.fit(scaler.transform(X[:cut]), y[:cut])
show("fresh scaler", logistic.predict(scaler.transform(X[cut:])))

# The same model, served with an old scaler left in the pipeline: one that
# was fitted on the first 10% of the rows only.
old = StandardScaler().fit(X[: int(0.1 * len(y))])
show("stale scaler", logistic.predict(old.transform(X[cut:])))

The Lab Report

A real terminal recording headed python stale_report.py, titled every table in this lesson, from the stored files and the data. It opens with the split, 36,249 training and 9,063 test half-hours, and 43 checks against the refits: all agree. Then six numbered sections: 1, the main run for the five cases; 2, the daily cache in eight 3-hour blocks, with the gap and how often the input differs; 3, the stale scaler's inputs as the model received them, the date at 211.89 to 244.53 against a training range of -1.30 to 1.95; 4, the previous chapter's checks on these inputs, a table for the five cases; 5, the follow-up, cache intervals from 1 to 336 half-hours and the scaler one input at a time, 0.4512 and 0.8030; 6, two cheap checks: a 1% recompute, 33 of 79 and 79 of 79, and the lag older than half an hour on 97.9% of rows; 7, review follow-ups: the daily cache against the trees with no lag, 652 against 687 rows, McNemar p 0.353; the range check by margin; the logistic model's 110 all-DOWN and 2 all-UP training days; and the weekly cache's gap by day. Beneath: the lab's own report. It fits the same two models again and stops unless every stored number comes back.

The report lives in scripts/labs/lifecycle/stale_report.py. It reads the lab's stored files, results/stale.json and results/stale_followup.json, lesson 1's results/baseline.json, and the Elec2 data from scikit-learn's local copy. The lab stored accuracies and shares, not each answer, so the report trains the lab's two models again with the same data, split and settings, rebuilds the five cases exactly as the lab did, and stops unless every stored accuracy, share and scaler average comes back exactly. It makes 43 checks in all, and they all agree. It changes nothing in the lab's files.

Its json mode writes every number to results/sa-report.json, which the figures read. The demo mode checks the student script's stored run, and the box mode writes the playground below and checks that it gives the lab's numbers.

Look for the Stale Piece Yourself

This box has no model in it. It holds, for every one of the 9,063 test rows, the real label, the fresh and the daily-cached lag input, and the answers of all five cases, as two hexadecimal digits per row (counting in sixteens, with the digits 0 to 9 and a to f). The report checked that every accuracy, share, time-of-day block, stuck share and weekly mix it gives matches the lab's files. It runs in your browser.

As it is, the box prints the five cases with their accuracy and share of UP answers, how often the cached input differs, and the stuck check on the fresh and the cached lag: 0.005 of whole days against 1.0.

Then look at time of day with by_time_of_day('lag_daily_cache') and compare it with by_time_of_day('lag_fresh'). Try hours=1 for a finer view. Look at the output mix week by week with said_up_by_week('lag_daily_cache') next to said_up_by_week('lag_fresh'), and then said_up_by_week('scaler_stale'). Ask yourself which of them you would have noticed on a dashboard.

The Code, Part by Part

Loading. fetch_openml("electricity", version=1) downloads Elec2 once and reads the local copy after that. The label column holds the words UP and DOWN, turned into 1 and 0. cut is 36,249: the models learn from the rows before it and are tested on the 9,063 after it.

show. Prints one case: the accuracy against the real labels of the test rows, and the share of answers that were UP.

The fresh lag. prev shifts the labels down by one row, so each row holds the label of the half-hour before. X.assign(prev_label=prev) adds it as a ninth input, and the trees, with seed 0, learn from it. This is lesson 1's trees plus lag.

The daily cache. day_start gives each row the index of the first row of its day, because every day in Elec2 has exactly 48 rows. prev[day_start] then gives every row of a day the lag of the day's first half-hour, which is the label of the last half-hour of the day before. The trained trees are not touched; only their input changes.

The scalers. StandardScaler().fit(X[:cut]) learns each input's average and spread from the training rows, and the logistic model learns from the standardised rows. old is a second scaler fitted on the first 10% of rows only. The last line serves the same logistic model through the old scaler.

In the lab file, stale_lab.py does the same for the five cases, adds the weekly cache, and stores the numbers and both scalers' averages in . Its mode adds the other cache intervals and the scaler one input at a time.

How to Keep a Pipeline's Pieces Fresh

A hand-sketched column of six boxes joined by arrows, headed keeping the pieces of a pipeline fresh, titled record it, expire it, test it, recompute it. 1, store when each input and file was computed. 2, give every cache an expiry, and refuse old values. 3, load artifacts by the model's own version. 4, run input checks on what the model receives. 5, watch the mix of answers over a week. 6, recompute a small sample fresh, and compare. Beneath: here: a 1% recompute found the cache in 33 of 79 sampled rows and the scaler in all 79.

Store when each piece was computed. For every cached input, keep the time it was computed next to the value. For every artifact, keep the data it was fitted on, the date, and the model version it belongs to. This is often called lineage, and the next lesson planned for this chapter measures what it buys.

Give every cache an expiry. Before choosing it, compute, on past data, how often a value of each age differs from a fresh one. Choose the longest age whose disagreement you can accept, write it into the pipeline's configuration, and make the pipeline recompute or refuse any value older than that.

Keep the model and its artifacts together. Save the scaler inside the same object as the model, or give both the same version and load them by it. Never let a retrain replace one without the other.

Check what the model receives. Run range checks after every transformation, on the numbers that reach the model, and a stuck check on inputs that should move within a batch.

Watch the output mix over a long window. A day can be unusual for good reasons. A week of one answer should not happen.

Recompute a sample. Every day, recompute 1% of served inputs from scratch, compare, and alarm on any disagreement.

Cache It, or Compute It Fresh?

A two-column table headed grounded in this lesson's numbers, titled cache it, or compute it fresh? Left, a cache is fine when: the value changes slowly next to how often it is used; a fresh sample agrees with the cache almost always; the input's age is stored and checked; a slightly old value costs little. Right, compute it fresh when: the value changes often: the lag differed on 44.6% of rows after a day; the input is one the model leans on heavily; nobody records when it was computed; a cache filled every hour already cost 0.7 points here. Beneath, left: cache what is slow to change and cheap to be wrong about. Beneath, right: compute fresh what the model depends on.

Use a cache when the value changes slowly compared with how often it is read. A customer's country, a product's category or last month's average barely change within a day. them for a day costs little and saves a lot of work.

Use a cache when a fresh sample agrees with it almost always. Measure it: if a fresh recompute and the cache disagree on a tiny share of rows, the cache is doing its job.

Use a cache only when its age is stored and checked. A cache with a recorded age and an expiry is a decision. A cache without them is a guess that nobody will revisit.

Do not cache an input that changes as fast as the model's answers. Here the lag changed within hours, and even a one-hour cache cost 0.72 points.

Do not cache an input the model leans on heavily without measuring the cost. Here a day-old lag left the trees about the same as trees that never had the lag, 0.7489 against 0.7527, and all 7.04 points the lag had added were gone.

Do not reuse a fitted artifact across retrains. A scaler, an encoder or a vocabulary fitted once belongs to the model it was fitted with. Here an old scaler turned a model that scored 0.6491 into one that said UP on every row.

What These Numbers Can and Cannot Tell You

A two-column page headed read before you trust these numbers, titled what these runs are, and what they are not. They are: one dataset: Elec2, June to Dec 1998; two models, one run each; staleness I built on purpose; the design came first; intervals, one input: after; checks computed after; one lag input cached. They are not: not a rate for other data; not every kind of model; not a bug found in the wild; not changed after the run; chosen knowing the results; not tuned, not tested ahead; not every input a pipeline caches.

One dataset, two models, one run each. Everything here is one electricity market in the second half of 1998, lesson 1's trees plus lag and its logistic model. With other data, other inputs and other models, the costs would move. No significance test was declared, so I do not call any gap significant.

Staleness I built on purpose. The caches and the old scaler are faults I put into the pipeline to measure them. They are the kind of fault that happens in real systems, but they are not a bug I found in someone's system, and a real stale piece might be stale in a messier way.

One input cached, one scaler stale. Real pipelines cache many inputs at many intervals. What a stale input costs depends on how much the model relies on it and how fast it changes; the lag here is an extreme case on both counts, which makes it a clear demonstration and a poor estimate for other inputs.

The follow-up came after. The cache intervals and the scaler one input at a time were designed after I saw the main results and written down before they ran.

The checks came after. The previous chapter's checks were computed on these inputs after the results, with rules I wrote knowing what I was looking for. They show which kind of check can see which kind of staleness; they are not a fair trial of the checks. The reasons for the time-of-day pattern and for the 0.8030 are my reading, labelled as such.

What to Do Next

A hand-drawn list headed before you trust a pipeline, titled five questions for every stored piece. When?: is the time each input and file was computed stored with it? How old?: what is the oldest value the pipeline will still serve? Which file?: is the scaler loaded by the model's own version, or by a path? Which check?: do the input checks look at what the model receives? Fresh sample?: does anyone recompute a few rows fresh and compare? Beneath: here: a lag cached once a day was older than half an hour on 97.9% of rows.

Take the pipeline of a model your team runs, and list every stored piece in it: every cache, every saved file that a step loads. For each one, ask the five questions on the card. If the answer to "when?" is "we do not know", start there. Adding a timestamp to a cached value is a small change, and it makes every later question answerable.

Then run a small version of this lab on your own pipeline. Take a day of served rows, recompute their inputs from scratch the slow way, and count how many disagree with what was served. If the answer is zero, you have learned your caches are healthy. If it is not zero, you have found a stale piece before it found you.

The chapter plan's next lesson is lineage and rollback: rebuilding last month's model exactly, and what breaks when you cannot.

A closing card headed to keep, titled a stale piece raises no error; store its age and check it. In large type: 0.8193, then 0.7489. Beneath: one input cached once a day. The mix of answers barely moved: UP 0.484, then 0.541. A stuck-input check caught it on every day. Then: an old scaler: 0.4512, UP on every row. Checks on the raw inputs saw nothing; a check on what the model received saw it every day. Then: store when each piece was computed, give caches an expiry, check what the model receives, and recompute a sample fresh. Last: one dataset, one run each: a way to check, not a law.

Knowledge Check

Knowledge Check

4 questions - Score 80% to pass

Q1

A lag input cached once a day cut the trees from 0.8193 to 0.7489, yet their share of UP answers moved only from 0.484 to 0.541. Why does that make it the more dangerous of the two stale pieces?

Q2

The stale scaler made the logistic model say UP on every row. Which check caught it on every test day?

Q3

In the follow-up, how did a cache refreshed every hour compare with the fresh lag input?

Q4

The daily cache's damage by time of day did not grow steadily after the midnight refresh. What did it follow more closely?

The five cases were fixed in the design. Two stale versions of the lag input, cached for a day or for a week, and one stale scaler, fitted on the first 10% of the rows only, 4,531 half-hours ending on 9 August 1996. That is the kind of file that gets left in a pipeline when someone builds the first version, saves the scaler, and later retrains the model without replacing it. No model was trained again for a stale case: each one is the same trained model, served with one old piece in front of it.

A sequence diagram with four columns: new row, input step, the cache, the model. Headed one served half-hour, daily cache, titled the cache answers first, and it answers with an old value. Step 1, the cache is filled once, before 00:00. Step 2, the new row sends the input step a new half-hour. Step 3, the input step asks the cache: lag input, please. Step 4, the cache sends back the stored value. Step 5, the input step sends the model the inputs. Step 6, the model sends back UP or DOWN. Beneath: fresh, step 3 would compute the label of the half-hour before. Nothing in this call knows how old the stored value is.

For each case the lab stored the accuracy, the share of answers that were UP, and, for the two caches, the share of rows where the cached input differed from the fresh one. One run each; the design described the numbers and declared no significance test, a calculation of how likely a difference is to come from chance alone.

The stale scaler was worse still. The logistic model with its own scaler scored 0.6491, as in lesson 1. With the old scaler it scored 0.4512 and said UP on every one of the 9,063 rows. 0.4512 is exactly the share of test rows that were really UP, which is what a model that always says UP must score.

I also measured the weekly cache by day, counting the refresh day as day 1, and there is a trap I met in the retraining lesson: its refresh always came on a Tuesday, so day 6 is always a Sunday. Its gap to the fresh input, in points, was 2.5, 7.9, 8.0, 3.9, 3.8, 17.7 and 9.9 for days 1 to 7. Day 1 had the smallest gap and day 6 the largest (0.665 against 0.843 fresh), but the gaps rise and fall rather than grow, and I cannot separate age from weekday.

Logistic regression adds up weight times input for every input, and says UP when the sum, its score, is above 0. The date's weight was 0.728. Multiplied by an input of more than 211, it added 160.3 to the average score, where with the right scaler it added 1.19. Nothing else in the sum comes close. The lowest score over all 9,063 test rows was 150.9, so every row was UP. This part is arithmetic on the stored model, not a guess.

This is the same input that decided lesson 6. There the trees treated every date past their training dates as the last one they knew. Here a scaler turned the same growing date into a number far outside anything the model had learned from. An input that only says "when" is dangerous in both places.

The daily output check missed the stale scaler completely, which surprised me. It compares each day's share of UP answers with the lowest and highest seen on training days. On the 755 training days the logistic model said DOWN on every row on 110 days and UP on every row on 2, so the training range of the daily share was the whole scale, 0.0 to 1.0. For this model the check could never fire, in either direction, whatever it said. Only over a whole week did it stand out: one answer on every row of a week happened on none of the 107 training weeks and on all 26 test weeks.

This is a real run in VS Code's terminal (python stale_demo.py).

A real screenshot of VS Code's terminal after running python stale_demo.py. It prints scikit-learn 1.9.1, 9,063 test half-hours; really UP: 0.451; fresh lag, accuracy 0.8193, said UP 0.484; daily-cached lag, accuracy 0.7489, said UP 0.541; the daily-cached input differs on 44.6% of test rows; fresh scaler, accuracy 0.6491, said UP 0.757; stale scaler, accuracy 0.4512, said UP 1.000.

When I ran it, it printed scikit-learn 1.9.1 and 9,063 test half-hours, really UP 0.451, then the four cases: fresh lag 0.8193 and UP 0.484, daily-cached lag 0.7489 and 0.541, fresh scaler 0.6491 and 0.757, stale scaler 0.4512 and 1.000, with the cached input different on 44.6% of rows. All of it matches the lab's stored stale.json, and the longest printed line was 52 characters. The report's demo mode checks all of it.

To try the follow-up yourself, change both 48s in the day_start line to 2 for a cache filled every hour, or to 12 for every six hours. When I did, it printed 0.8121 and 0.7577 for the cached lag, the follow-up's stored numbers.

What came before the run, in stale_lab.py: the data, the two models, the five cases and what to report. What came after I saw the main results: the follow-up's cache intervals and the scaler one input at a time, designed after the results and written down before it ran. What came after all the results, in the report: the time of day, the rows where the input differs, the sketched day and its selection rule, the scaled inputs and scores, the checks and their exact rules, the weekly output mix, the sample recompute and the input's age. After a review of this lesson, the report also compares the daily cache with the trees that never had the lag, row by row, measures the range check with margins, counts the logistic model's one-answer training days, and gives the weekly cache's gap by day.

Four brand cards headed the tools, with their logos, titled what ran where. scikit-learn: both models, both scalers, the data download. pandas: the table of half-hours. NumPy: the caches, the days and weeks. Python: the report, the demo, the box.

results/stale.json
followup