Why Production Breaks

Watching Inputs Before the Answers Arrive: Drift Measures Against the Real Error

0 of 21 complete

0%

Contents

Back|Why Production BreaksWatching Inputs Before the Answers Arrive: Drift Measures Against the Real Error
1/21
61 min left
Prerequisites
Failing Silently: Which Checks Catch a Broken Input Before the Answers Arriverequired
Related Topics
Drift Alarms Against Real Harm: When the Alarm Rings, Is the Model Worse?Data Engineering for MLWhy ML Models Fail in Production: The Production GapCore ConceptsMonitoring and Drift Detection: Catching a Model That Fails Without an ErrorCore ConceptsHandling Imbalanced and Messy Data: Why Your 99% Accuracy Is a LieData Engineering for MLA Repaired Golden Set Reported 100%. The System Earned 70%.LLM Evaluation and Error Analysis
1 of 21

Watching the Sky Instead of the Sales

Imagine a woman who owns three coffee carts in a city. Staff run the carts, and the sales figures go to a bookkeeper, who sends her the accounts once a month. Every week she orders milk and cups for each cart. To decide how much, she uses a notebook from last year that says how much sold on cold mornings, warm mornings and rainy ones.

She cannot see this month's sales until the accounts arrive, so she watches what she can see: the weather. Each morning she compares it with the same week of last year. This spring it looks almost the same as last spring, so she feels safe and orders what the notebook says.

An illustration of a woman standing, her hair in a bun, wearing a sweater and wide trousers, holding a paper coffee cup. Headed watching what you can see, titled the weather looked like last year. The riders did not. Beside her: bike rentals, 2012 against 2011. The weather, all year: 20.1 then 20.7 C on average, 65% then 66% of hours clear. The riders, every month: 1.41 to 2.53 times as many as the same month a year before. A model trained on 2011 missed 2012 by 61.6 to 110.7 rentals an hour, month by month.

The accounts arrive at the end of May, and they are a shock. The carts sold far more than last year, and ran out of cups on most afternoons. A new tram line had opened, and it stopped next to two of her carts. Nothing about the weather could have told her that. The weather was fine. The customers had changed.

Two more things are worth noticing. One morning in June she compared the week's weather with her whole notebook, all twelve months, and it looked alarming: only warm days, no cold ones. Of course. It was June. And what would have warned her early was not the sky. It was a phone call to one cart at closing time, on a few days, asking how many cups were sold. This lesson measures all three of those ideas on the chapter's bike-rental data.

Where This Lesson Starts

This is the fifth lesson of the chapter, and it uses the same public data and the same kind of program as the first four: two years of hourly bike rentals from Washington, D.C., and a program that guesses how many bikes will be rented in an hour. Lesson 2 trained that program on 2011 and scored it month by month through 2012. The version that was never updated missed by 89.5 rentals an hour over the year, almost always too low, because many more people rented bikes in 2012. Lesson 2 ended with a rule: the error on the newest month whose real answers have arrived is the number to decide on.

Lesson 4 looked at the time before the real answers arrive and tested simple checks on the inputs. One of them, a measure of how far the mix of weather values had moved (it is called psi, and the words slide explains it), fired on every clean day. Its follow-up found why: a day is too small a batch, and one day's weather against a year and a half of training always looks different.

This lesson asks the question psi is really built for. Take whole months, big enough batches, and measure how far the inputs moved. Does that rise and fall with the real error, so that it could warn about the error until the answers arrive?

Eight Words for This Lesson

A hand-drawn list headed eight words for this lesson, titled watching a model before the answers arrive. Label: the real answer for one case, here the bikes really rented in an hour. Label delay: the time between a guess and its real answer: an hour here, months for a loan. Input drift: the model's inputs move away from the data it learned from. Output drift: the model's own guesses change, for example their average. Performance drift: the real error grows. Needs labels, so it is known last. Reference: what a drift measure compares with: all of 2011, or the same month of 2011. psi: population stability index: one number for how far a column's mix of values moved. Rank correlation: Spearman: do two lists rise and fall in the same order? 1 same, 0 no same-order relation, -1 opposite. Beneath: input and output drift are known at once. The real error is known last.

A label is the real answer for one case: here, the number of bikes really rented in an hour. The label delay is the time between a guess and its label. For bikes it is one hour. For a loan, whether it is repaid can take years; for a fraud check, a customer may dispute a payment weeks later. When labels are late, teams watch other things in the meantime.

Drift means something moving away from what the model learned from. It comes in three places. Input drift is the model's inputs moving: this month's weather is not like the training weather. Output drift is any change in the model's own guesses, for example their average rising or falling. Performance drift is the real error growing, and it needs labels, so it is known last.

Every drift measure compares today with a reference: here, either all of 2011 or only the same month of 2011. psi, from lesson 4, is one number for how far the mix of values in a column moved. Last, a rank correlation (the kind called Spearman) says whether two lists rise and fall in the same order: 1 means exactly the same order, 0 means no same-order relation, and -1 means the opposite order.

Three Kinds of Drift, and When Each One Speaks

The three kinds of drift sit in three different places, and they become known at three different times.

A hand-drawn sketch headed sketched: three places to watch, and when each one speaks, titled input drift, output drift, performance drift. A row of three boxes joined by arrows: inputs: weather, hour, day; the model; its guesses. Under the inputs box: input drift, known at once. Under the guesses box: output drift, known at once. A dashed arrow runs down from its guesses to a box, real counts, later, which points left to a box, guess against real. Under it: performance drift: the real error, known only once the labels arrive. Beneath: here a label arrives one hour after the guess. In many systems it takes days or months, and the two drifts on top are all there is until then.

Input drift can be measured without any labels: you have this hour's weather, you have last year's, and you can compare them. Output drift needs no labels either, because the model answers at once. Performance drift needs the labels. Until they arrive, nobody knows the real error. One catch: a measure over a whole month, like the ones in this lab, still has to wait for the month to end, and for bikes the labels exist by then too. The early warning is real only when labels take much longer than the batch you measure.

That timing is the whole reason teams watch inputs. If input drift always rose when the real error rose, you would have an early warning for free. But the three can come apart, and it helps to see why before looking at any numbers. The inputs describe what the model is asked about. The error depends on the answer too. If the answers change while the inputs stay the same (more riders on days with the same weather), the error grows and the inputs say nothing. And if the inputs change in a way the model already knows how to handle (a warm month, which it has seen before), the inputs move a lot and the error need not move at all.

A sequence diagram with four columns: the city, the model, the monitor, the records. Step 1, the city sends the model this hour's inputs. Step 2, the city sends the monitor the same inputs. Step 3, the monitor computes psi against last year. Step 4, the model sends the monitor its guesses. Step 5, the monitor computes output drift. Step 6, the records send the monitor the real counts, later. Step 7, the monitor computes the real error. Headed one month in production, as the lab plays it, titled two drifts need no labels; the real error does. Beneath: steps 3 and 5 need no answer. A monthly psi still waits for the month to end, and for bikes the labels exist by then too. Step 7 needs the labels.

You will also meet two older names for the same ideas. Many guides call a change in the inputs , and a change in how the inputs relate to the answer : the same weather, hour and day now bring a different number of riders. With the year column set aside (the frozen model could not use it), concept drift is what happened in 2012. The same kind of hour brought more riders than it did a year before. Data drift can be measured from inputs alone. Concept drift, by its definition, needs the answers.

What the Lab Ran

I wrote the design at the top of the lab file, watch_lab.py, before I ran it, and the chapter plan, CHAPTER-PLAN.md, records it the same day ("WATCH LAB (batch 5) designed before running").

A flowchart headed what the lab ran, titled two models, twelve months, four measures. Two boxes, frozen: trained once on 2011, 8,645 hours, and monthly: retrained before each month, lead to each month of 2012, which leads to two boxes: real error: MAE, and psi_year, psi_same, weather_mix, pred_shift. Both lead to rank correlation over 12 months. Beneath: designed before it ran, in watch_lab.py: the same boosted trees as lessons 1 to 4, default settings, random_state 0. 12 points per correlation: a description, no significance test.

The two models are lesson 2's, and both are scikit-learn's boosted trees (HistGradientBoostingRegressor, default settings, random_state=0): many small decision trees, each asking yes-or-no questions about the inputs and each correcting the ones before. The frozen model was trained once on all of 2011, 8,645 hours, and never again. The monthly model was retrained before each month of 2012 on every hour before that month. Each month's real error, the MAE (mean absolute error: the average size of the miss, in rentals an hour), was computed again here the same way as in lesson 2, and the report checks it matches lesson 2's stored number exactly.

The four drift measures, fixed before the run, are explained on the next slide. For each one, the lab computed the rank correlation with the monthly MAE across the 12 months, once for each model.

What the design said about the result. A significance test is a calculation of how likely a pattern this strong would be by luck alone. Twelve points are far too few for one. So the design declared, in writing, that the lesson would describe the pattern and nothing more. Everything after the headline, from the weather comparison to the labelled hours, was chosen after I saw the results, and each figure says so.

Four Ways to Say the Inputs Moved

Each measure compares one month of 2012 with 2011.

An editorial page in four labelled zones, headed the four drift measures, fixed before the run, titled each one compares a 2012 month with 2011. psi_year, input drift, against the whole year: the month's temperature, feels-like temperature, humidity and wind, each against all of 2011; the largest of the four. January 2012: 3.19. psi_same, input drift, against the same month: the same four columns against the same calendar month of 2011; the largest. January 2012 against January 2011: 0.73. weather_mix, input drift in the weather label: how far the shares of clear, misty and rain hours moved from the same month of 2011 (0 = the same mix, 1 = nothing in common). January: 0.053. pred_shift, output drift: the frozen model's average guess for the month, minus its average guess for the same month of 2011. January: +6.4 rentals an hour. Beneath: psi's usual line is 0.25. The lab also scored the real error of each month, and ranked each measure against it over the 12 months.

psi_year is psi as most monitoring guides use it: the month's values of the four weather numbers (temperature, "feels like" temperature, humidity and wind speed) against all of 2011, the model's whole training year. Each column is cut into ten bands, each holding a tenth of 2011's hours, and psi adds up how far the share of hours in each band moved. The measure keeps the largest of the four columns. psi_same does the same against only the same calendar month of 2011: January 2012 against January 2011, using the same ten bands, cut from all of 2011. That is the fix lesson 4 tried for single days.

weather_mix looks at the weather label instead (clear, misty, rain). It is the share of hours that would have to change label to make the month's mix match the same month of 2011: 0 means the same mix, 1 means nothing in common.

pred_shift is the one measure of output drift. It is the frozen model's average guess for the month, minus its average guess for the same month of 2011. It needs no labels: both are the model's own guesses.

psi's usual line is 0.25, a rule of thumb from lesson 4: a simple figure people use because it often works, not one worked out for this data. A measure that could warn about the error should be high in the months where the error is high and low where it is low. The rank correlation measures exactly that.

No Measure Moved in Step With the Error

Here is the headline, from the lab's stored file, results/watch.json.

A bar chart headed rank correlation of each drift measure with the monthly MAE, 12 months, titled no drift measure moved in step with the error. Four pairs of bars, frozen model then monthly model, on a scale from -1.0 to 1.0, each bar growing up or down from a dashed line at 0 labelled no relation: psi_year about -0.34 (down from 0) and 0.05; psi_same about 0.41 and 0.42; weather_mix about 0.39 and 0.06; pred_shift about 0.58 and 0.19. Beneath: frozen: -0.34, +0.41, +0.39, +0.58. Monthly: +0.05, +0.42, +0.06, +0.19. 1 would mean the measure ranks the 12 months exactly as the error does; 0, no same-order relation. 12 points: a pattern, not a proof.

For the frozen model, the rank correlations with its monthly MAE were -0.34 for psi_year, 0.41 for psi_same, 0.39 for weather_mix and 0.58 for pred_shift. For the monthly model they were 0.05, 0.42, 0.06 and 0.19. A measure that could take the error's place would sit near 1. None came close. The usual psi, psi_year, even ran slightly the wrong way for the frozen model: its months with the highest psi tended to be months with a lower error.

A rank correlation is easier to feel with the months themselves. The three months with the frozen model's largest errors were March (110.7), September (109.4) and October (108.3). The three months with the largest pred_shift were March, September and April: two of the three match. For psi_same the top three were March, September and November: again two of three. For psi_year the top three were July, August and November: none of the three. For weather_mix they were September, December and April: one. So pred_shift and psi_same did pick out the two worst months, and missed others; psi_year pointed at the hottest summer months, which were not the worst. I read this table after the results; the design only asked for the correlations.

I use a rank correlation rather than the ordinary kind on purpose, and the design fixed it before the run. It asks only about order: is the month with the third-highest drift also the month with the third-highest error? One extreme month cannot dominate it, and it does not care whether psi and the MAE are on the same scale, which they are not.

These are 12 points each, and the design declared that no significance test could be run on 12 points. So I read them as a description of this year only. With 12 points, a correlation of 0.4 or 0.6 can appear from a pattern that is not really there; what the numbers can say is that no measure tracked the error closely here.

psi_year Measures the Season

psi_year was far over the line every month. Before calling that drift, I asked (after the results) what it would say about months that certainly had not drifted: the 2011 months the frozen model learned from, each scored against 2011 as a whole.

A line chart headed psi_year, added after the results: 2011's own months, beside 2012's, titled psi_year measures the season, not a change. Two lines over January to December on a scale up to 6, with a dashed line at 0.25 labelled usual line: 2011, the months the model learned from, starts near 5.0 in January, falls to about 1.5 in April, rises to about 5.1 in July and falls to about 1.8 in October before rising to about 3.4 in December; 2012, dashed, follows a similar path, lowest near 1.3 in March and highest near 4.9 in July. Both lines stay far above 0.25. Beneath: 2011 against itself: 1.53 to 5.10, all 12 over the line. 2012: 1.34 to 4.94. Rank correlation of the two lines +0.78.

They scored just as high. 2011's own months ran from 1.53 to 5.10, all 12 over the line, and they rise and fall with 2012's months in almost the same order: the rank correlation between the two lines is 0.78. July is the highest in both years and April or March the lowest. The column that scores highest is always the temperature or the "feels like" temperature. And psi_year has a rank correlation of 0.81 with how far the month's average temperature sits from 2011's yearly average of 20.1 C. So psi_year mostly measures how far a month is from an average day of the year: the season. A July is hot compared with a whole year. That is not a change in the world.

Two panels headed is it the batch size? added after the results, titled a month is big enough; the season is not the same. 720 random hours of 2011: 0.016: median psi against all of 2011, 500 draws; highest 0.034. One month of 2012: 1.34 to 4.94: psi_year against all of 2011, each of the 12 months. Beneath: lesson 4's follow-up: at 24 hours, 82.55% of random training draws crossed 0.25 on temperature; at 720 hours, 0%.

Could the batch size explain it, as it did in lesson 4? That follow-up found that at 24 hours, random draws from the training hours themselves crossed 0.25 on temperature 82.55% of the time, and at 720 hours never. I checked the same here, after the results: 500 random draws of 720 hours from 2011, each against all of 2011, with a fixed seed (a starting number for the random choices, so a rerun picks the same hours). The median psi (the middle value when all 500 are lined up) was 0.016, and the highest 0.034. A month is big enough. The 1.34 to 4.94 is not noise from a small batch; it is the season, measured correctly against the wrong reference.

Over the Year, the Weather Looked Like 2011. The Riders Did Not.

If the inputs did not explain the frozen model's error, what did? I compared the same month in both years (after the results).

A line chart headed average temperature each month, added after the results, titled the weather in 2012 looked like the weather in 2011. Two lines over January to December on a scale up to 35 C: 2011 rises from about 8 in January to about 31 in July and falls to about 13 in December; 2012, dashed, lies almost on top of it, a little higher in January and March and lower in November. Beneath: same month, 1.5 C apart or less in 9 of 12 months; more in Jan (+3.2), Mar (+4.9), Nov (-2.7). All of 2012 against all of 2011: the largest psi of the four columns is 0.048.

Over the whole year, the weather of 2012 looked very much like 2011's. Month by month it moved more. The average temperature of the same month differed by 1.5 C or less in 9 of 12 months. March 2012 was warmer, by 4.9 C, and January by 3.2; November was colder by 2.7. Over the whole year, 2011 averaged 20.1 C and 2012 20.7; 65% of 2011's hours were labelled clear and 66% of 2012's. All of 2012 against all of 2011 gives a psi of at most 0.048 in any column, far under the line. Single months differed more: psi_same crossed 0.25 in 9 of 12 months, and the weather label moved most in September (0.269) and December (0.214, clear hours falling from 67% to 45%). So the fair summary is: similar on yearly averages, not identical month by month.

An isometric drawing of eight blocks in four pairs, headed rentals per hour, the same month in both years, heights to scale, titled similar yearly weather, many more riders. Mar 11, 88; Mar 12, 222; Jul 11, 190; Jul 12, 274; Sep 11, 178; Sep 12, 304; Dec 11, 118; Dec 12, 167. In each pair the 2012 block is the taller. Beneath: average temperature, 2011 then 2012: Mar 13.6 and 18.4 C; Jul 31.1 and 30.8 C; Sep 25.1 and 25.4 C; Dec 13.3 and 13.2 C. Rentals: 2.53 times, 1.44 times, 1.71 times, 1.41 times. March has the largest temperature gap of the year.

The riders did not look like 2011. Every month of 2012 had between 1.41 and 2.53 times as many rentals an hour as the same month of 2011, 234.7 an hour against 143.8 over the year. The four months in the drawing include March on purpose: it had the largest temperature gap of the year, and still most of its extra riders are left over once the weather is counted, as the next figure shows. That is the change lesson 2 measured, and it is a change in the target, the number the model guesses, not in its inputs.

A dot chart headed each month of 2012, added after the results, titled the frozen error followed the extra riders. Twelve dots, one per month, with the extra rentals an hour over the same month of 2011 across (40 to 160) and the frozen MAE up the side (50 to 120). The dots rise from lower left to upper right; seven are labelled beside their own dot: Dec near 49 and 62, Jan near 75 and 69, Feb near 75 and 75, Jun near 82 and 89, Jul near 84 and 94, Sep near 126 and 109, Mar near 134 and 111. Beneath: rank correlation +0.97, close to arithmetic: the frozen miss is the growth less the model's own weather response (pred_shift, at most 24.5). The growth is known only once the counts arrive.

Output Drift Moved a Little, the World a Lot

pred_shift, the output drift, had the highest correlation of the four, 0.58. It is worth seeing what it actually measured.

A bar chart headed output drift against the real change, rentals an hour, added after the results, titled the guesses moved by at most 24.5; the real counts by 48.9 or more. Twelve pairs of bars growing from 0, January to December 2012, on a scale from -20 to 140: the model's average guess change (pred_shift) stays between about -5 and 25, going below 0 in June, July and December, while the real rentals change is between about 49 and 134 in every month, highest in March and September. Beneath: largest output moves: Mar +24.5, Sep +22.0. Three months went down: Jun -1.7, Jul -5.0, Dec -4.7. The real counts rose in every month.

The frozen model's average guess for a 2012 month was never more than 24.5 rentals an hour away from its average guess for the same month of 2011. The real counts moved by 48.9 to 134.2 in the same months. The lab did not store the model's guesses on 2011, so the report works out the 2011 average as the 2012 average minus pred_shift, and says so.

This is not a fault in the measure. A frozen model's guesses can only move as far as its inputs push them. In March 2012, which was 4.9 C warmer than March 2011, the model raised its guess by 24.5, partly, it seems, because warm days bring riders. That reason is untested: March 2012 was also clearer than March 2011, with different humidity and a different calendar. It could not raise its guess for whatever brought the extra riders, because nothing in its inputs described it. Output drift is input drift seen through the model: it tells you how much the model thinks the world changed, never how much it really did.

Why did pred_shift still rank the months somewhat like the error did? Its rank correlation with the real change in riders was 0.65 (measured). My guess, not tested: the months with better weather than last year were also months with the biggest growth, perhaps because good weather brings out more of a growing number of riders. With 12 points I cannot tell that apart from chance.

The Same-Month Reference as an Alarm

psi_same is the fairer reference: January against January. At a month's size, batch size alone no longer makes psi cross the line (lesson 4's follow-up, and the random draws on the season slide). So I read it, after the results, the way a team would: as an alarm at the usual 0.25.

Two panels headed psi_same read as an alarm at 0.25, added after the results, titled it fired in good months and stayed quiet in bad ones. Frozen, 3 quiet months: 89.3, 93.6, 61.6: MAE in Jun, Jul, Dec, when psi_same stayed under the line. Monthly, best month: 28.7, psi 0.31: Feb: its best month of the year, with psi_same over the line. Beneath: psi_same fired in 9 of 12 months. Mean MAE when it fired against when it was quiet: frozen 92.2 and 81.5; monthly 44.5 and 40.9.

It fired in 9 of the 12 months. For the frozen model, the months it fired averaged an MAE of 92.2, and the three quiet months, June, July and December, averaged 81.5. So the quiet months were not good months: the frozen model missed them by 89.3, 93.6 and 61.6 rentals an hour, between 32% and 37% of each month's real average. An alarm that stays quiet during those months would have told a team that all was well while the model was far behind.

It fails the other way too. The monthly model, retrained every month, had its best month of the year in February, with an MAE of 28.7. psi_same was 0.31 in February, over the line. So the alarm fired on the best month of the healthy model. For the monthly model, the months it fired averaged 44.5 and the quiet months 40.9: close to no difference.

psi_same measured something real. March 2012 was warmer than March 2011, and September 2012 was much clearer (75% of hours clear against 48%, the largest weather_mix of the year, 0.269). Those are true changes in the weather. They are just not the same thing as the model being wrong. A model that has seen warm Marches and clear Septembers can handle them.

One more detail changes the verdict itself, and a review of this lesson found it. Lesson 4 also scored whole months against the same month of 2011 and found 2 of its 5 test months over the line (September 0.87 and November 1.01), where this lab finds 4 of the same 5. The difference is how the ten bands are cut: this lab cuts them from all of 2011, while lesson 4 cut them from the same month of 2011 itself (and used only the last 588 hours of August). I recomputed psi_same lesson 4's way for every month (after the results; the report checks it against lesson 4's stored values): it crosses 0.25 in 7 of 12 months instead of 9, because August falls from 0.28 to 0.11 and October from 0.44 to 0.15. So whether a month crosses 0.25 depends on how the bands are cut, which is one more reason not to treat that line as an alarm.

The Real Alarm: A Few Labels, Early

If the inputs could not show the frozen model's problem, what could, and how soon? Everything on this slide was added after the results, and it reads lesson 2's stored guesses; no model was trained for it.

A table headed the frozen model, read on the morning of 1 February 2012, added after the results, titled only the labels said which way it was wrong. psi_year: 3.19: over the line, as it was in every 2011 month too. psi_same: 0.73: over the line; January was warmer than January 2011. weather_mix: 0.053: about the same mix of clear, misty and rain. pred_shift: +6.4: the guesses were a little higher than a year before. January's labels: MAE 68.9, signed -67.8: 53% of the month's average, too low. 24 labelled hours: MAE between 47.3 and 91.6 in 90% of 1,000 draws; too low in 100% of them. Beneath: the drift numbers said 'something differs'. The labels said 'too low, by about half'.

Put yourself on the morning of 1 February 2012, with the frozen model in service. psi_year says 3.19, but it said as much about every month of 2011. psi_same says 0.73: January was warmer than a year before. weather_mix says 0.053, about the same weather. pred_shift says +6.4: the model's guesses were a little higher than last January. None of them says which way the model is wrong, or by how much. January's labels do: an MAE of 68.9 and a signed error of -67.8, 53% of the month's average, too low. That is lesson 2's newest labelled month, and it said the same thing every month after.

A two-column table headed 24 labelled hours drawn at random from each month, 1,000 draws, added after the results, titled a few labels showed the frozen model's lean. Left, frozen: month; 24 hours; right, monthly (sign); agree. Twelve rows, January to December. The frozen whole-month MAE runs from 61.6 (Dec) to 110.7 (Mar), and its 24-hour ranges from 41.3 to 84.4 (Dec) up to 78.2 to 149.4 (Mar); Jan reads 68.9; 47.3 to 91.6. The monthly column reads, for example, Jan 68.9 (-67.8); 100%, Feb 28.7 (-9.6); 87%, Apr 39.6 (+3.0); 61%, and May 42.9 (+16.8); 93%. Beneath: frozen: too low in every draw, every month; 85.5% of its hours were too low. Monthly: draws agree with the month's own sign in 61% to 100%; least in Apr (+3.0).

You may not have a whole month of labels. So I asked how few would do. For each month I drew 24 hours at random, 1,000 times, and scored the model on those 24 alone. Those hours come from across the whole month, so all 24 exist only at its end; a team would collect them as a small daily sample, a little under one labelled hour a day, adding up to 24 by the end of the month. The lab tested the 24, not the daily routine. For the frozen model, every one of the 1,000 draws in every month had a signed error below zero: 24 labels were enough to see that it guessed too low. For January, 90% of the draws gave an MAE between 47.3 and 91.6, against the full month's 68.9: rough, but in the right place. The frozen model leaned so far that a few labels found it. The monthly model leaned much less. Its draws agreed with the sign of the whole month in 61% to 100% of cases. In May only 7% of the draws said "too low", and that is right: May's whole-month signed error was +16.8, too high. The draws split only where the month's own lean was small: April (+3.0, 61% agree) and June (-6.2, 71%). A small lean needs more labels.

Try It Yourself

This script is the lab made small. It downloads the same data, trains the frozen model once on 2011, and for five months of 2012 prints the frozen model's real error, psi_same and pred_shift side by side. It does not need a GPU.

A real screenshot of VS Code with watch_demo.py open, showing the top of the file: the docstring, the imports, the code that loads the data and numbers the months, the lists of input columns and the start of the frozen model. The rest of the file is in the code box on this slide. Beneath: copy it from the box on the slide.

Before you run this lab. You need Python 3 and three libraries: pip install scikit-learn pandas scipy. The demo itself needs only scikit-learn and pandas (its docstring lists scipy too); scipy comes with scikit-learn anyway, and the full lab and report use it for the rank correlation. scikit-learn holds the model and the download, and pandas holds the table. The first run downloads the Bike Sharing data from OpenML, so it needs an internet connection once; scikit-learn keeps a copy in a folder in your home directory (scikit_learn_data) for later runs. I ran it with scikit-learn 1.9.1 on a Mac. scikit-learn runs the same way on Windows and Linux, but I have not checked the numbers there, and another version of scikit-learn may give slightly different decimals.

"""Input drift, output drift and the real error, month by month, for a model trained on 2011.

This is the lab of lesson 5 of "Why Production Breaks", made small. It needs Python 3 with
scikit-learn, pandas and scipy (pip install scikit-learn pandas scipy). The first run downloads
the UCI Bike Sharing data from OpenML (under 1 MB) and keeps a copy for later runs.
    python watch_demo.py

Author: Roni Das
Created: 2026-09-29
"""
import numpy as np
from sklearn.compose import ColumnTransformer
from sklearn.datasets import fetch_openml
from sklearn.ensemble import HistGradientBoostingRegressor
from sklearn.metrics import mean_absolute_error
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import OrdinalEncoder

# 17,379 hours from 2011 and 2012, in time order. The answer is "count": bikes rented that hour.
data = fetch_openml("Bike_Sharing_Demand", version=2, as_frame=True, parser="auto").frame
X = data.drop(columns="count")
y = data["count"].astype(float)
month = X["year"].astype(int) * 12 + X["month"].astype(int)   # 1 to 12 = 2011, 13 to 24 = 2012

WORDS = ["season", "holiday", "workingday", "weather"]            # columns that hold words
NUMBERS = ["year", "month", "hour", "weekday", "temp", "feel_temp", "humidity", "windspeed"]
WEATHER = ["temp", "feel_temp", "humidity", "windspeed"]          # the columns psi watches
NAMES = ["Jan", "Feb", "Mar", "Apr", "May", "Jun", "Jul", "Aug", "Sep", "Oct", "Nov", "Dec"]

# The frozen model: trained once, on all of 2011, and never again.
columns = ColumnTransformer([
    ("words", OrdinalEncoder(handle_unknown="use_encoded_value", unknown_value=-1), WORDS),
    ("numbers", "passthrough", NUMBERS)])
train = month <= 12
frozen = make_pipeline(columns, HistGradientBoostingRegressor(random_state=0))
frozen.fit(X[train], y[train])

# psi cuts each weather column into 10 bands, each holding a tenth of 2011's hours.
bands = {}
for c in WEATHER:
    edges = np.unique(np.quantile(X[train][c].astype(float), np.linspace(0, 1, 11)))
    edges[0], edges[-1] = -np.inf, np.inf
    bands[c] = edges


def psi(before, now, edges):
    # population stability index: how far the share of hours in each band moved
    b = np.clip(np.histogram(before, edges)[0] / len(before), 1e-4, None)
    n = np.clip(np.histogram(now, edges)[0] / len(now), 1e-4, None)
    return float(((n - b) * np.log(n / b)).sum())


print("month      real error   input drift    output drift")
print("           frozen MAE   psi_same       pred_shift")
for m in [13, 15, 19, 21, 24]:                                    # Jan, Mar, Jul, Sep, Dec 2012
    now, last_year = X[month == m], X[month == m - 12]            # this month, same month of 2011
    mae = mean_absolute_error(y[month == m], frozen.predict(now))
    drift = max(psi(last_year[c].astype(float), now[c].astype(float), bands[c])
                for c in WEATHER)
    shift = frozen.predict(now).mean() - frozen.predict(last_year).mean()
    print(f"{NAMES[m - 13]} 2012   {mae:8.2f}     {drift:8.2f}       {shift:+8.1f}")

The Lab Report

A real terminal recording headed python watch_report.py, titled every table in this lesson, from the stored files. Eight numbered sections: each month of 2012 with both models' MAE and the four drift measures, and the eight rank correlations, declared before the run; what psi_year measures, with 2011's own months against 2011 beside 2012's, the 720-hour random draws and the whole-year psi per column; the same month in both years, with temperature, humidity, wind, share of clear hours, rentals, their ratio and the frozen MAE, and the rank correlation with the rental gap; output drift, the frozen model's mean guess against the real counts; psi_same read as an alarm at 0.25; and the real alarm, the newest labelled month and 24 labelled hours drawn at random, with the first day's signed error and the monthly model's agreement with each month's own sign; psi_same with its bands cut from all of 2011 and from the same month, with lesson 4's partial August; and the frozen signed error taken apart into pred_shift, the rider gap and the rest, with the share of hours too low. Every explanation line is marked measured or guess. Beneath: the lab's own report. It trains no model.

The report lives in scripts/labs/prodbreaks/watch_report.py. It reads the lab's stored file, results/watch.json, lesson 2's stored guesses in results/drift.json, lesson 4's batch-size follow-up in results/checks_psi_size.json and its same-month psi in results/checks_followup.json, and the Bike Sharing data itself from scikit-learn's local copy. It trains no model.

Before it prints anything, it checks the files against the data and against each other. Each month's MAE in watch.json must equal lesson 2's stored MAE for the same model and month. psi_year, psi_same and weather_mix are recomputed from the data and must match. The eight rank correlations are recomputed and must match. The one number it cannot recompute is pred_shift, because the model's guesses on 2011 were never stored; the report reads it from watch.json, cross-checks its 2012 half against lesson 2's guesses, and says so. If anything differs, the report stops. Every explanation it prints starts with MEASURED or GUESS.

Rank the Measures Yourself

This box has no model in it. It holds the lab's four drift measures for each month of 2012, the average rentals of each month of 2011, and every hour of 2012 with its real count and the stored guesses of both models (to two decimals). It runs in your browser.

As it is, the box prints the table from the results slide and the frozen model's four rank correlations: -0.34, +0.41, +0.39 and +0.58, the lab's numbers. The report's box mode checks every printed row and all eight correlations against watch.json; storing the guesses to two decimals moves a monthly MAE by at most 0.00017.

Try agreement(MONTHLY) for the monthly model: +0.05, +0.42, +0.06, +0.19. Try label_a_few(0) to score January from 24 random hours: with seed 0 it gives an MAE of 44.8 and a signed error of -42.7, against 68.9 for the whole month. That is an unusual draw, just below the 47.3 to 91.6 that held 90% of the report's draws (the box picks hours with Python's own random module, so its draws are not the report's). label_a_few(0, seed=1) gives 75.9 and -74.8. Different hours, the same direction. Then try the monthly model in May, label_a_few(4, guess=MONTHLY): 24 hours gave a signed error of +42.8 and n=100 gave +28.7, against +16.8 for the whole month: the right direction, the size still rough. The spearman function takes any two lists of 12 numbers, so you can build your own measure and rank it against [mae(FROZEN, m) for m in range(12)].

The Code, Part by Part

Loading and training. fetch_openml("Bike_Sharing_Demand", version=2, as_frame=True) downloads the data (or reads the local copy). The line month = X["year"].astype(int) * 12 + X["month"].astype(int) numbers the months 1 to 24, so "all of 2011" is month <= 12. The frozen model is the pipeline from the earlier lessons: an OrdinalEncoder turns the four word columns into number codes, and HistGradientBoostingRegressor(random_state=0) is the boosted trees. It is trained once, on 2011.

The bands. For each weather column, np.quantile finds the values that cut 2011's hours into ten equal groups. The first and last edges are set to minus and plus infinity, so every value lands in some band.

psi. psi(before, now, edges) counts the share of hours in each band for the reference and for this month, gives an empty band a share of 0.0001 so the logarithm is defined (the log, written np.log, cannot be taken of 0), and adds up the difference in shares times the log of their ratio. It is the lab's own function.

The loop. For five months of 2012, now is this month's rows and last_year the same month of 2011. The real error is mean_absolute_error of the frozen model's guesses. psi_same is the largest psi over the four columns, last_year against . pred_shift is the model's mean guess on minus its mean guess on .

How to Watch a Model Whose Answers Arrive Late

A hand-sketched column of six boxes joined by arrows, headed what held in this one lab, titled watching a model whose answers arrive late. 1, find out how long the labels take. 2, compare inputs with the same season, not the year. 3, treat input and output drift as a prompt to look. 4, label a random sample: 24 hours a month, tested. 5, score the newest labels: the MAE and its sign. 6, let the real error raise the alarm. Beneath: one dataset, two models, one year: a place to start, not a law.

Find out how long the labels take. Before you choose what to watch, measure the label delay. If labels arrive within an hour, as they do for bike rentals, score the real error every day and treat everything else as a detail. If they take months, you need something else in the meantime, and you need to know how long "in the meantime" is.

Compare inputs with the same season, not the whole year. A reference that mixes all seasons makes every month look like drift. Here psi_year fired on every month of 2012 and on every month of 2011, the model's own training year. Compare a July with past Julys, and check on known-good data how often your measure crosses its line.

Treat a drift number as a reason to look. When a drift number moves, someone should look at what moved and ask whether the model has seen anything like it. Here psi_same rose in a warm March and a clear September: real changes, and ones the model could mostly handle.

Label a small random sample, a little each day. Pay for it if you must. 24 labelled hours a month, a little under one a day, were enough to show the frozen model's lean in every draw, every month; the lab tested the 24, not the daily routine. The sample must be random: labelling only the cases that are easy to check gives a biased answer.

Score the newest labels: the MAE and its sign. As in lesson 2, compare with the error you accepted at launch, and read the sign.

Let the real error raise the alarm. Calling someone, retraining, going back to the old model: tie them to the real error on labelled data, not to a drift number alone.

When Drift Measures Help, and When They Do Not

Use input drift checks when labels are slow and inputs can break. Lesson 4 showed simple input checks catching unit changes, new labels, stuck sensors and missing values with few false alarms. Those are breaks in the inputs themselves, and there the inputs are exactly the right place to look. A value no training hour ever had is a real fault, so those checks can raise an alarm on their own. A shift in the mix of values, the drift this lesson measured, should not be the only alarm.

Use a same-season reference when your data has seasons. Weather, shopping, travel, school terms. Against a whole-year reference, a season looks like drift every time: here 1.34 to 4.94 in 2012, and 1.53 to 5.10 in the very year the model learned from.

Use output drift to see what the model believes changed. It is cheap and needs no labels. Here it showed the model raising its March guesses by 24.5 rentals an hour for a warmer March. Just remember that it can only move as far as the inputs push it.

Do not let input drift take the place of the error when the target itself can move. More customers, new prices, a competitor closing, a change in what people want. Here the error came from more riders, which none of the weather inputs carried, and the best rank correlation of any measure was 0.58 on 12 points.

Do not call people out on a drift line you have not tested on clean data. psi_same fired in 9 of 12 months, including the healthy model's best month. Measure the false alarms first, as lesson 4 did.

Do not skip labels because labels are expensive. A small random sample is cheaper than a year of a model that misses by 30% to 53% of each month's average, and here it gave the direction of the miss from 24 hours.

What These Numbers Can and Cannot Tell You

A two-column page headed read before you trust these numbers, titled what these runs are, and what they are not. They are: one dataset: bike rentals in one city; two models, default settings; 12 months: 12 points per correlation; measures and correlations fixed before the run; season, weather, labels chosen after; a drift that was mostly more riders. They are not: not a rule for every dataset; not tuned; tuning may move every score; not a significance test; not chosen after seeing results; not tests declared in advance; not a world whose inputs changed.

One dataset, one year. Hourly bike rentals in one city, and the twelve months of 2012. The drift here was mostly one kind: many more riders, with weather much like the year before. In data where the inputs themselves change a lot (a new kind of customer, a new sensor, a new region), input drift could track the error much better than it did here. I make no claim that it fails everywhere.

Two models, default settings. The frozen and monthly boosted trees of lesson 2, not tuned. Another model might react to the same months differently.

12 points per correlation. The design declared that 12 months are far too few for a significance test, so none was run. -0.34, 0.41, 0.39 and 0.58 describe this year; they are not estimates of anything general.

The design was declared; the reading was not. The two models, the four measures and the rank correlations were written in watch_lab.py before the run. 2011's own months, the random draws, the weather and rental comparisons, the rank link with the rental gap, the alarm reading and all the label sampling were chosen after I saw the results. The 24-hour size and the 1,000 draws were my choice then, not something I tested several values of.

The reasons are partly guesses. That the weather looked like 2011 over the year, and the riders did not, is measured. That the frozen model's error came from the growth is close to arithmetic: its signed error is the growth less its own weather response, give or take 2.2, and lesson 2 found it never used the year. Why pred_shift ranked the months at 0.58 is a guess I have not tested.

What to Do Next

A hand-drawn list headed before you trust a drift dashboard, titled five questions. Labels?: how long after a guess does its real answer arrive? Reference?: what does each drift number compare with: the whole year, or the same season? Size?: does a clean batch of this size stay under the line? Sample?: can someone label a small random sample, a few cases a day, before the rest arrive? Alarm?: does the real error, not a drift number, decide who is called? Beneath: here: psi_year 1.3 to 4.9 every month, while the error came from more riders.

Take a model that you or your team runs, and look at its drift dashboard, if it has one. For each number on it, find out what it compares with. If the answer is "the training data" and the data has seasons, run the same measure on a stretch of the training data itself, a month at a time, and count how often it crosses its line. That count is how often it will raise an alarm for nothing.

Then find out how long the labels take, and whether you could label a small random sample sooner. About one labelled case a day adds up to the 24 a month tested here, and even several a day is a small job for a person. Score the model on them as they arrive, read the sign, and let that number decide who is called.

The next lesson in this chapter looks at a harder case: what happens when the model's own guesses change the data it will later learn from, a feedback loop, as a clearly labelled simulation built on the same real data.

A closing card headed to keep, titled drift says where to look; the real error is the alarm. In large type: +0.58 at best. Beneath: the best rank correlation of any drift measure with the frozen model's monthly error, over 12 months of 2012. The error came from more riders: rentals were 1.41 to 2.53 times 2011's, while the weather looked like 2011. 24 labelled hours said 'too low' in every draw. Then: compare like with like, label a few cases early, and let the real error raise the alarm.

Knowledge Check

Knowledge Check

4 questions - Score 80% to pass

Q1

psi_year was above 0.25 in every month of 2012. What did the lab find it mostly measured?

Q2

Why could no input measure show the frozen model's main error?

Q3

On the morning of 1 February 2012, which reading showed which way the frozen model was wrong?

Q4

What does the lesson recommend doing with input and output drift numbers?

data drift
concept drift

The diagram plays one month of the lab as a service would run it. The city sends inputs; the model guesses; a monitor compares the inputs and the guesses with last year, with no labels needed; and the real counts arrive later, when the real error can finally be scored.

A two-column table headed each month of 2012, from the lab's stored file, titled the error and the four measures, month by month. Left, month: MAE, frozen; monthly; right, psi_year; psi_same; mix; shift. Twelve rows, January to December, for example Jan 2012: 68.9; 68.9, and 3.19; 0.73; 0.053; +6.4. Mar 2012: 110.7; 56.8, and 1.34; 1.25; 0.107; +24.5. Jul 2012: 93.6; 42.6, and 4.94; 0.08; 0.137; -5.0. Sep 2012: 109.4; 41.9, and 3.00; 1.12; 0.269; +22.0. Dec 2012: 61.6; 42.6, and 3.07; 0.12; 0.214; -4.7. Beneath: psi_year over 0.25 in 12 of 12 months. psi_same over 0.25 in 9 of 12.

The table shows every month. psi_year was above 0.25 in all 12 months, between 1.34 and 4.94, while the frozen model's error ranged from 61.6 to 110.7. psi_same was above the line in 9 of 12. The next slides take the measures one at a time and ask what each was really seeing.

That does not make a whole-year psi worthless; it changes how to read it. If you keep it, compare each month's value with the same month's value in a year you trust, not with 0.25. Here 2011's own months give that normal range: from 1.53 in April to 5.10 in July. A month that sat far outside its own season's usual value would be worth a look. That is advice from reading these tables, not something the lab tested. The simpler route is the one psi_same takes: build the season into the reference, so that 0.25 means something again.

Now put that gap next to the frozen model's error. The extra rentals an hour over the same month of 2011 and its monthly MAE have a rank correlation of 0.97. That is close to arithmetic, and it is worth seeing why (a review of this lesson pointed it out, and the report now prints it). The frozen model's signed error for a month, the average of guess minus real, splits into three parts: its pred_shift, minus the rider gap, plus how far its guesses for the same month of 2011 sat from 2011's real counts. That last part is its fit on its own training months, and it was never more than 2.2 in size. So the signed error is the model's own weather response less the growth. In March, +24.5 against a gap of 134.2 leaves -109.5. The weather response covered at most 18% of the growth, in March.

The signed error was between 87% and 99% of the MAE in size, always below zero, and 85.5% of all hours of 2012 were guessed too low (78% to 92% by month). A model whose misses are nearly all on one side is said to lean that way; this one leaned low. Against the ratio of the two years instead of the gap, the rank correlation is only 0.29: the MAE is counted in rentals, so it follows the size of the growth, not its rate. The growth lives in the labels. The one input that could have said "this is 2012", the year column, was 0 on every row the frozen model learned from, so, as lesson 2 measured, no tree could use it.

The word random matters. A sample made of the cases that are easiest to label, or the ones someone happened to look at, can lean in its own direction and tell you the wrong thing. Here the frozen model's lean was so large (85.5% of its hours too low) that almost any hours would show it, as the first days below do. Random picks matter for a smaller lean, like the monthly model's, where a biased handful could point the wrong way. In practice teams get early labels in a few common ways, none of which this lab tested: a person reviews a small random slice of cases each day (a support agent reads 20 tickets the model sorted); a quick early sign of the real outcome is used where one exists (a first missed payment long before a loan is written off); or a few customers are simply asked. Each costs something, and each is still cheaper than finding out a year late.

Even one day helps. The frozen model's signed error on just the first day of each month was below zero in 11 of 12 months; June's first day was the one exception, at +0.4. In this dataset labels arrive within the hour. In a system where they take weeks, 24 hours of labels means paying someone to label a small random sample now, or asking a few customers directly. That cost was not measured here.

This is a real run in VS Code's terminal (python watch_demo.py).

A real screenshot of VS Code's terminal after running python watch_demo.py. Two header lines: month, real error, frozen MAE; input drift, psi_same; output drift, pred_shift. Then five lines. Jan 2012: 68.89, 0.73, +6.4. Mar 2012: 110.68, 1.25, +24.5. Jul 2012: 93.59, 0.08, -5.0. Sep 2012: 109.41, 1.12, +22.0. Dec 2012: 61.63, 0.12, -4.7.

When I ran it, every printed number matched the lab's stored file, watch.json: the MAE (68.89, 110.68, 93.59, 109.41 and 61.63), psi_same (0.73, 1.25, 0.08, 1.12 and 0.12) and pred_shift (+6.4, +24.5, -5.0, +22.0 and -4.7). The report's demo mode checks this line by line. Look at July: psi_same 0.08, far under the line, while the model missed by 93.59 rentals an hour.

What came before the run, in watch_lab.py: the two models, the four measures, the monthly error and the rank correlations, and the statement that 12 points allow no significance test. What came after I saw the results: 2011's own months scored as a year, the month-sized random draws, the same-month weather and rentals, the rank link with the rental gap, the output drift beside the real change, psi_same read as an alarm, the newest labelled month, the first day and the 24 labelled hours. Two sections came later still, prompted by a review of this lesson: psi_same with its bands cut lesson 4's way, and the frozen signed error taken apart. The json mode writes the numbers to results/wa-report.json, which the figures read.

Five brand cards headed the tools, with their logos, titled what ran where. scikit-learn: the two models, every fit, the data download. pandas: the table of hours, split into months. NumPy: psi's bands and shares, the random draws. SciPy: the rank correlation, spearmanr. Python: the lab, the report, the demo, the box.

now
now
last_year

In the lab file, watch_lab.py does this for all twelve months and for the monthly model too, adds psi_year and weather_mix, and computes the rank correlations with scipy.stats.spearmanr.