Why Production Breaks

When the Data Answers Back: A Feedback Loop, Simulated on Real Demand

0 of 21 complete

0%

Contents

Back|Why Production BreaksWhen the Data Answers Back: A Feedback Loop, Simulated on Real Demand
1/21
65 min left
Prerequisites
Watching Inputs Before the Answers Arrive: Drift Measures Against the Real Errorrequired
Related Topics
Why ML Models Fail in Production: The Production GapCore ConceptsMonitoring and Drift Detection: Catching a Model That Fails Without an ErrorCore ConceptsHandling Imbalanced and Messy Data: Why Your 99% Accuracy Is a LieData Engineering for ML
1 of 21

The Baker Whose Notebook Was Always Right

Imagine a man who runs a small bakery on a busy street. Every evening he decides how many loaves to bake for the next morning. To decide, he opens his notebook, where he writes down how many loaves he sold each day. If the notebook says he sold about 60 on Tuesdays, he bakes about 60 for next Tuesday.

One spring, a new office opens across the road. On Tuesday morning a queue forms. By nine o'clock all 60 loaves are gone, and for the rest of the morning people come in, see the empty shelf, and walk out. That evening he writes in his notebook: sold 60. It is true. He did sell 60.

An illustration of a bearded man standing with his arms folded, in a shirt over a plain top, beside text. Headed a simulation on real bike-rental counts, titled every record was right. The riders who left are not in it. Beside him: bikes placed to match the forecast, 2012. Real demand, invented rule. The loop that learned from its own records (margin 1.0) wrote down 62.5% of the riders and turned away 769,320. Scored on its own records, it missed by 2.6 bikes an hour. Scored on the riders who really came: 90.7.

Next Tuesday he looks at the notebook, sees 60, and bakes 60 again. The shelf empties again, the same people walk out again, and the notebook says 60 again. Nothing in the notebook is wrong. Every number in it is a real sale. But the notebook can never say more than the number of loaves he baked, because nobody writes down the customer who found an empty shelf. His own decision about how much to bake has set a ceiling on what he can ever learn.

Now imagine a second baker on another street with the same notebook habit, except that she always bakes a few more loaves than the notebook says. Some mornings she throws a few away. But on the morning the office opens, her extra loaves sell, her notebook says 70, and the next week she bakes a little more. Her extra loaves are how she finds out that the street changed. This lesson measures both bakers, with bikes instead of bread.

Where This Lesson Starts

This is the sixth lesson of the chapter, and it uses the same public data and the same kind of program as the first five: two years of hourly bike rentals from Washington, D.C., and a program that guesses how many bikes will be rented in an hour. Lesson 2 showed that a program trained on 2011 and never updated fell far behind in 2012, because many more people rode, and that retraining it every month on the newest real counts kept it close. Lesson 5 showed that the real counts, even a few of them, were the best alarm.

Both lessons made one quiet assumption: that the real counts are there to learn from. In the bike data they are, because every rental was written down. But many programs do more than guess. They decide something, and the decision changes what gets written down. A guess about bike demand can decide how many bikes are put out, and nobody can rent a bike that is not there.

This lesson is a simulation, and I want to say that plainly before anything else. The hourly counts in the data are real, and I treat them as the true demand: the number of people who wanted a bike in that hour. The rule that turns a guess into bikes on the street is invented by me. No bike-share system in the data worked this way. Everything in this lesson measures what that invented rule does to a real program learning from real demand. The riders who were "turned away" are riders the simulation turned away, not riders who really left.

Eight Words for This Lesson

A hand-drawn list headed eight words for this lesson, titled when a model's decisions shape its own data. Feedback loop: the model's decisions change the data it later learns from. Demand: how many people wanted a bike in an hour; here, the real count. Supply: the bikes placed for that hour; here, set from the forecast. Recorded: the rentals written down: never more than the bikes placed. Censored: a record cut off at a limit: you only see up to what was there. Turned away: riders who wanted a bike and found none; nobody records them. Margin: extra bikes on top of the forecast: 1.2 means 20% more. Exploration: giving up a little now to see what you would otherwise never see. Beneath: the records can only show what the model's decision allowed to happen.

A bike-share system is a city service where people rent a bike from a street rack and return it later; the data in this chapter comes from one in Washington, D.C. The operator is whoever runs it and decides how many bikes to put out. A policy here is one fixed way of running the loop, set by one number, the margin; the lab compares four policies. A model is the program that guesses; here it guesses bikes rented in an hour, and its guess is called the forecast. A feedback loop is what happens when the model's decisions change the data it will later learn from: today's forecast shapes tomorrow's training rows.

Demand is how many people wanted a bike in an hour. Supply is how many bikes were there. The recorded rentals are what gets written down, and they can never be more than the supply, because a rider cannot rent a bike that is not there. When demand is bigger than supply, the extra riders are turned away, and nobody writes them down.

A record cut off at a limit like that is called censored. In plain words: you only see rentals up to the number of bikes that were there. If 568 people wanted a bike and 417 bikes were there, the record says 417, the same number it would say if exactly 417 people had wanted one.

A margin is extra supply on top of the forecast. A margin of 1.2 means 20% more bikes than the forecast. More generally, exploration means giving up a little now (some bikes stand unused) so that you can see what you would otherwise never see (whether more people would have come). The , from the earlier lessons, is the mean absolute error: the average size of the miss, in bikes an hour.

How the Loop Closes

The loop runs once a month in the simulation. The sketch draws it as a cycle of five boxes; the sequence diagram further down splits the same month into seven steps.

A hand-drawn sketch headed sketched: the loop, one month at a time, titled the forecast decides what can be recorded. Across the top, three boxes joined by arrows: the model's forecast; bikes placed: forecast x margin; riders take what is there. An arrow runs down from the third box to a box, the records: at most placed, which points left to a box, retrain before next month, which points back up to the forecast. Under the records box: riders who found no bike: never written down. Under the retrain box: no new real rider numbers enter this loop. Beneath: in the lab the real hourly counts play the riders, and the rule for placing bikes is invented. The control model skips the cap and learns from every rider.

The model makes a forecast for every hour of the coming month. The operator places bikes to match: the forecast times the margin, rounded up to a whole bike. Riders come, and each one who finds a bike rents it. The rentals are written down. Before the next month, the model is retrained on everything recorded so far, and the loop begins again.

Look at what enters the loop from the outside world. The true 2011 rows stay in training, but nothing new about real demand comes in: only the recorded rentals. And the recorded rentals are capped by the bikes placed, which were set by the forecast. So the model is, in part, learning from its own earlier guesses. If it guessed too low, the records come back too low, and it learns that the low guess was right.

A sequence diagram with four columns: the model, the operator, the street, the records. Step 1, the model sends the operator a forecast, each hour. Step 2, the operator sends the street the bikes placed. Step 3, on the street, riders arrive. Step 4, the street sends the records the rentals, capped. Step 5, on the street, the rest leave. Step 6, the records send the model this month's records. Step 7, the model retrains. Headed one month of the simulation, step by step, titled the model learns from records its own forecast capped. Beneath: step 5 leaves no trace in the records. The lab knows how many left only because it holds the real counts, which no operator has.

The sequence diagram lays out the same month step by step. Step 5 is the important one, and it is the one nobody sees: the riders who find no bike simply leave. In a real system there is no row for them. The lab can count them only because it holds the real hourly counts from the data, which in the simulation play the part of what riders wanted. A real operator would have only the records.

What the Lab Ran

I wrote the design at the top of the lab file, loop_lab.py, before I ran it, and the chapter plan, CHAPTER-PLAN.md, records it the same day ("LOOP LAB (batch 6, a SIMULATION on real demand) designed before running").

A flowchart headed what the lab ran, titled four policies, twelve months, one invented rule. A box, all of 2011, true demand: the start for every policy, leads to control, loop 1.0, loop 1.2, loop 1.5, which leads to each month of 2012: forecast, place, record. That box leads to two boxes: retrain on the records before the next month, and score against true demand: MAE. Beneath: designed before it ran, in loop_lab.py: the same boosted trees as lessons 1 to 5, default settings, random_state 0. One run per policy over one year: a description, no significance test.

The model is the chapter's: scikit-learn's boosted trees (HistGradientBoostingRegressor, default settings, random_state=0), many small decision trees that each ask yes-or-no questions about the inputs and each correct the ones before. Every policy starts from the same place: a model trained on all of 2011 with the true counts. Then, before each month of 2012, each policy's model is retrained on every earlier hour.

The four policies differ only in what the 2012 rows they have already served hold. The control keeps the true demand, as if every rider had been recorded; it is exactly lesson 2's monthly model, and the report checks that its errors match lesson 2's stored ones. Its mean monthly MAE is 43.6 (lesson 2 quoted 43.7, its error over all hours). Loop 1.0 keeps what was recorded when bikes were placed at the forecast with no margin. Loop 1.2 and loop 1.5 do the same with 20% and 50% more bikes.

An editorial page in four labelled zones, headed what is real and what is invented, titled real riders, an invented rule for bikes. True demand, real: the real count of bikes rented in each hour of 2012, from the Bike Sharing data: 2,049,576 rentals over 8,734 hours. The forecast, the model: the policy's own boosted trees, retrained before each month on everything before it. Bikes placed, invented: the forecast times the margin, rounded up to a whole bike. Margins 1.0, 1.2 and 1.5. Rentals recorded, invented: the smaller of true demand and the bikes placed. The next month's model learns from this. Beneath: only the first zone is data. The control learns from true demand, as if every rider had been recorded.

What a Censored Record Looks Like

Before the results, it helps to see the censoring on a real day. This is Monday 2 July 2012 under loop 1.0, the loop with no margin, at three hours.

A hand-sketched column of three boxes, headed loop 1.0, Monday 2 July 2012, three real hours, titled when bikes ran out, the record says exactly the number placed. Hour 3: 5 riders wanted a bike; placed 13; recorded 5; 8 bikes not used; nobody turned away. Hour 8: 568 riders wanted a bike; placed 417; recorded 417; 151 turned away, not in the record. Hour 17: 747 riders wanted a bike; placed 544; recorded 544; 203 turned away, not in the record. Beneath: nothing in the records of hour 8 or hour 17 says more riders came. Loop 1.2 placed 617 and 852 and recorded all 568 and 747.

At 3 in the morning, 5 people wanted a bike and 13 were there. All 5 rode, and the record says 5: the truth. At 8 in the morning, 568 people wanted a bike, the loop had placed 417, and the record says 417. The 151 who found an empty rack are not in it. At 5 in the afternoon, 747 wanted a bike, 544 were placed, and the record says 544.

Now read those records the way the next model will. Hour 8: 417 rentals, and the model had guessed 416.1. Hour 17: 544 rentals, and the model had guessed 543.7. The guesses look almost perfect. The records hold no sign that anything was wrong, because the only thing that could show it, the riders who left, was never written down.

Loop 1.2, with 20% more bikes, had its own forecast for the same hours (the loops' forecasts drift apart over the months) and placed 617 at hour 8 and 852 at hour 17. Both were more than the demand, so its records hold the true 568 and 747. That is what a margin gives the model: the hours where the record is not cut off are hours where the model can learn the truth.

The Loop Missed by Twice as Much

Here is the headline, from the lab's stored file, results/loop.json. Every number is scored against the true demand.

A bar chart headed mean monthly MAE over 2012, against true demand (lower is better), titled learning from its own records: 90.7; from every rider: 43.6. Four bars on a scale up to 100: control about 44, loop 1.0 about 91, loop 1.2 about 60, loop 1.5 about 46, with a dashed line near 89.5 labelled lesson 2's frozen model, 89.5. Beneath: control 43.6, loop 1.0 90.7, loop 1.2 60.0, loop 1.5 46.4. Riders recorded: loop 1.0 62.5%, loop 1.2 87.6%, loop 1.5 97.1%. One run each.

Over the 12 months of 2012, the control missed true demand by 43.6 bikes an hour on average. Loop 1.0, the same model retrained just as often but on its own capped records, missed by 90.7: more than twice as much. It recorded 62.5% of the 2,049,576 riders who wanted a bike, and turned away 769,320. Demand was bigger than supply in 84.7% of its hours.

A margin helped a lot. Loop 1.2 missed by 60.0, recorded 87.6% of riders and turned away 254,040. Loop 1.5 missed by 46.4, close to the control, recorded 97.1% and turned away 59,804. These are single runs over one year, so they describe this simulation; they are not estimates of what a margin does in general.

The dashed line on the chart is lesson 2's frozen model, trained on 2011 and never updated, at 89.5 (its mean of the 12 monthly errors; its error over all hours of 2012 also rounds to 89.5). Loop 1.0 was retrained twelve times and ended up almost exactly where a model that never saw 2012 was. The next slides ask why.

A line chart headed MAE against true demand, each month of 2012, titled loop 1.0 never caught up; loop 1.2 closed most of the gap by June. Four lines over January to December on a scale up to 120, all starting at about 69 in January. Loop 1.0 stays high, between about 63 and 111, peaking in March, September and October. Loop 1.2, dotted, rises to about 99 in March, then falls to about 48 in June and stays between about 37 and 56. Loop 1.5 falls to about 41 by April and then runs close to the control, dashed, which drops to about 29 in February and stays between about 35 and 57. Beneath: January is 68.9 for all four: the same model trained on 2011. After it, loop 1.0 runs from 63.3 to 111.2; the control from 28.7 to 56.8.

Month by month the picture has three shapes. January is 68.9 for all four, because all four used the same model, trained on 2011, before any 2012 record existed. After that, loop 1.0 never came close to the control. Loop 1.2 stayed far behind until April, closed most of the gap by June, and from then on stayed within 15 bikes an hour of it (5.4 to 14.5 worse, and 5.6 better in December). Loop 1.5 reached the control by April.

Why Nobody Would Notice

The lab could score every policy against the true demand, because it holds the real counts. A real operator could not. The operator has only the records. So I asked, after the results, what the error looks like when each model is scored the only way an operator could score it: against the rentals it recorded.

Two panels headed loop 1.0, the same forecasts scored two ways, added after the results, titled scored on its own records, it looked nearly perfect. Against the rentals it recorded: 2.6: mean monthly MAE, bikes an hour: what the operator can compute. Against the riders who came: 90.7: mean monthly MAE, bikes an hour: only the lab can compute it. Beneath: the records hit the bikes placed in 86.1% of hours, so a forecast of about the placed number looks right almost every time.

Scored on its own records, loop 1.0 missed by 2.6 bikes an hour. Scored on the riders who came, it missed by 90.7. It is the same model and the same forecasts; only what they were compared with changed. In 86.1% of hours the record equals the bikes placed, and the bikes placed are the forecast rounded up. So the record is, most of the time, just the forecast again, and a forecast compared with itself looks excellent.

A feedback loop is invisible in the recorded data. Every recorded number is right. None of them is a mistake that a data check could find, and none of them is missing in a way that shows up as an empty cell. The error lives entirely in what was never written down.

A two-column table headed all four policies, each scored two ways, added after the results, titled the records rank the policies in exactly the wrong order. Left, policy: MAE on its records. Right, MAE on true demand; hours that ran out. Control: 43.6, and 43.6; no cap. Loop 1.0: 2.6, and 90.7; 86.1%. Loop 1.2: 30.9, and 60.0; 49.4%. Loop 1.5: 39.5, and 46.4; 17.4%. Beneath: by the records, best first: loop 1.0, loop 1.2, loop 1.5, control. By true demand: control, loop 1.5, loop 1.2, loop 1.0.

Now compare the policies. Ranked by the error against their own records, the four policies come out in the order loop 1.0 (2.6), loop 1.2 (30.9), loop 1.5 (39.5), control (43.6). Ranked by the error against true demand, the order is control, loop 1.5, loop 1.2, loop 1.0: exactly reversed. A team that picked its policy by the error on its own records would have picked the worst one here. That is my reading of these four numbers, not a test of how teams choose.

A line chart headed rentals an hour, each month of 2012, added after the results, titled loop 1.0's records grew too, so they looked healthy. Three lines over January to December on a scale up to 350: true demand, dashed, rises from about 131 in January to about 304 in September and falls to about 167 in December; loop 1.2 recorded, dotted, starts near 75 and rises to follow just under true demand from June; loop 1.0 recorded starts near 63, rises to about 190 from May to September and falls to about 108 in December. Beneath: loop 1.0 recorded 62.9 an hour in January and 194.4 at its highest; true demand ran from 130.6 to 303.6. Seen alone, the records show a business growing.

Why Loop 1.0 Stayed Stuck

Loop 1.0 was retrained every month on the newest data, like the control. Yet it ended where lesson 2's frozen model was. I compared the two month by month, after the results, using lesson 2's stored guesses (results/drift.json).

A line chart headed MAE against true demand, added after the results, titled loop 1.0 stayed where a model that never saw 2012 was. Two lines over January to December on a scale up to 120 that lie almost on top of each other: lesson 2's frozen model, trained on 2011 only, dashed, and loop 1.0, retrained every month. Both start near 69 in January, rise to about 111 in March, dip to about 79 in May, rise to about 111 in September and October, and fall to about 62 in December. Beneath: loop 1.0 minus frozen, each month: 0.0 to +2.8. From February their forecasts for the same hour differ by 4.3 to 9.0 bikes on average.

The two lines lie almost on top of each other. In every month, loop 1.0's error was between 0.0 and 2.8 bikes an hour above the frozen model's. Their forecasts for the same hour differed by 4.3 to 9.0 bikes on average from February on, and over 2012 the frozen model's forecasts averaged 149.7 an hour and loop 1.0's own forecasts 148.2, while true demand averaged 234.7 (loop 1.0's records averaged 146.6). All of that is measured.

Why? The measured part: in all 8,734 hours of 2012, loop 1.0 never recorded more rentals than its own forecast rounded up (the report counts 0 such hours; the rule makes it impossible). So every 2012 row the loop learned from said "demand was at most what I guessed". My explanation, which the report labels a guess: a model retrained on rows like that is taught its own old level again. The only true rows it ever saw were from 2011, and 2011 was quieter, so they pulled its guesses down, not up. Retraining could not help, because the new rows only repeated what the model already believed.

This is the same result as lesson 2's frozen model, reached a different way. The frozen model never saw 2012 because nobody retrained it. Loop 1.0 saw 2012 every month, but only the part of 2012 its own forecast let through. Retraining on the wrong data is, here, no better than not retraining.

How a 20% Margin Let the Model Climb

Loop 1.2 also started far behind. It then closed most of the gap by June. I read the month-by-month data after the results to see when and why. To compare forecasts with demand in one number, I use the forecast level: the mean forecast of a month divided by the mean true demand. 1.00 means the forecasts were right on average; 0.60 means they were 40% too low.

A line chart headed mean forecast divided by mean true demand, each month, added after the results, titled loop 1.2 climbed; loop 1.0 stayed at about two thirds. Four lines over January to December on a scale from 0 to 1.2, with a dashed line at 1.0 labelled forecast = demand at its left end. All start at about 0.48 in January. The control, dashed, jumps to about 0.94 in February, dips to 0.78 in March, then stays near 1.0. Loop 1.5 climbs to about 0.96 by April and stays near 1.0. Loop 1.2 climbs slowly, to about 0.67 in April, 0.80 in May and 0.90 in June, then stays near 0.9 to 0.97. Loop 1.0, dotted, rises to about 0.73 in May and then settles near 0.62 to 0.67. Beneath: loop 1.2: 0.56 in March, 0.80 in May, 0.90 in June. Loop 1.0: 0.73 in May, 0.66 in December.

In January every policy's forecast level was 0.48: the 2011 model guessed about half of 2012's demand. The control, which learned January's true counts, jumped to 0.94 in February. Loop 1.0 rose slowly with the seasons, like the frozen model, to 0.73 in May, and fell back to about two thirds. Loop 1.2 climbed a little each month, 0.53, 0.56, 0.67, then 0.80 in May and 0.90 in June.

A table of four rows, headed loop 1.2 in spring, measured; the link between the rows is my reading, titled loop 1.2 closed most of the gap in May, and more by June. Mar 2012: forecast level 0.56; short of bikes in 81% of hours; recorded 67% of riders; MAE 99.4, control 56.8. Apr 2012: forecast level 0.67; short of bikes in 69% of hours; recorded 77% of riders; MAE 87.6, control 39.6. May 2012: forecast level 0.80; short of bikes in 50% of hours; recorded 91% of riders; MAE 62.0, control 42.9. Jun 2012: forecast level 0.90; short of bikes in 30% of hours; recorded 96% of riders; MAE 48.1, control 37.7. Beneath: the gap to the control fell most into May (+48.1 to +19.1), under a model trained before any month had 90% of riders recorded. From June it was at most +14.5.

The measured chain for loop 1.2 is this. In March its forecast level was 0.56, it was short of bikes in 81% of hours, and it recorded 67% of the riders. In April, 0.67, short in 69% of hours, 77% recorded. In May, 0.80, short in 50% of hours, and 91% of the riders recorded: the first month over 90%. In June, trained on those May records among the rest, its level was 0.90 and its MAE 48.1, against the control's 37.7. From June on, its gap to the control was at most 14.5 bikes an hour; from February to May it had been 43.2, 42.6, 48.1 and 19.1. In December it was even 5.6 below the control.

What the Margin Cost

A margin is not free. Bikes placed for riders who never come stand idle. I counted them after the results from the stored hourly values: for every hour, the bikes placed minus the rentals recorded.

An isometric drawing of six blocks in three pairs, headed over 2012, thousands, heights to scale, added after the results, titled a margin trades riders turned away for bikes that stand idle. 1.0 idle, 19; 1.0 away, 769; 1.2 idle, 197; 1.2 away, 254; 1.5 idle, 820; 1.5 away, 60. The 1.0 away and 1.5 idle blocks are the tallest. Beneath: idle = bikes placed but not rented, counted each hour. Loop 1.0: 19,123 idle, 769,320 turned away; loop 1.2: 196,770 idle, 254,040 turned away; loop 1.5: 820,304 idle, 59,804 turned away.

Loop 1.0 placed 1,299,379 bikes over the year's hours and 19,123 of them were not rented: 1.5%. It wasted almost nothing, and turned away 769,320 riders. Loop 1.2 placed 1,992,306 and left 196,770 idle, 9.9%. Loop 1.5 placed 2,810,076, more than the 2,049,576 riders who wanted one, and left 820,304 idle, 29.2%, while turning away 59,804.

There is an important caveat, and it comes from the invented rule. The rule places bikes hour by hour, as if each hour were a new day. A bike unused at 3 in the morning is counted as idle at 3 and placed again at 4. A real bike-share system would not work like that: the bike would still be in the rack. So "idle bike-hours" is a cost unit of this simulation, not a count of bikes a real operator would have bought.

Two panels headed what each step up in margin bought, over 2012, added after the results, titled the first 20% was cheap; the next 30% was not. From 1.0 to 1.2: 0.34 idle: bike-hours per extra rider served: 515,280 more riders, 177,647 more idle. From 1.2 to 1.5: 3.21 idle: bike-hours per extra rider served: 194,236 more riders, 623,534 more idle. Beneath: the rule places bikes hour by hour, so a bike left unused this hour is counted again next hour. In a real bike-share system it would still be in the rack.

The trade is uneven. Going from no margin to 20% served 515,280 more riders for 177,647 more idle bike-hours: 0.34 idle bike-hours for each extra rider. Going from 20% to 50% served 194,236 more for 623,534 more idle: 3.21 for each extra rider, almost ten times the price. In this simulation, the first bit of margin did most of the work. Which margin is "right" depends on what an idle bike and a lost rider cost, which the lab did not measure.

A line chart headed riders turned away over 2012, by hour of the day, added after the results, titled the loop failed riders most at the rush hours. Three lines over hours 0 to 23 on a scale up to 80 thousand: loop 1.0 has two tall peaks, about 72 at hour 8 and about 78 at hour 17, and stays above about 28 through the middle of the day; loop 1.2, dotted, has the same shape at about a third of the height, peaking near 32 and 25; loop 1.5 stays under about 8 all day. Beneath: loop 1.0: hour 17, 78,194; hour 8, 71,700. Loop 1.5 at the same hours: 4,211 and 7,739. The busiest hours had the most riders to lose.

The Same Shape in Other Systems

Bikes are a small example of a common shape. What follows is general knowledge about how such systems are built, not something this lab measured, and I make no claim about any particular company.

An editorial page in four labelled zones, headed the same shape elsewhere: general knowledge, not measured here, titled a decision limits what gets recorded. Recommendations: people can only click what they were shown. The next model learns that what it showed is what people like. Fraud rules: a blocked payment never completes, so nobody learns whether it was fraud. Blocked cases leave the data. Credit decisions: a refused loan is never repaid or missed. The lender learns outcomes only for people it approved. Stock ordering: a shop records no sales once a shelf is empty. Missing sales look like low demand next time. Beneath: in each case the records are true and incomplete, and the missing part is the part the decision ruled out.

Recommendations. A shop's website or a video app shows each person a short list chosen by a model. People can only click what they were shown. The clicks are recorded, the next model is trained on them, and it learns that the things it showed are the things people like. Things it never showed never get a click, so it never learns they were good. This is the bike loop with "shown" in place of "bikes placed".

Fraud rules. A rule or model blocks payments that look risky. A blocked payment never completes, so nobody finds out whether it was really fraud. If the next model is trained only on completed payments, the cases the old rule blocked are missing, and the new model knows least about exactly the kind of payment the old one was worried about.

Credit decisions. A lender approves some loan applications and refuses others. It learns who repaid and who did not only for the people it approved. A refused applicant who would have repaid leaves no record of it. A model trained on approved loans alone learns about a narrower group of people than it will be asked about.

Stock ordering. A shop orders stock from a forecast of sales. When a shelf runs empty, sales stop, and the sales record shows a quiet afternoon rather than a busy one. The next forecast, trained on those sales, orders less. This one is the closest to the bike simulation, and it is the baker from the first slide.

In all four, every recorded number is true. What is missing is the outcome of the choice the system did not make. That is why the loop cannot be found by checking the data that exists; it has to be found by asking what data the decision prevented.

Try It Yourself

This script is the lab made small. It downloads the same data, and for January to June 2012 it runs the control and loop 1.0 side by side: it trains both, places bikes at loop 1.0's forecast, records the capped rentals, and feeds them into the next month's training. It prints both models' real error, loop 1.0's error on its own records, and the riders it turned away. It does not need a GPU.

A real screenshot of VS Code with loop_demo.py open, showing the top of the file: the docstring that says it is a simulation, the imports, the code that loads the data and numbers the months, the lists of input columns and the start of the model function. The rest of the file is in the code box on this slide. Beneath: copy it from the box on the slide.

Before you run this lab. You need Python 3 and two libraries: pip install scikit-learn pandas. scikit-learn holds the model and the download, and pandas holds the table. The first run downloads the Bike Sharing data from OpenML, so it needs an internet connection once; scikit-learn keeps a copy in a folder in your home directory (scikit_learn_data) for later runs. I ran it with scikit-learn 1.9.1 on a Mac. scikit-learn runs the same way on Windows and Linux, but I have not checked the numbers there, and another version of scikit-learn may give slightly different decimals.

"""A feedback loop, simulated on real bike-rental counts: a model that learns from its own records.

This is the lab of lesson 6 of "Why Production Breaks", made small. It is a SIMULATION: the real
hourly counts are treated as the true demand, and the rule for placing bikes is invented. It needs
Python 3 with scikit-learn and pandas (pip install scikit-learn pandas). The first run downloads
the UCI Bike Sharing data from OpenML (under 1 MB) and keeps a copy for later runs.
    python loop_demo.py

Author: Roni Das
Created: 2026-09-29
"""
import numpy as np
from sklearn.compose import ColumnTransformer
from sklearn.datasets import fetch_openml
from sklearn.ensemble import HistGradientBoostingRegressor
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import OrdinalEncoder

# 17,379 hours from 2011 and 2012, in time order. "count" = bikes rented that hour.
data = fetch_openml("Bike_Sharing_Demand", version=2, as_frame=True, parser="auto").frame
X = data.drop(columns="count")
demand = data["count"].to_numpy(dtype=float)        # treated as the TRUE demand
month = (X["year"].astype(int) * 12 + X["month"].astype(int)).to_numpy()  # 13 = Jan 2012

WORDS = ["season", "holiday", "workingday", "weather"]
NUMBERS = ["year", "month", "hour", "weekday", "temp", "feel_temp", "humidity", "windspeed"]
NAMES = ["Jan", "Feb", "Mar", "Apr", "May", "Jun", "Jul", "Aug", "Sep", "Oct", "Nov", "Dec"]


def new_model():
    # the lab's model: scikit-learn's boosted trees, default settings, random_state 0
    columns = ColumnTransformer([
        ("words", OrdinalEncoder(handle_unknown="use_encoded_value", unknown_value=-1), WORDS),
        ("numbers", "passthrough", NUMBERS)])
    return make_pipeline(columns, HistGradientBoostingRegressor(random_state=0))


control = demand.copy()   # what the control model learns from: every rider
loop = demand.copy()      # what the loop learns from: 2012 rows become its own records

print("month        control   loop 1.0   loop 1.0 on   riders turned")
print("                  MAE        MAE   its records     away (loop)")
for m in range(13, 19):                                  # January to June 2012
    before, now = month < m, month == m
    c = new_model().fit(X[before], control[before]).predict(X[now])
    f = new_model().fit(X[before], loop[before]).predict(X[now])
    bikes = np.ceil(np.clip(f, 0, None) * 1.0)           # bikes placed: the forecast
    seen = np.minimum(demand[now], bikes)                # riders can only take bikes that are there
    loop[now] = seen                                     # next month, the loop learns from this
    true_c = np.abs(c - demand[now]).mean()
    true_f = np.abs(f - demand[now]).mean()
    own = np.abs(f - seen).mean()                        # the error the operator would see
    away = np.clip(demand[now] - bikes, 0, None).sum()
    print(f"{NAMES[m - 13]} 2012  {true_c:9.2f}  {true_f:9.2f}  {own:12.2f}  {away:14,.0f}")

The Lab Report

A real terminal recording headed python loop_report.py, titled every table in this lesson, from the stored files. Seven numbered sections: each month of 2012 scored against true demand, with the MAE of all four policies, the share of short hours and loop 1.0's riders turned away, declared before the run; why loop 1.0 stays stuck, beside lesson 2's frozen model; when loop 1.2 recovers, with the forecast level, short hours, recorded share and gap to the control each month; what the margin costs, bikes placed, rented, idle and turned away; what the operator could see, the error against the records and the hours that ran out; riders turned away by hour of the day; and one real day, three hours. Section 3 also splits July's fewer riders turned away into the part from the extra bikes on the day and the part from what the model learned. Every explanation line is marked measured or guess. Beneath: the lab's own report. It trains no model.

The report lives in scripts/labs/prodbreaks/loop_report.py. It reads the lab's stored file, results/loop.json, the hourly rerun in results/loop_detail.json, lesson 2's stored guesses in results/drift.json, and the Bike Sharing data itself from scikit-learn's local copy. It trains no model.

Before it prints anything, it checks the files against the data and against each other. The stored demand must be the data's 2012 counts. Every monthly MAE, signed error, share of short hours, riders turned away and rentals recorded is recomputed from the hourly file and must match loop.json (the error to 0.001, because the hourly forecasts are stored to four decimals; the counts exactly). Every recorded count must be the smaller of demand and bikes placed, and the control must equal lesson 2's monthly model. If anything differs, the report stops. Every explanation it prints starts with MEASURED or GUESS.

What came before the run, in loop_lab.py: the four policies, the rule, the monthly retraining, the error against true demand, the short hours and the riders turned away, and the statement that one run per policy allows no significance test. What came after I saw the results: the rerun, the comparison with the frozen model, the forecast level, the idle bikes and the price per rider, the error against the records and the hours that ran out, the hour-of-day view, the example day and the July split of the margin's benefit. The mode writes the numbers to , which the figures read.

Run the Loop's Numbers Yourself

This box has no model in it. It holds, for every hour of 2012, the true demand, the hour of the day, and for loops 1.0 and 1.2 the stored forecast (to two decimals) and the bikes placed, plus the lab's monthly MAE for all four policies. It runs in your browser.

As it is, the box prints loop 1.0's year month by month: the MAE against true demand (68.9 in January up to 111.2 in March and September), the MAE against its own records (1.1 to 4.5), the riders turned away, the idle bike-hours and the share of hours that ran out (80% to 92%). The report's box mode checks every month's riders turned away, idle bikes and ran-out share against the stored files, and every printed MAE against loop.json.

Try table('loop_1.2') to watch the change: the share of hours that ran out falls from 88% in January to 31% in June, and the idle bike-hours rise from 813 to 23,728. Try compare() for all four policies side by side, and by_hour('loop_1.2') for the hour-of-day view.

Then try offer. It takes one month's stored forecasts and places bikes with a different margin, without retraining anything, so it is not the loop; it shows only the cost of that month's decision. offer('loop_1.0', 1.2, 6) puts a 20% margin on loop 1.0's July forecasts: 42,740 riders turned away and 2,967 idle bike-hours. Loop 1.2 itself, whose forecasts had climbed through the loop, turned away 10,193 that July. The report does the same sum from the unrounded forecasts: of the 58,074 fewer riders turned away in July, 32,544 (a little more than half) came through what the model learned and 25,530 through the extra bikes on the day. That is one month of one run. Because offer rounds up the stored two-decimal forecasts, at a loop's own margin it can differ from the lab's count by a few riders: at most 5 in a month, which the report measures.

The Code, Part by Part

Loading. fetch_openml("Bike_Sharing_Demand", version=2, as_frame=True) downloads the data (or reads the local copy). demand holds the real hourly counts, which the simulation treats as the true demand. The line month = X["year"].astype(int) * 12 + X["month"].astype(int) numbers the months 1 to 24, so January 2012 is 13.

The model. new_model() is the pipeline from the earlier lessons: an OrdinalEncoder turns the four word columns (season, holiday, working day, weather) into number codes, the eight number columns pass through, and HistGradientBoostingRegressor(random_state=0) is the boosted trees.

Two copies of the answers. control and loop both start as copies of the true demand. The control's copy is never changed. The loop's copy gets its 2012 rows replaced, month by month, with what the loop recorded. That one difference is the whole experiment.

The month. For each month from January to June 2012, both models are trained on every earlier row of their own copy and forecast the month. bikes = np.ceil(np.clip(f, 0, None) * 1.0) places bikes at the forecast, rounded up and never below zero; change 1.0 to 1.2 to add a margin. seen = np.minimum(demand[now], bikes) is the censoring: riders can only take bikes that are there. writes the records into the loop's copy, so the next month's model learns from them.

How to Run a Model Whose Decisions Shape Its Data

A hand-sketched column of six boxes joined by arrows, headed what held in this one simulation, titled running a model whose decisions shape its data. 1, log what was offered, not only what was taken. 2, count how often the offer ran out. 3, keep a margin or a small slice outside the model. 4, measure the demand you turned away. 5, score on the real outcome, never the records alone. 6, keep one control that learns from full data. Beneath: one simulation, one model, one year: a first check, not a rule for every system.

Log what was offered, not only what was taken. Write down the bikes placed, the items shown, the payments allowed, the stock ordered. Without the offer, a record of 417 rentals cannot be told apart from a quiet hour. With it, you can see that 417 were placed and 417 went.

Count how often the offer ran out. The share of hours when every bike went needs no knowledge of demand. Here it was 86.1% for the loop with no margin. If most of your records sit exactly at the limit your own decision set, your data is censored most of the time.

Keep a margin or a small slice outside the model. Some supply above the forecast, or a small random share of decisions made without the model, keeps some records uncut. Here a 20% margin took the yearly error from 90.7 to 60.0.

Measure the demand you turned away, where you can. Some systems can: a rider who opens the app and sees no bike, a search that found nothing in stock, a queue that grew. Those are signs of the riders who left, and they are worth logging even when they are rough.

Score on the real outcome, never on the records alone. An error computed against censored records can look perfect: 2.6 here, against 90.7 on the truth.

Keep one control that learns from full data. Where you can get the full outcome for a slice (a random sample of decisions made generously, or a region where supply is never short), train or score a model on it too. If the two drift apart, the loop is doing it.

When a Margin or Exploration Helps, and When It Does Not

Use a margin or exploration when the model's decision limits what gets recorded. Supply set by a forecast, items shown by a ranking, applications approved by a score. Here, without a margin, the model learned almost nothing the frozen model did not already know (0.0 to 2.8 bikes an hour worse each month).

Use it early, when the model is most likely to be wrong. A new model, a new market, a change in the world. Here every policy started at half of 2012's demand, and in this run, as far as I can tell, the margin was what let loop 1.2 close most of the gap by June.

Use a small amount first. In this simulation, going from no margin to 20% cost 0.34 idle bike-hours per extra rider served, and going from 20% to 50% cost 3.21. The first step bought the most.

You do not need it when the decision does not limit the record. If every rider is counted whether or not a bike was there (for example, because riders book ahead and every booking is logged), the records hold the full demand and the model can learn it directly, as the control did.

You do not need it when you already log the demand you turned away. If the lost riders are counted, the censoring can be undone in the data, and the model can learn from demand rather than from records.

Do not judge a margin by the idle cost alone. Loop 1.0 had the lowest idle share, 1.5%, and the worst error and the most riders turned away. Low waste can be the sign of a loop that stopped learning.

What These Numbers Can and Cannot Tell You

A two-column page headed read before you trust these numbers, titled what these runs are, and what they are not. They are: a simulation: the rule is invented; real demand: 2012's hourly counts; one model, default settings; one run per policy, one year; policies and scores fixed before the run; idle bikes, levels, records view: after. They are not: not a real operator's rule; not riders who react to empty racks; not tuned; tuning may move every score; not a significance test; not chosen after seeing results; not tests declared in advance.

It is a simulation. The real hourly counts are the demand, and the rule for placing bikes is mine. A real bike-share system moves bikes between stations, keeps unused bikes in the rack, and has riders who walk to the next station. None of that is here. The numbers describe what this rule does to this model on this demand.

The demand does not react. In the simulation, a rider turned away at 8 in the morning does not come back at 9, and does not give up on bike-share next week. Real riders might. That could make a real loop worse (demand shrinks to match the supply) or milder (riders wait), and the lab cannot say which.

One model, default settings, one run per policy, one year. The boosted trees of the earlier lessons, not tuned, each policy run once. The design declared no significance test, so 90.7, 60.0 and 46.4 describe this run. A different model, or a different start, might climb faster or slower.

The design was declared; the reading was not. The four policies, the rule and the scores were written in loop_lab.py before the run. The hourly rerun, the comparison with the frozen model, the forecast level, the idle bikes, the error against the records, the hours that ran out and the example day were chosen after I saw the results.

The reasons are partly guesses. That loop 1.0's records never exceeded its forecast is measured, and so is how close it stayed to the frozen model. That this is why it stayed stuck is my explanation. The chain of numbers for loop 1.2 in spring is measured; the claim that one caused the next is my reading. Why loop 1.5 beat the control in five months, I do not know.

The real-world examples are general knowledge. Recommendations, fraud rules, credit and stock ordering share the shape of this loop; none of them was measured here.

What to Do Next

A hand-drawn list headed before you trust a model that acts, titled five questions. Decides?: does the model's answer limit what can happen next? Offered?: do you log what was offered, not only what was taken? Ran out?: how often did the offer run out, or get refused? Outside?: is any slice of decisions made without the model? Truth?: is the model scored on the real outcome, or on its own records? Beneath: here: 2.6 on its own records, 90.7 on the truth.

Take a model that you or your team runs, and ask the first question: does its answer limit what can happen next? If it only predicts, and nothing is decided from the prediction, the loop does not exist. If it sets stock, ranks a list, approves or refuses, or places anything, the loop may exist, and the rest of the questions matter.

Then look at what is logged. If you find only what was taken (rentals, clicks, completed payments), add what was offered (bikes placed, items shown, payments blocked). With both, you can count how often the offer ran out, which is the one number here that an operator could see and that pointed straight at the problem.

The next lesson in this chapter wraps it up: a production readiness check, run against everything the chapter broke, from the flattering test split in lesson 1 to the loop in this one.

A closing card headed to keep, titled a model that decides what gets recorded can hide its own error. In large type: 2.6 on its records, 90.7 on the truth. Beneath: loop 1.0, a simulation on real bike demand: the model placed bikes from its own forecast, learned from what was rented, and turned away 769,320 riders in 2012. A 20% margin brought the error from 90.7 to 60.0; 50% to 46.4, for 820,304 idle bike-hours. Then: log what was offered, keep some exploration, and score on the real outcome.

Knowledge Check

Knowledge Check

4 questions - Score 80% to pass

Q1

Loop 1.0 missed true demand by 90.7 bikes an hour. Why did its error on its own records look so small, only 2.6?

Q2

Why did retraining every month not help loop 1.0 catch up?

Q3

Which number could a real operator see, without knowing demand, that pointed at the censoring?

Q4

What did going from a 20% to a 50% margin cost in this simulation?

MAE

The scores, fixed before the run: for each month, the MAE and the signed error (the average of forecast minus demand, so below zero means too low), both against the true demand, never against the records; the share of hours where demand was bigger than supply; and the number of riders turned away. The headline is the mean of the 12 monthly MAEs.

What the design said about the result. One run per policy over one year gives one number per policy, so the design declared that the lesson would describe the results and run no significance test (a calculation of how likely a difference this big would be by luck alone). Everything after the headline numbers (the comparison with lesson 2's frozen model, the idle bikes, the error against the records, the hour-of-day view) was chosen after I saw the results, and each figure says so.

One thing I added later. The first run stored only monthly totals. To count idle bikes and compare forecasts hour by hour, I added a detail mode to loop_lab.py, with its own design note written before it ran and labelled as added after the results. It reruns the same four policies, stores every hour's forecast, bikes placed and rentals recorded in results/loop_detail.json, and stops without writing anything unless every monthly number it gets is exactly the one in loop.json. All four matched exactly.

A two-column table headed each month of 2012, from the lab's stored file, titled MAE against true demand, month by month. Left, month: control; loop 1.0. Right, loop 1.2; loop 1.5. Twelve rows, for example Jan 2012: 68.9; 68.9, and 68.9; 68.9. Feb 2012: 28.7; 75.8, and 71.9; 55.3. Apr 2012: 39.6; 100.8, and 87.6; 40.9. Jun 2012: 37.7; 90.3, and 48.1; 38.0. Sep 2012: 41.9; 111.2, and 49.6; 41.0. Dec 2012: 42.6; 63.3, and 36.9; 39.3. Beneath: mean: control 43.6, loop 1.0 90.7. Mean: loop 1.2 60.0, loop 1.5 46.4.

The signed errors tell the direction. Loop 1.0 was too low in every month, from -56.0 in December to -110.0 in March: it placed too few bikes every month of the year. The table has every month for all four policies.

The records do not even look flat. Loop 1.0 recorded 62.9 rentals an hour in January and 194.4 at its highest month. On a dashboard, that is a bike-share system that tripled its business. The true demand went from 130.6 to 303.6 in the same months, so the growth was real, but each month the records held only 48% to 71% of true demand. Nothing on such a dashboard points at the missing part.

There is one number the operator could see, and I added it for that reason: the share of hours when every bike placed was rented. The operator knows how many bikes were placed and how many were rented, so this needs no knowledge of demand. For loop 1.0 it was 86.1% of hours; for loop 1.2, 49.4%; for loop 1.5, 17.4%. An hour where every bike went is not proof that riders were turned away (demand may have matched supply exactly), but when it happens in most hours, the records are cut off in most hours.

Two measured facts put the order right. The biggest one-month rise in loop 1.2's level was into May, from 0.67 to 0.80, and the biggest fall in its gap to the control was also into May, from 48.1 to 19.1 bikes an hour. So most of the gap closed under the May model, which was trained on April's records (77% of riders recorded), before any month had 90% recorded. The June model, trained on May's records as well, closed the gap further, from 19.1 to 10.4. In May only half of the hours (50.1%) were not cut off, yet the month recorded 91% of its riders, because even in the cut-off hours the bikes placed covered 86.8% of the riders who came.

My reading of that chain, which I have not tested: a 20% margin lets the records rise up to 20% above the forecast, so each retrain can see a little more demand than the last forecast allowed. In spring the model's level was also rising with the season, as the frozen model's did. April's records, three quarters complete, already moved the May model most of the way; May's, nine tenths complete, moved the June model a good part of the rest.

Loop 1.5 did better still: its level was 0.96 by April. It even had a lower MAE than the control in five months (May, July, August, September, December), while its year was higher (46.4 against 43.6). I have no tested explanation for those five months. One possible reason is that its records cut off a few unusually busy hours that the control learned from, and that this helped in some months; with one run per policy I cannot tell that apart from chance.

The riders turned away were not spread evenly. Loop 1.0 turned away the most at 5 in the afternoon (78,194 over the year) and 8 in the morning (71,700): the commuting hours, when the most people want a bike. Loop 1.5 turned away 4,211 and 7,739 at those two hours. A model that guesses low by a fixed share loses the most riders when the numbers are biggest.

What a margin or exploration gives, and what it costs. In the bikes, the margin gave hours where the record was not cut off, and those hours were how the model learned that demand had grown. The same idea elsewhere is often called exploration: show some items the model did not rank first, approve a small random share of borderline applications, stock a little more than the forecast. Each gives up something now (a less relevant item, a riskier loan, some waste) to see outcomes the system would otherwise never see. The cost is real and it is paid every day; the benefit is information, and it arrives later.

This is a real run in VS Code's terminal (python loop_demo.py).

A real screenshot of VS Code's terminal after running python loop_demo.py. Two header lines: month, control MAE, loop 1.0 MAE, loop 1.0 on its records, riders turned away (loop). Then six lines. Jan 2012: 68.89, 68.89, 1.26, 50,118. Feb 2012: 28.72, 75.75, 1.54, 51,355. Mar 2012: 56.85, 111.18, 1.07, 81,809. Apr 2012: 39.57, 100.76, 2.72, 70,391. May 2012: 42.90, 79.71, 4.47, 55,982. Jun 2012: 37.66, 90.33, 2.45, 63,269.

When I ran it, every printed number matched the lab's stored files: the control's MAE (68.89, 28.72, 56.85, 39.57, 42.90 and 37.66) and loop 1.0's MAE (68.89, 75.75, 111.18, 100.76, 79.71 and 90.33) match loop.json, the riders turned away match it to the rider, and the error on its own records (1.26 to 4.47) matches what the report computes from loop_detail.json. The report's demo mode checks this line by line. Look at March: 111.18 bikes an hour against the riders who came, 1.07 against the records.

detail
json
results/lp-report.json

Four brand cards headed the tools, with their logos, titled what ran where. scikit-learn: every model fit, the data download. pandas: the table of hours. NumPy: the rule: bikes placed, rentals recorded. Python: the lab, the report, the demo, the box.

loop[now] = seen

The scores. true_c and true_f compare each forecast with the true demand. own compares loop 1.0's forecast with its own records, the only error an operator could compute. away adds up the riders who found no bike.

In the lab file, loop_lab.py does this for all twelve months and for margins 1.2 and 1.5 as well, and its detail mode stores every hour.