Ml Lifecycle

Canary and Shadow: How Much Live Traffic It Takes to See a Worse Model

0 of 26 complete

0%

Contents

Back|Ml LifecycleCanary and Shadow: How Much Live Traffic It Takes to See a Worse Model
1/26
72 min left
Prerequisites
The Promotion Gate: How Often a New Model That Is Not Better Gets the Jobrequired
Related Topics
Tomorrow Is Different: A Model Frozen on 2011, Scored Through 2012Why Production BreaksTrain/Serve Skew: One Input Computed Two WaysWhy Production BreaksFailing Silently: Which Checks Catch a Broken Input Before the Answers ArriveWhy Production BreaksWatching Inputs Before the Answers Arrive: Drift Measures Against the Real ErrorWhy Production BreaksWhen the Data Answers Back: A Feedback Loop, Simulated on Real DemandWhy Production Breaks
1 of 26

The New Loaf

Imagine a bakery with twenty shops that has sold the same bread for years. The head baker has a new recipe that is cheaper to make. She thinks it tastes just as good, but she is not sure, and a bad loaf in every shop at once would cost her customers. So she has two careful ways to try it.

The first way: sell the new loaf in one shop out of twenty and keep the old loaf everywhere else. Then count the complaints. If the one shop gets more complaints than the other nineteen, the new loaf is probably worse. The trouble is that the shops are different. One shop is near a school and gets fussy children; another is near an office and gets people in a hurry. A few extra complaints in one shop may say more about its customers than about the bread.

An illustration of a woman in a shirt and trousers, standing and waving one hand, next to text. Headed the new loaf, titled how long before you notice the new model is worse? Beside her: a canary sends a share of live rows to the new model; shadow runs the new model silently on every row. A known 4.09-point regression, 20 starts: a 20% canary raised the alarm in 12; shadow in 20, after 241.5 half-hours (median). The same model checked after every half-hour: a canary raised a false alarm in 9.8% of runs, against 1.1% when checked once. Last: compare the same rows, and decide in advance when to look.

The second way: in every shop, bake both loaves for every order. The customer gets the old loaf, as always. A taster in the back tries both loaves from the same order and writes down which one is better. No customer ever eats the new loaf, so no customer is harmed. And because the taster compares two loaves made for the same order, the fussy school children and the busy office workers no longer spoil the comparison. The cost is simple to see: the bakery bakes twice as much bread, and the taster's opinion is not the same as a paying customer's.

Both ways have a delay. A complaint arrives the next day, not at the counter. So the baker has to wait for the answers, and she has to decide how often to look at them. This lesson measures both ways for a program that learns from examples, and asks how long each one takes to notice that the new version is worse.

Where This Lesson Starts

This chapter follows one model through its life on the same data. Start with the dumbest model set the baselines. Same code, different model showed that two runs of the same code give two different models. Picking the best of many showed that the winner of a big search can do worse on later months. The promotion gate measured how a new model is judged against the old one on stored rows, and found that a test which looks only at the rows where the two disagree held its false promotions near 5%, where "higher score wins" promoted equal models about half the time.

Now the new model has passed the gate, and it meets real traffic. The course's lesson on model serving and inference has a slide called "Shipping a New Version Without Betting All Your Traffic". It explains in words what a canary and a shadow rollout are, and in the middle of the slide it gives a short rule: shadow proves it is safe, canary proves it is better. I will not repeat it. One thing to add to that rule: when the right answer arrives for every row, shadow can also measure accuracy, and that is what this lesson measures. This lesson measures one question that slide does not: how many rows of live traffic each way needs before it can see that a new model is worse.

A flowchart headed step 7 of the life of a model: meet live traffic, titled where canary and shadow sit. A box, the gate passed: a new model to roll out, leads to a diamond, canary or shadow: watch it on live rows. From the diamond, alarm leads to roll back: the old model keeps every row, and no alarm by the end leads to full rollout: the new model serves everyone. Beneath: lesson 4 measured the gate on stored rows. This lesson measures what live traffic can show, and how fast.

Why test on live traffic at all, if the gate already passed? Because the gate saw stored rows from the past, and live rows come from now. Lesson 3 found that months can differ a lot on this data. A model that passed on last season's rows can still be worse on this season's, and live traffic is the first place that shows.

Nine Words for This Lesson

A hand-drawn list headed nine words for this lesson, titled putting a new model in front of real users. Rollout: the steps by which a new model replaces the old one for real users. Regression: a new version that is worse than the old one; not the word in 'logistic regression'. Live traffic: the real requests arriving now, one row per half-hour here. Canary: a share of live rows goes to the new model; the rest stay with the old one. Shadow: the new model answers every row silently; users only see the old one. Label delay: the time until the right answer is known: half an hour here. Alarm: the test says the new model is worse, so roll it back. False alarm: an alarm when the new model is not worse. Paired: both models judged on the same rows, row by row. Beneath: a canary compares different rows; shadow compares the same rows.

A rollout is the set of steps by which a new model replaces the old one for real users. To roll back is to send all traffic back to the old model. A regression is a new version that is worse than the old one. That is not the same word as the "regression" in "logistic regression", which is the name of a kind of model. This lesson uses both, and one of its challengers is both at once: the logistic regression model is the "big regression", because it is much worse than the champion. Live traffic is the stream of real requests arriving now. In this lab, one request is one half-hour of the electricity market, so the traffic is one row every half-hour. As in lesson 4, the old model that is live now is the champion and the new one is the challenger.

A canary rollout sends a share of live rows to the challenger, chosen at random, and keeps the rest on the champion. The name comes from the birds that miners once carried underground: if the air went bad, the small bird showed it first. A shadow rollout sends every row to both models. Users only ever see the champion's answer; the challenger's answer is written down and compared, never shown.

The label is the right answer for a row, and the label delay is how long it takes to arrive. Here, whether the price went UP or DOWN is known half an hour later. An alarm is the monitoring test saying "the challenger is worse, roll it back". A false alarm is an alarm when the challenger is not worse. A comparison is paired when both models are judged on the same rows, row by row. is the share of rows a model got right, and a is one hundredth of accuracy.

Canary and Shadow, Row by Row

Here is the difference between the two rollouts on real rows. The sketch shows the first 16 live half-hours of the lab, with the lab's real random draw for a 20% canary and the real right-or-wrong outcomes of two models: the champion, which is boosted trees, and the logistic model from lesson 1, which is worse.

A hand-drawn sketch headed sketched: the first 16 live half-hours, the lab's real 20% draws and real outcomes, titled canary: one model per row. Shadow: both on every row. Four rows of 16 boxes. Canary, who: O, O, N, N, then O seven times, N, O, N, O, N. Canary, result: R in every box except the seventh and the fourteenth, which are w. Shadow, old: R in every box except the seventh. Shadow, new: R in every box except the seventh and the fourteenth. Beneath the boxes: O = old model answered, N = new model, R = right, w = wrong. Beneath the sketch: old: the boosted trees. New: the logistic model. In these 16 rows the canary sent 5 to the new model; shadow saw both answers on all 16, and they differed on 1.

In the canary, each half-hour is answered by one model only. The draw sent 5 of these 16 rows to the new model. On each row, we learn how one model did, and never how the other would have done. To judge the new model, the canary compares its mistakes on its rows with the old model's mistakes on the other rows. Those are different half-hours, some easy and some hard.

In shadow, every half-hour is answered by both. On 15 of these 16 rows the two gave the same result, both right or both wrong. On one row, the fourteenth, only the old model was right. That single row is all the evidence these 16 half-hours hold about which model is better, and shadow sees it directly.

Look at the seventh half-hour. Both models got it wrong. In the canary, that row happened to go to the old model, so it counts as one mistake for the old model and tells us nothing about the new one. In shadow, it counts for neither, because the two agree. That is the whole idea of pairing, and the next slides measure what it is worth.

Two Alarm Tests

Each rollout needs a rule that turns the rows so far into a yes-or-no alarm. The lab fixed both rules in its file before it ran.

A table of three rows headed the lab's two alarm tests, fixed before the run, titled two ways to say the new model is worse. Canary: compare the new model's error rate on its rows with the old model's on its rows; alarm when the gap is too big for luck (z above 2.33, about 1 in 100); wait for 30 rows each. Shadow: count the rows where only one model was right; alarm when the split is too uneven for coin flips (McNemar, p below 0.01); wait for 10 such rows. Both: run again after every half-hour, on all the rows so far. Beneath: one-sided: both only ask whether the new model is worse. The 1-in-100 bar is stricter than lesson 4's 0.05, because a rollout is checked many times.

The canary test compares two error rates, the share of rows each model got wrong, measured on two different groups of rows. To judge the gap, it first puts both groups together and computes one pooled error rate, the share of all rows so far that were answered wrong. That pooled rate says how much two error rates on groups of this size would normally wobble by luck if the models were equal. The test asks how big the real gap is compared with that wobble. That ratio is called z. A z above 2.33 happens by luck about 1 time in 100 when the two models are equally good, so the lab sounds the alarm there. The test waits until each group has at least 30 rows, because with fewer rows the wobble is too wild to measure.

The shadow test is McNemar's test from lesson 4, turned around. It counts only the rows where exactly one model was right. If the two models were equally good, each such row would be a fair coin toss between them. The test asks how often fair coins would split this unevenly against the challenger. That chance is the p-value, and the lab sounds the alarm when it falls below 0.01, again about 1 in 100. It waits for at least 10 disagreeing rows. The bar of 0.01 is stricter than the 0.05 of lesson 4's gate, because a rollout is checked many times, and every check is another chance to be fooled.

Both tests are one-sided: they only ask whether the challenger is worse, because a rollout only rolls back in one direction. Both use the same 1-in-100 bar, so their false alarms can be compared fairly. And both were run again after every half-hour, on all the rows so far; a later slide measures what that does.

Why Comparing the Same Rows Needs Fewer of Them

Before the results, here is the reason to expect shadow to be faster, and it has nothing to do with this lab.

A model's mistakes come from two things: the model, and the row. Some half-hours are simply hard. A price that jumps in a way the inputs could not foresee will fool almost any model. When a canary compares two error rates on different rows, both numbers carry the difficulty of their own rows. If the challenger's share of rows happens to hold a few more hard half-hours, it looks worse even if it is not. The canary has to wait until that luck of the rows averages out.

Three panels headed what each comparison has to see through, all live rows; after the results, titled shadow reads only the rows where the two differ. Canary: 25%, of rows the old model gets wrong anyway: noise both groups carry. Shadow, logistic: 1,697 to 758, rows only the old right, to only the new right. Shadow, injected 5%: 371 to 0, every difference points the same way. Beneath: out of 9,063 live rows. Rows where both were right or both wrong say nothing, and shadow ignores them.

Shadow sees both models on the same half-hour. A hard row that fools both of them is a row where both were wrong, and it drops out of the test. So does an easy row both got right. What is left are the rows where the models truly behaved differently. On the 9,063 live rows, the champion got about 25% wrong. For the logistic model, the rows where only one was right numbered 1,697 against 758; for the injected challenger of the follow-up, 371 against 0. I counted these after the results. The canary has to find a difference hidden in the full 25% of noise; shadow looks at the difference itself.

This is the same point as the cafe customers in lesson 4 who liked both cups equally. They carry no information about which blend is better, and a paired test simply stops listening to them. The price of pairing is that both models must answer the same rows, which is exactly what shadow does and a canary cannot.

What the Lab Ran

I wrote the lab's design at the top of its file, canary_lab.py, before it ran. The dated record is in the chapter plan (2026-09-29, "CANARY LAB (batch 5) designed before running").

An editorial page in four labelled zones, headed what the lab ran: canary_lab.py, designed before it ran, titled one champion, three challengers, four ways to watch. The data: Elec2: train on the first 36,249 half-hours; the last 9,063 play live traffic, 1998-06-01 to 1998-12-06. Each label is known half an hour later. The champion: boosted trees, seed 0: 0.7527 on all live rows. Three challengers: trees without two input columns; the logistic model of lesson 1; the champion itself. Four ways to watch: canary at 5%, 20% and 50% of rows, and shadow. 20 starts across the live rows, 2,000 half-hours from each. Beneath: the injected regressions and the swapped equal pair came after the results.

The data is Elec2 again: 45,312 half-hours of the New South Wales electricity market, May 1996 to December 1998, in time order. Each row asks whether the price is UP or DOWN against its average over the last 24 hours. The models train on the first 36,249 rows. The last 9,063, from 1 June to 6 December 1998, play the part of live traffic, one row per half-hour, and each row's label is known half an hour later, when the next row arrives.

The champion is the boosted trees with seed 0, the model of lessons 1 to 4. It scored 0.7527 on all the live rows. Three challengers were fixed in advance. A "small regression": the same trees trained without two inputs, vicprice and vicdemand, the price and demand in the neighbouring state of Victoria, as if a data pipeline had dropped two columns. A "big regression": the logistic model of lesson 1, about 10 points worse. And an "equal" challenger: the champion itself, so that any alarm would be a false one.

A sequence diagram with four columns: live row, router, one model, the monitor. Headed one half-hour of the lab, canary, titled a row, an answer, and the label half an hour later. Step 1, the live row sends the router a new half-hour. Step 2, the router sends one model the row: old or new, by a draw. Step 3, the model sends back UP or DOWN. Step 4, the model sends the monitor its answer, logged. Step 5, the live row sends the monitor the real label, later. Step 6, the monitor marks right or wrong and tests: alarm? Beneath: in shadow, step 2 sends the row to both models and only the old answer goes back. The draws use seed 0, the same at every start.

The Main Run

Here is what the lab stored in results/canary.json. For each challenger and each way of watching: in how many of the 20 starts an alarm came, and the median half-hour of the first alarm among those that came. The median is the middle value when they are lined up.

A two-column table headed the main run, canary.json: alarms of 20 starts, median half-hour, titled the planned small regression was not a regression. Left, challenger and rule; right, alarms and median. Small regression (in fact 0.61 points better), canary 5 / 20 / 50%: 2 / 0 / 2. Small regression, shadow: 0 of 20. Big regression (10.36 points worse), canary 5 / 20 / 50%: 14 (1,001.5) / 15 (202) / 14 (87). Big regression, shadow: 18 of 20, 115.5. Equal (the champion itself), canary 5 / 20 / 50%: 1 / 0 / 4. Equal, shadow: 0 of 20. Beneath, left: the small regression scored higher: 0.7588 against 0.7527. Beneath, right: every alarm on the small regression and on equal was a false one.

The big regression, 10.36 points worse over the whole live period, was caught by a 5% canary in 14 of 20 starts, with a median of 1,001.5 half-hours, about 21 days. A 20% canary caught it in 15, median 202 half-hours, about 4 days. A 50% canary caught it in 14, median 87 half-hours, under 2 days. Shadow caught it in 18 of 20, median 115.5 half-hours, about 2.4 days.

At first sight, a 50% canary looks faster than shadow: 87 against 115.5. A later slide shows that this comparison is misleading, because the two medians count different starts.

The small regression was not a regression. It scored 0.7588 on the live rows, 0.61 points above the champion. Dropping the two Victorian inputs made the trees slightly better here. Shadow never raised an alarm on it, which is correct. The canary did, twice at 5% and twice at 50%. In those starts the challenger was between 0.30 and 0.75 points better than the champion over the 2,000 half-hours, so these were false alarms against a model that was not worse. I measured those per-start gaps after the results.

The equal challenger, the champion itself, got 1, 0 and 4 false alarms from the canary. Shadow gave 0.

Two Faults in My Design

The main run answered some of what I asked, but two of my three challengers were built wrong, and I want them on the record rather than hidden.

Two panels headed two faults in the design, found in the results, titled what the main run could not tell me. Fault 1: +0.61, points: the 'small regression' was better, not worse. Fault 2: 0 of 20, shadow alarms on the champion against itself: it can never disagree. Beneath: the follow-up fixed both: a regression of a known size, and two models that are equal but not identical.

Fault 1: the small regression was not a regression. I guessed that removing two inputs would make the model a little worse, and wrote "its size is whatever the data gives" into the design. The data gave a model 0.61 points better. So the main run has only one real regression, the logistic model, and no small one. Its question, "how much traffic does it take to see a small regression?", was left without an answer.

Fault 2: the equal challenger made shadow's zero meaningless. I used the champion itself as the equal challenger, so its answer matched the champion's on every row. For a canary that is a fair test: the two groups are still different rows, so luck can still make one look worse, and the 1, 0 and 4 false alarms are real. But shadow counts only the rows where the two models disagree, and here there were none. Shadow could not have raised an alarm whatever its rule. Its 0 of 20 says nothing about how often shadow raises false alarms.

So I designed a follow-up after I saw the main results, and I wrote its design into the lab file's followup mode before it ran. It fixed both faults. For regressions of a known size, it took the champion's own live results and gave each right answer a 2% or 5% chance of being turned wrong. For an equal pair that can disagree, it used lesson 4's trick: two seeds of the trees, with their results swapped row by row by a coin, so that neither can be better by design. Its results are in results/canary_followup.json.

The Big Regression, Start by Start

Before the follow-up, the main run's one real regression deserves a closer look, because its 20 starts are not alike. After the results, I measured how far the logistic model was behind the champion inside each start's own 2,000 half-hours.

A dot chart headed the big regression, each of 20 starts: how far behind it was, and when the alarm came, titled the further behind, the sooner the alarm, mostly. Dots for canary 20% and shadow on a scale of logistic behind the champion in that start's 2,000 rows, points, from 0 to 20, against half-hour of the first alarm, from 0 to 2,000. Near 0.5 to 0.7 points there are only three shadow dots, at about 160, 290 and 1,000. Near 4 points there are two dots, near 1,575 and 1,730. Near 11 points, dots between about 750 and 1,700. From about 13 to 18 points, most dots sit below 600, and the shadow dots are the lowest, several under 100. Beneath: the first five starts, which begin in June and early July, were only 0.5 to 1.4 points behind: no canary alarm. Shadow missed two, at 0.6 and 1.4 points. Measured after the results.

The "10-point regression" is an average over the whole live period, and it hides a lot. In the first five starts, which begin in June and early July, the logistic model was only 0.5 to 1.4 points behind over its 2,000 half-hours. In the last fourteen starts it was 10.7 to 17.7 points behind. No canary of any size raised an alarm in the first five starts. Shadow raised one in three of them and missed the other two. In the late starts, where the gap was large, everything saw it, and shadow saw it first.

Three panels headed the big regression, the same start, both alarmed; after the results, titled where both raised the alarm, shadow was first every time. Against canary 5%: 14 of 14, shadow first; shadow alone on 4 more. Against canary 20%: 15 of 15, shadow first; shadow alone on 3 more. Against canary 50%: 14 of 14, shadow first; shadow alone on 4 more. Beneath: medians alone mislead: canary 50% had a median of 87 and shadow 115.5, because they count different starts.

This explains the strange medians. The 50% canary's median of 87 half-hours comes only from the 14 starts where it caught anything, and those were the late starts with a big gap. Shadow's median of 115.5 includes slow catches from the early starts, where the canary caught nothing at all. Compared start by start, which I did after the results, shadow was first in every start where both raised an alarm: 14 of 14 against the 5% canary, 15 of 15 against 20%, and 14 of 14 against 50%. It was never second.

The lesson for reading any rollout report: a median of "time to alarm" means little unless you know which runs had an alarm. A method that only catches the easy cases can post a faster median than one that catches everything.

A Regression of Known Size

The follow-up gives two regressions whose size is known exactly. Each of the champion's 6,822 right answers on the live rows had a 2% or a 5% chance of being turned wrong. The draw turned 159 of them (2.3%) and 371 (5.4%), which made challengers 1.75 and 4.09 points worse. I counted those after the results. The same 20 starts, the same four ways of watching, the same alarm tests.

A bar chart headed follow-up, designed after the results: known regressions, 20 starts, titled shadow caught both in every start; a canary missed most of the small one. Two groups of four bars, canary 5%, canary 20%, canary 50% and shadow, at 1.75 points and 4.09 points, on a scale of starts with an alarm, of 20, from 0 to 20. At 1.75 points: 2, 4, 7 and 20. At 4.09 points: 6, 12, 15 and 20. Beneath: canary 5 / 20 / 50%: 2 / 4 / 7 and 6 / 12 / 15. Shadow median: 524.5 and 241.5 half-hours.

The 1.75-point regression: the canary raised an alarm in 2, 4 and 7 of 20 starts at 5%, 20% and 50%. Shadow raised one in all 20, with a median of 524.5 half-hours, about 11 days. The 4.09-point regression: the canary in 6, 12 and 15 of 20; shadow in all 20, median 241.5 half-hours, about 5 days.

The equal pair, two seeds with their results swapped by a coin, came out 0.08 points apart over the live period, which is as close to equal as real models get. The canary raised false alarms in 1, 0 and 2 of 20 starts. Shadow raised none. This time shadow's zero means something, because these two models did disagree: on 325 of the 9,063 live rows, and the swap made each such row a fair coin.

Start by start, again after the results: against the 4.09-point regression, shadow alarmed first in 6 of the 6 starts where the 5% canary also alarmed, 11 of 12 at 20%, and 11 of 15 at 50%. So a big canary sometimes beat shadow to the alarm here, which never happened with the logistic model. But in every case the canary missed starts that shadow caught: 14, 8 and 5 of them.

Why Shadow Caught All of Them, and Why That Is the Easiest Case

Twenty out of twenty, twice, is an unusually clean result, and an unusually clean result means inspect the inputs first. So I did, after the results.

Two panels headed why shadow caught every injected regression; after the results, titled an injected error only ever points one way. Injected: 10 to 0, at the alarm: p 0.001, the 10th disagreeing row, every start. The real logistic: 1,697 to 758, it is also right where the champion is wrong. Beneath: the injected challenger is the easiest case a paired test can meet. A real regression is not.

The injected challenger was built from the champion by turning some right answers into wrong ones. It never turns a wrong answer into a right one. So on every row where the two disagree, it is the champion that is right. The split of disagreeing rows is always "all of them to the champion, none to the challenger". For McNemar's test, that is the most lopsided split possible: the chance of 10 fair coins all landing the same way is 0.5 to the power 10, about 0.001. The alarm fired at exactly the 10th disagreeing row in every start, for both injected regressions; the test was only waiting for the 10-row minimum. The report checks this for all 40 alarms.

A real worse model does not behave like that. The logistic model was right on 758 live rows where the champion was wrong, against 1,697 the other way. Its disagreements point both ways, so shadow has to wait for the uneven split to build up. That is why shadow took a median of 115.5 half-hours on the logistic model and missed two starts, while a much smaller injected regression was caught every time.

So the follow-up's shadow numbers are the best case for a paired test. They show how fast shadow can be when every disagreement is a real loss. They do not show how fast it is on a real regression of 4 points; that would depend on how often the real challenger also wins some rows. The canary's numbers are affected much less, because the canary never looks at disagreements: it compares error rates, and a 4.09-point rise in errors is a 4.09-point rise however it was made. That last point is reasoning, not a measurement: errors that bunch up in time could still change what a canary sees.

Does a Real Regression Bunch Up in Time?

The injected errors were scattered at random, one row here, one row there. A real regression might not be. It might be fine for weeks and then fail badly when the weather or the market changes. That matters because a test that counts rows as separate coin flips assumes they are scattered. So I measured it on the stored outcomes, after the results.

A bar chart headed how far behind the champion, month by month, all live rows; after the results, titled the real gap opened in August; the injected one was there every month. Two bars per month, logistic and injected 5%, for June to December 1998, on a scale of champion minus challenger, points, from -5 to 30, with a dashed line at 0. The logistic bar is under 1 in June and just below 0 in July, then about 15 in August, 11 in September, 17 in October, 16 in November and 27 in December. The injected bar sits between about 3 and 5 in every month. Beneath: logistic: 0.9, -0.5, 15.4, 10.8, 16.7, 15.5, 27.1. Injected 5%: 2.7 to 4.7. In Jul 1998 the logistic model was ahead. Dec holds 288 rows.

It does bunch up. Month by month, the logistic model was 0.9 points behind the champion in June and 0.5 points ahead in July. Then the gap opened: 15.4 points in August, 10.8 in September, 16.7 in October, 15.5 in November and 27.1 in the six days of December. The injected 5% regression was 2.7 to 4.7 points behind in every month, as you would expect from errors thrown in at random. The injected 2% was 1.0 to 2.1.

Two panels headed the rows only the champion got right, in time order; after the results, titled the real regression's errors come in runs. Logistic: 2.84, rows in a run on average, longest 27; next-row correlation 0.57. Injected 5%: 1.05, rows in a run on average, longest 3; next-row correlation 0.01. Beneath: neighbouring rows of the real regression go wrong together. Tests that count rows as separate coin flips assume they do not.

It bunches up row by row too. Take the rows where only the champion was right, in time order. For the logistic model, those 1,697 rows came in runs of 2.84 rows on average, the longest 27 in a row, and whether one half-hour was such a row predicted the next one strongly: a correlation of +0.57, where 0 means no link and 1 means always the same. For the injected 5% errors, the runs averaged 1.05 rows and the correlation was +0.01, which is no link at all.

Why did the logistic model fall behind from August? I measured that it did, not why. One possible reason, and it is a guess I have not tested: something in the inputs changed from August in a way the straight-line model handles worse than the trees. Lesson 3's comparison of months cannot settle it, because its later months include June and July, when the logistic model was fine. What the measurement does settle is the caveat the follow-up needed. Random injected errors suit a paired test that treats rows as separate coin flips. A real regression came in runs and in seasons, and a test that assumes separate coin flips may be too sure of itself on such rows. This lab did not measure how much.

How Many Rows a Canary Sees

There is a simpler reason a small canary is slow, and it is plain counting. I counted it after the results, from the lab's own random draws.

An isometric drawing of four blocks, heights to scale, headed rows the new model answers in 2,000 half-hours; heights to scale, titled a small canary sees very few rows. A flat block labelled 110, canary 5%; a small cube labelled 419, canary 20%; a taller block labelled 1,007, canary 50%; and the tallest block labelled 2,000, shadow. Beneath: the canary may test only once both sides hold 30 rows: at half-hour 702, 158 and 62. Shadow gets the new model's answer on every row. Counted after the results.

In 2,000 half-hours, almost 42 days, a 5% canary sent just 110 rows to the challenger. A 20% canary sent 419, a 50% canary 1,007. Shadow gets the challenger's answer on all 2,000. A canary's test needs rows from both models, and the fewest it gets are from the challenger. With 110 rows, the error rate of the challenger is only known to within several points either way, so only a large regression can be told apart from luck.

The 30-row minimum has a second effect. A 5% canary reached 30 challenger rows only at half-hour 702, almost 15 days in; a 20% canary at 158; a 50% canary at 62. That first allowed test is also the one with the fewest rows, when luck swings the most. Of the big regression's alarms, 5, 6 and 5 fired at that very first allowed check. The same first check shows up over and over in the table of starts, because the lab used the same random draw at every start. That was a simplification of my design, and a reason the 20 starts are less independent than they look.

This is where the real world differs most from this lab. Elec2 has one row per half-hour. A real service may get thousands of requests a minute, so a 5% canary could see 110 labelled rows within minutes. What matters is not the clock but the number of labelled rows each model has answered. The numbers in this lesson are in half-hours because that is how the data arrives; read them as rows.

Checking After Every Half-Hour

Both alarm tests were run again after every new half-hour, on all the rows so far. Lesson 4 measured what repeated checking does to a promotion gate: on equal models, checking after every 200 rows turned 2 false promotions out of 20 into 6. Monitoring a rollout is exactly the case where a team wants to look at every new batch, so I measured it here too, after the results.

Two panels headed equal models, 4,000 runs each, checked after every half-hour; after the results, titled checking after every half-hour raised false alarms. Canary 20%: 9.8%, false alarms; once at the end: 1.1%. Shadow: 4.4%, false alarms; once at the end: 0.6%. Beneath: canary 5% and 50%: 6.0% and 9.4% against 0.9% and 0.9%. The test aims at 1%. 200 new seeds x 20 starts that overlap.

For the canary I ran the champion against itself, which is a true null for a canary because the two groups are still different rows, with 200 new random draws at each of the 20 starts: 4,000 runs per share. For shadow I used the swapped pair of seeds with 200 new coins, also 4,000 runs. Each run was tested two ways: after every half-hour, and once, at the end of the 2,000.

Tested once at the end, the false alarms landed where the 1-in-100 bar says they should: 0.9%, 1.1% and 0.9% for the three canaries and 0.6% for shadow. Tested after every half-hour, they rose to 6.0%, 9.8% and 9.4% for the canaries and 4.4% for shadow. So watching continuously multiplied the false alarms by about seven to eleven: 6.9, 8.9 and 10.7 times for the three canaries, and 7.0 times for shadow. Shadow's rates were lower only in absolute terms; its ratio was the same as the 5% canary's. One likely reason its rates are lower, which is reasoning rather than a measurement: with few disagreeing rows an exact test cannot land exactly on its bar and stays below it, 0.6% here when checked once, against the 1% it aims at.

The swapped pair also makes shadow's side a best case. The coin makes every disagreeing row a fair and separate toss, which is exactly what McNemar's test assumes. Real equal models whose disagreements come in runs, like the logistic model's errors, could give shadow more false alarms than 4.4%, and this lab did not measure how many.

This also explains the main run's false alarms. The canary's 1, 0 and 4 false alarms on the equal challenger, and the 2 and 2 against a slightly better model, all came from checking after every half-hour. The 4,000 runs are not independent: the 20 starts overlap, and neighbouring starts share about 81% of their rows. For the canary I also counted where those false alarms fired: 50 of 240, 49 of 391 and 35 of 375 came at the very first allowed check (21%, 13% and 9%), and the rest a median of 322, 267 and 220 half-hours after it. Read the rates as a clear direction, not as exact values.

What Each Way of Watching Costs

Shadow looks better on every measurement so far. It is not free, and its costs are of a different kind from the canary's.

A two-column table headed the big regression, medians of 20 starts, up to the alarm or the end; after the results, titled what each way of watching costs. Left, canary: users get it; right, shadow: nobody does. Canary 5%: 72.5 rows answered by the new model, 9 of them wrong where the old one was right. Shadow: 0 rows answered for real. Canary 20%: 104 rows answered by the new model, 14.5 of them wrong where the old one was right. 178.5 extra predictions, up to the alarm or the end. Canary 50%: 168.5 rows answered by the new model, 29 of them wrong where the old one was right. And no real user reaction is ever seen. Beneath, left: a canary costs wrong answers to real people. Capped at 2,000 half-hours: a missed regression goes on costing after that. Beneath, right: shadow costs two models on every row.

A canary's cost is paid by users. Every row the challenger answers is a real answer to a real person. With the big regression, I counted after the results how many rows the challenger answered before the alarm, or before the end of the 2,000 half-hours when no alarm came, and how many of those it got wrong where the champion would have been right. The medians over the 20 starts: 72.5 rows and 9 extra wrong answers for a 5% canary, 104 and 14.5 for 20%, 168.5 and 29 for 50%. A bigger canary sees the problem sooner, and hurts more people while it does.

These counts stop at 2,000 half-hours. A regression the canary missed goes on giving wrong answers after that, for as long as the rollout runs, so the counts make a small canary look cheaper than it is.

Shadow's first cost is compute. Every row is answered twice. Counted the same way as the canary, up to the alarm or to the end of the 2,000, shadow ran a median of 178.5 extra predictions on the big regression. For small boosted trees that is nothing. For a large model that needs an expensive graphics card for each answer, running two copies on all traffic can double what the service costs to run, and many teams shadow only a sample of traffic for that reason.

Shadow's second cost is blindness. The challenger's answers are never shown, so nothing a user does in reply to them is ever seen. In this lab that does not matter: the price goes up or down whatever the model says. In a real product it often does matter. A recommendation changes what people click; a fraud score changes which payments go through; a search ranking changes what people read. Those effects only exist when real people get the new answers. Shadow can show that a model is less accurate on the same rows. Only a canary can show what users do with it.

What I Would Do Instead

Here is what the lab points to, in the order a rollout meets it.

Start in shadow, and pair the rows. Run the challenger on every live row next to the champion, and compare them only on the rows where exactly one was right. On every regression in this lab, shadow raised the alarm in more starts than any canary, and on the real one it was first in every start where both did.

Count labelled rows, not hours. Before a rollout, work out how many labelled rows each model will answer per day. A 5% canary here saw 110 rows in almost 42 days, which could only see a very large regression. If a canary is too small to see the regression you care about, it is only a slow way of harming a few users.

Decide in advance when to look. Choose the points at which you will test, and test only there, or use a method built for repeated looks, which demands a stricter bar at each look. Here, testing after every half-hour turned a 1% false-alarm rate into 6 to 10% for a canary.

Be careful with the first look. The first test a canary may run is the one with the fewest rows. Here more than a third of the big regression's canary alarms fired at that very first check: 5 of 14, 6 of 15 and 5 of 14. A real regression that big deserves the alarm. For false alarms, the first check is a single look on the fewest rows: in the 4,000 runs on equal models it held 21%, 13% and 9% of the canary's false alarms, and none of the lab's own 12.

Look at time slices, not only the total. The logistic model was a 10-point regression on average, and less than 1.5 points behind in the starts that began in June and early July. A canary that ran in June would have passed it; shadow caught it in 3 of the 4 June starts. If your data has seasons, watch long enough to cover the ones that matter.

Then use a canary for what only a canary can see. Once shadow shows no accuracy problem, a small canary is the way to see how real users respond. Keep it small, keep the old model ready to take all traffic back, and log every alarm with the rows and counts behind it, next to the gate record of lesson 4.

Try It Yourself

This script is the lab made small. It downloads the same data, trains the champion, builds the known 4.09-point regression exactly as the follow-up did (its comment says 5% of the right answers; each one had a 5% chance, and 5.4% were turned), and then watches the first 2,000 live half-hours twice: once as a 20% canary, and once in shadow. It prints the half-hour of each first alarm, or says that none came. It does not need a GPU.

A real screenshot of VS Code with canary_demo.py open, showing the docstring that says what the script is and how to run it, the imports, the lines that load Elec2 and turn the label into 1 for UP and 0 for DOWN, the cut at 80%, the champion trained with seed 0 and marked right or wrong on each live row, and the challenger made by turning a random 5% of its right answers wrong; the canary and shadow functions are further down. Beneath: copy it from the box on the slide.

Before you run this lab. You need Python 3 and two libraries: pip install scikit-learn pandas. scikit-learn holds the model and the download, and brings NumPy with it; pandas holds the table. The first run downloads Elec2 from OpenML (under 1 MB compressed), so it needs an internet connection once; scikit-learn keeps a copy in a folder in your home directory (scikit_learn_data) for later runs. Both alarm tests need no extra library: they are a few lines of arithmetic with Python's own math. I ran it with scikit-learn 1.9.1 on a Mac. scikit-learn runs the same way on Windows and Linux, but I have not checked the numbers there, and another version of scikit-learn may give different decimals, which is why the first line printed is the version.

"""Canary and shadow: how soon does each one see a model that got worse?

Lesson 5 of 'The ML & AI Lifecycle', made small. It needs Python 3 with scikit-learn
and pandas (pip install scikit-learn pandas). The first run downloads Elec2 from
OpenML (under 1 MB) and keeps a copy for later runs.
    python canary_demo.py

Author: Roni Das
Created: 2026-09-29
"""
from math import comb, sqrt

import numpy as np
import sklearn
from sklearn.datasets import fetch_openml
from sklearn.ensemble import HistGradientBoostingClassifier

# 45,312 half-hours in time order. The label: is the NSW price UP or DOWN
# against the average of the last 24 hours? It is known half an hour later.
data = fetch_openml("electricity", version=1, as_frame=True, parser="auto").frame
X = data.drop(columns="class").astype(float)
y = (data["class"].astype(str) == "UP").astype(int).to_numpy()
cut = int(0.8 * len(y))  # train on the first 80%; the rest is "live"
print(f"scikit-learn {sklearn.__version__}, {len(y) - cut:,} live half-hours")

model = HistGradientBoostingClassifier(random_state=0).fit(X[:cut], y[:cut])
champion = model.predict(X[cut:]) == y[cut:]  # right or wrong, each live row

# The challenger: the champion with 5% of its right answers turned wrong,
# so we know for certain that it is worse.
spoil = np.random.default_rng(4).random(len(champion)) < 0.05
challenger = champion & ~spoil
gap = 100 * (champion.mean() - challenger.mean())
print(f"champion {champion.mean():.4f}, challenger {challenger.mean():.4f}")
print(f"the challenger is worse by {gap:.2f} points")

WATCH = 2000  # half-hours watched, from the first live half-hour
to_new = np.random.default_rng(0).random(WATCH) < 0.20  # the canary's 20%


def canary():
    # each row is answered by ONE model; compare the two error rates
    wrong_new = n_new = wrong_old = n_old = 0
    for t in range(WATCH):
        if to_new[t]:
            n_new += 1
            wrong_new += int(not challenger[t])
        else:
            n_old += 1
            wrong_old += int(not champion[t])
        if n_new >= 30 and n_old >= 30:
            pool = (wrong_new + wrong_old) / (n_new + n_old)
            se = sqrt(pool * (1 - pool) * (1 / n_new + 1 / n_old))
            if (wrong_new / n_new - wrong_old / n_old) / se > 2.33:
                return t + 1, n_new
    return None, n_new


def shadow():
    # both models answer EVERY row; count the rows where only one was right
    b = c = 0  # b: only the champion right, c: only the challenger right
    for t in range(WATCH):
        b += int(champion[t] and not challenger[t])
        c += int(challenger[t] and not champion[t])
        n = b + c
        if n >= 10:
            p = sum(comb(n, i) for i in range(b, n + 1)) / 2 ** n
            if p < 0.01:
                return t + 1, b, c
    return None, b, c


at, served = canary()
print("\ncanary: 20% of rows go to the challenger")
if at:
    print(f"  first alarm at half-hour {at} ({at / 48:.1f} days)")
else:
    print(f"  no alarm in {WATCH:,} half-hours ({WATCH / 48:.1f} days)")
print(f"  rows the challenger answered: {served}")

at, b, c = shadow()
print("\nshadow: the challenger answers every row, silently")
if at:
    print(f"  first alarm at half-hour {at} ({at / 48:.1f} days)")
else:
    print(f"  no alarm in {WATCH:,} half-hours ({WATCH / 48:.1f} days)")
print(f"  rows where only the champion was right: {b}")
print(f"  rows where only the challenger was right: {c}")

The Lab Report

A real terminal recording headed python canary_report.py, titled every table in this lesson, from the stored files and the data. It opens with Elec2, 9,063 live half-hours, 1998-06-01 to 1998-12-06, and 123 checks against the refits and the stored files: all agree. Then nine numbered sections: 1, the main run, alarms of 20 and median half-hour for each challenger and rule; 2, the follow-up, with injected_2 at +1.75 and injected_5 at +4.09 and shadow at 20 of 20; 3, how many rows a canary sees: 110, 419 and 1,007, first checks at half-hours 702, 158 and 62; 4, head to head, same start; 5, the month table and the runs; 6, why shadow caught every injected regression; 7, what a canary costs before its alarm; 8, checking after every half-hour: canary 0.060, 0.098 and 0.094 against 0.009, 0.011 and 0.009 at the end only, and shadow 0.044 against 0.006, with 50 of 240, 49 of 391 and 35 of 375 canary false alarms at the first check; 9, review follow-ups: 159 and 371 of 6,822 right answers turned, the June starts, and shadow's median of 178.5 counted up to 2,000. Beneath: the lab's own report. It fits the same models again and stops unless every stored alarm comes back.

The report lives in scripts/labs/lifecycle/canary_report.py. It reads the lab's stored files, results/canary.json and results/canary_followup.json, lesson 1's stored file, and the Elec2 data from scikit-learn's local copy. The lab stored the half-hour of each first alarm, not each model's right and wrong rows, so the report trains the lab's four models again with the same data, split and settings, rebuilds the follow-up's challengers with its own seeds, and replays every rule on every start. It stops unless every stored first alarm, count, median and gap comes back exactly. It makes 123 checks in all, and they all agree. It changes nothing in the lab's files.

Its json mode writes every number to results/cn-report.json, which the figures read. The demo mode checks the student script's stored run, and the box mode writes the playground below and checks that it gives the lab's alarms.

Watch a Rollout Yourself

This box has no model in it. It holds the real right-or-wrong outcome of four models on all 9,063 live rows: the champion, the logistic model, and the follow-up's two injected challengers, one hexadecimal digit per row (counting in sixteens, with the digits 0 to 9 and a to f). It also holds the lab's own random canary draws. The report checked that every canary and shadow alarm it gives matches the lab's stored files. It runs in your browser.

As it is, the box prints the four rows of the lab for the 4.09-point regression: 6, 12, 15 and 20 alarms of 20, with the lab's medians. Try table('logistic', CHAMP, LOGISTIC) for the main run's big regression, and table('2%', CHAMP, INJ2) for the smaller one.

Then look for false alarms: table('the same model', CHAMP, CHAMP) gives the canary's false alarms on the champion against itself. Change the random draw with canary(CHAMP, CHAMP, 0, my_assign(50, 7)), and try a few seeds; None means no alarm in 2,000 half-hours. Every start in STARTS is a half-hour of the live period where the lab began watching.

The Code, Part by Part

Loading. fetch_openml("electricity", version=1) downloads Elec2 once and reads the local copy after that. The label column holds the words UP and DOWN, turned into 1 and 0. cut is 36,249: the champion learns from the rows before it, and the 9,063 rows after it are the live traffic.

The champion. One boosted-trees model with seed 0. champion is a list of True and False, one per live row: right or wrong. This list is all either rollout needs.

The challenger. spoil gives each live row a 5% chance of being marked, with the follow-up's own seed, 4. champion & ~spoil keeps the champion's result everywhere except on the marked rows, where a right answer becomes wrong and a wrong one stays wrong. The draw turned 371 of the champion's 6,822 right answers, 5.4%, which is why the challenger is worse by 4.09 points, and it is never right where the champion is wrong.

The canary draw. to_new marks the rows that go to the challenger: a random 20% of the first 2,000, with the lab's seed, 0.

canary. Walks through the half-hours in order. Each row adds to the counts of the model that answered it. Once both models have at least 30 rows, it computes the pooled error rate (pool, the share of all rows so far answered wrong), the typical wobble se, and the z value, and returns the first half-hour where z is above 2.33.

How to Roll Out a New Model

A hand-sketched column of six boxes joined by arrows, headed rolling out a new model, titled shadow first, then a canary, with the checks fixed in advance. 1, run the new model in shadow, on every row. 2, pair the rows: count where only one was right. 3, decide in advance when to look. 4, if shadow is clean, start a small canary. 5, watch what users do, not only the labels. 6, roll back on an alarm; log every decision. Beneath: here, on a known 4.09-point regression: shadow 20 of 20, canary 20% 12 of 20.

Run the new model in shadow first. Send every live row, or a large sample of rows, to both models. Show users only the champion's answer, and store the challenger's answer with the row.

Pair the rows. When each label arrives, mark both models right or wrong on that row. Count the rows where only the champion was right and the rows where only the challenger was right, and test that split with McNemar's test, as in lesson 4.

Decide in advance when to look. Choose the checkpoints before the rollout starts, for example once a day, and test only there, or use a method designed for repeated looks. Write the checkpoints, the bar and the minimum number of rows into the rollout's configuration.

If shadow is clean, start a small canary. Send a small share of real traffic to the challenger. Work out first how many labelled rows that share will give per day, and how long it will need to see the size of problem you care about.

Watch what users do. In the canary, measure the things shadow cannot see: clicks, returns, complaints, payments. Accuracy on labels is only part of the picture.

Roll back on an alarm, and log every decision. Keep the champion ready to take all traffic back at once: that is the roll back. Record each check, its counts and its result.

Shadow or Canary, When?

A two-column table headed grounded in this lesson's numbers, titled shadow or canary, when? Left, shadow fits when: you can afford to run two models on every row; the right answer arrives later, whatever answer was shown; a regression may be small: a 20% canary caught 1.75 points in 4 of 20; a wrong answer to a real user is costly. Right, a canary is needed when: what matters is how users react to the answer; the new model changes what users do next; two models on every row cost too much; you have shadow results and want real use. Beneath, left: shadow sees accuracy on the same rows. Beneath, right: a canary sees real effects on real people.

Use shadow when you can afford two models on every row. For small models, the extra compute is cheap. For large ones, shadow a sample of the traffic rather than none of it.

Use shadow when the right answer arrives later and does not depend on which answer was shown. Here the price goes up or down whatever the model says, so every row gets a label for both models, and that is what makes shadow possible to score. In many products the label exists only for the answer users saw: whether someone clicked a recommendation can only be known for the recommendation they were shown. There shadow cannot measure accuracy at all. And when a label takes weeks, every alarm waits for it, in shadow and canary alike.

Use shadow when the regression you fear is small. Here a 20% canary caught a 1.75-point regression in only 4 of 20 starts, where shadow caught it in all 20. The injected errors are the easiest case for shadow, as the earlier slide showed; the canary's count is affected much less by how the errors were made.

Use shadow when a wrong answer to a real user is costly. Shadow shows no one the new answers.

Use a canary when what matters is how users react. Shadow never shows its answers, so it can never see a click, a purchase or a complaint.

Use a canary when the model changes what users do next. If the answer changes the next request, the rows shadow sees are not the rows the new model would really meet.

Do not use a small canary to look for a small regression. Here a 5% canary answered 110 rows in almost 42 days.

Do not check a canary after every batch with a fixed bar. Here that raised the false alarms from about 1% to 6 to 10%.

What These Numbers Can and Cannot Tell You

A two-column page headed read before you trust these numbers, titled what these runs are, and what they are not. They are: one dataset: Elec2, June to Dec 1998; one row per half-hour; 20 starts that overlap; the design came first; injected and swapped: after; swapped rows are fair coins; months and runs: after. They are not: not a rate for other data; not a real service's volume; not 20 separate trials; two faults, kept on record; injected errors are random; shadow's false alarms: a best case; not a cause for August.

One dataset, one champion. Everything here is boosted trees on Elec2, June to December 1998. The results depend on how often the two models disagree and on how the data changes over the months. With other data they would move.

One row per half-hour. A real service may have far more labelled rows per hour. Read every time in this lesson as a number of rows, not as a time a real service would take.

Twenty starts that overlap. Each start watches 2,000 half-hours, and neighbouring starts are about 372 half-hours apart, so they share about 81% of their rows. The lab also used the same canary draw at every start. The counts out of 20 are counts, not 20 separate trials, and I did not test them.

Two faults, kept on the record. The planned small regression was 0.61 points better, and the planned equal challenger could not test shadow. The follow-up replaced them, and it was designed after the main results and written down before it ran.

Injected errors are random, and so are the swaps. They suit a paired test that treats rows as coin flips, and they never point the other way, which is the easiest case for shadow. The swapped equal pair likewise makes every disagreeing row a fair, separate coin, so shadow's 0 of 20 and 4.4% false alarms are a best case too. The real regression came in runs and in seasons. This lab did not measure how much that changes shadow's speed or false alarms.

After the results. The first allowed check, the head-to-head, the per-start and per-month gaps, the runs, the tenth-row alarm, the rows served and the repeated-checking rates were all measured after I saw the results. Why the logistic model fell behind from August is a guess, not a measurement.

What to Do Next

A hand-drawn list headed before your next rollout, titled five questions for your rollout. Same rows?: can both models answer the same live rows, so you can pair them? How many?: how many labelled rows will the new model see per day? How late?: how long until the right answer is known? When to look?: is the checking schedule fixed, or will you look after every batch? What else?: what do users do next, which accuracy alone cannot show? Beneath: here: a 5% canary answered 110 rows in 41.7 days.

Find how your team rolls out a new model, and ask the five questions on the card. If the answer to "same rows?" is yes and you are not running shadow, you are not using the fastest check this lab found. If the answer to "when to look?" is "whenever someone opens the dashboard", your false alarms are several times higher than the bar you think you set.

Then run a small version of this lab on your own traffic. Take your current model's stored right-and-wrong results on recent rows, make a known regression by turning a few percent of its right answers wrong, and measure how many rows your canary and your shadow check need to see it. Then do the same with a real older model in place of the injected one, since a real regression may not be as tidy.

The chapter plan's next lesson asks when to retrain: on a schedule or when something triggers it, and what each costs, on the same time-ordered data.

A closing card headed to keep, titled compare the same rows, and decide in advance when to look. In large type: first in 15 of 15. Beneath: the real 10.36-point regression: shadow raised the alarm before a 20% canary in every start where both did, and alone in 3 more. Then: on an injected 4.09-point regression: shadow 20 of 20, the canary 12. Injected errors are the easiest case for shadow. Then: pair the rows in shadow first, decide in advance when to look, then use a small canary to see what users do. Last: one dataset: a measurement here, not a rule for all data.

Knowledge Check

Knowledge Check

4 questions - Score 80% to pass

Q1

On a known 4.09-point regression, shadow raised the alarm in all 20 starts and a 20% canary in 12. What is the main reason shadow needs fewer half-hours?

Q2

The lab tested equal models after every half-hour instead of once at the end. What happened to a 20% canary's false alarms?

Q3

Shadow caught every injected regression at exactly the 10th disagreeing row. Why is that the easiest case for shadow?

Q4

What can a canary show that shadow cannot?

Accuracy
point

Four ways to watch: a canary with 5%, 20% or 50% of rows going to the challenger, drawn at random with seed 0, and shadow. The starts: a real rollout begins at some moment, and the moment matters, so the lab ran every rollout from 20 starting points spread over the live period. From each start it watched the next 2,000 half-hours, about 41.7 days, and recorded the half-hour of the first alarm, or no alarm. One model per challenger, one run each; the design described the counts and declared no test of them.

This is a real run in VS Code's terminal (python canary_demo.py).

A real screenshot of VS Code's terminal after running python canary_demo.py. It prints scikit-learn 1.9.1, 9,063 live half-hours; champion 0.7527, challenger 0.7118; the challenger is worse by 4.09 points. Then canary: 20% of rows go to the challenger; no alarm in 2,000 half-hours (41.7 days); rows the challenger answered: 419. Then shadow: the challenger answers every row, silently; first alarm at half-hour 378 (7.9 days); rows where only the champion was right: 10; rows where only the challenger was right: 0.

When I ran it, it printed scikit-learn 1.9.1 and 9,063 live half-hours; champion 0.7527 and challenger 0.7118, worse by 4.09 points; for the 20% canary, no alarm in 2,000 half-hours (41.7 days), with 419 rows answered by the challenger; and for shadow, a first alarm at half-hour 378 (7.9 days), with 10 rows where only the champion was right and 0 where only the challenger was. Both alarms match the lab's stored first start in canary_followup.json, and the longest printed line was 50 characters. The report's demo mode checks all of it.

This is one start, the first one, and it happens to be a start where the 20% canary missed. The canary caught this regression in 12 of the lab's 20 starts, so another start can go the other way. To see that, change every champion[t] and challenger[t] into champion[3345 + t] and challenger[3345 + t], and leave to_new[t] as it is. That watches from the lab's tenth start. When I did, the canary's first alarm came at half-hour 424 and shadow's at 174, the lab's stored values for that start. The other starts are listed in the playground below.

What came before the run, in canary_lab.py: the data, the champion, the three challengers, the four ways of watching, the two alarm tests, the 20 starts and the 2,000-half-hour horizon. What came after I saw the main results: the follow-up's injected regressions and swapped equal pair, designed after the results and written down before it ran. What came after all the results, in the report: the first allowed check, the head-to-head by start, the gap per start and per month, the runs, the tenth-row alarm, the rows a canary served, and the repeated-checking rates with 200 new seeds. After a review of this lesson, the report also counts how many right answers the injection really turned, which starts begin in June, where the null runs' false alarms fired, and shadow's median counted the same way as the canary's.

Four brand cards headed the tools, with their logos, titled what ran where. scikit-learn: the four models, the data download. pandas: the table of half-hours. NumPy: the canary draws, the injected errors. Python: both alarm tests: math, nothing else.

shadow. Walks through the same half-hours. b counts rows only the champion got right, c rows only the challenger got right. Once there are at least 10 such rows, it computes McNemar's p-value with math.comb and returns the first half-hour where p is below 0.01.

In the lab file, canary_lab.py does the same for three challengers in the main run and three in the follow-up, for four ways of watching and 20 starts, and stores every first alarm.