Ml Lifecycle

The Promotion Gate: How Often a New Model That Is Not Better Gets the Job

0 of 24 complete

0%

Contents

Back|Ml LifecycleThe Promotion Gate: How Often a New Model That Is Not Better Gets the Job
1/24
67 min left
Prerequisites
Picking the Best of Many: How Much of the Winner's Lead Survives on New Monthsrequired
Related Topics
The Score That Lied: A Random Split Against a Split in TimeWhy Production BreaksTomorrow Is Different: A Model Frozen on 2011, Scored Through 2012Why Production BreaksTrain/Serve Skew: One Input Computed Two WaysWhy Production BreaksFailing Silently: Which Checks Catch a Broken Input Before the Answers ArriveWhy Production BreaksWatching Inputs Before the Answers Arrive: Drift Measures Against the Real ErrorWhy Production Breaks
1 of 24

The Taste Test

Imagine a small cafe that has sold the same house coffee for years. A supplier offers a new blend. The owner does not want to change without a good reason, so she runs a taste test: ten regular customers try both cups without knowing which is which, and she will switch if more of them prefer the new one. Six of the ten pick the new blend. Does she switch?

Think about what the six really tells her. If the two blends tasted exactly the same, each customer would be picking at random, like tossing a coin. Six heads out of ten tosses happens all the time. So would five, and so would seven. Ten customers are simply too few to tell a real difference from chance.

An illustration of a woman in a blazer and glasses, standing and raising one hand, next to text. Headed the taste test, titled a new model must earn the job. Beside her: a gate compares the live model with a new one and promotes the new one if it scores higher. Two models that are equal by construction: higher wins promoted the new one in 0.43 to 0.51 of evaluation windows. A paired test on the rows where they disagree: 0.018 to 0.022. A real 6.7-point gain: every rule caught it in every window of 1,000 rows or more. Last: higher is not better until enough rows say so.

There is a second, quieter point in her test. Some customers like both cups equally, or dislike both. They tell her nothing about which blend is better. Only the customers who clearly preferred one cup carry any information. If she has 100 customers and 90 of them cannot tell the cups apart, her decision really rests on the other 10. Asking the same people about both cups, and counting only those who preferred one, is called a paired test: both things are judged on the same cases.

This lesson is about the same decision for a program that learns from examples. A new version of the program wants to replace the one in use. I measured how often three ways of deciding let a new version through when it was no better, and how often they caught one that really was.

Where This Lesson Starts

This chapter follows a model through its life, one step at a time, on the same data. Start with the dumbest model showed that a rule with no learning beat every trained model on the future. Same code, different model trained the same model 20 times, changing only the seed, and got 20 different models. Picking the best of many showed that the winner of a large search can do worse on later months than a setting picked at random.

Now the model is trained, and the question is whether it should replace the one already running. The course's lesson on training pipelines and orchestration has a slide called "The Gate: When to Ship, When to Hold". It explains, in words, the idea of a gate that compares the new model with the live one on held-out rows (rows neither model learned from), and a margin set in advance. I will not repeat it. This lesson measures how well such a gate works.

A flowchart headed step 6 of the life of a model: ship it, or not, titled where the promotion gate sits. Two boxes at the top, retrain: a new model, the challenger, and the live model, the champion, both lead to a diamond, the gate: same rows, a rule set in advance. From the diamond, passes leads to promote: the challenger serves from now on, and fails leads to keep the champion; discard the challenger. Beneath: the pipelines lesson explains this step in words. This lesson measures how often each kind of rule gets the decision wrong.

Lesson 2 matters most here. A seed is the starting number for a model's random choices; scikit-learn calls it random_state. Lesson 2 found that two seeds of the same code give two genuinely different models, sometimes several points apart. So a pipeline that retrains on a schedule produces a new, different model every time, even when nothing about the data or the code changed. Each of those models arrives at the gate. The gate is what decides whether that difference is progress or luck.

Eight Words for This Lesson

A hand-drawn list headed eight words for this lesson, titled deciding whether a new model replaces the old one. Champion: the model that is live now, answering real requests. Challenger: a newly trained model that wants the champion's job. Promotion gate: the rule that decides whether the challenger replaces the champion. Evaluation window: the rows both models are scored on for one decision. Naive rule: promote if the challenger scores higher, by any amount. Margin rule: promote only if it scores higher by at least 1 point. McNemar's test: a check that looks only at the rows where the two disagree. False promotion: promoting a challenger that is not better. Beneath: a gate is only as good as the rows it looks at.

The champion is the model that is live now, answering real requests. The challenger is a newly trained model that wants the champion's job. The promotion gate is the rule that decides whether the challenger replaces the champion. To promote a model is to put it live in place of the old one.

The gate scores both models on the same rows. I call one such set of rows an evaluation window, and N is the number of rows in it. Accuracy is the share of rows a model got right, and a point is one hundredth of accuracy.

The naive rule promotes the challenger if it scores higher, by any amount. The margin rule promotes it only if it is ahead by a fixed amount, here 1 point. McNemar's test is a check that looks only at the rows where the two models disagree; the next slides explain it. A false promotion is promoting a challenger that is not better. A null, in this lesson, means a pair of models where neither is better, so any promotion is a false one. A paired comparison scores both models on the same rows and compares them row by row, which is what all three rules here do.

Three Rules for One Decision

All three rules look at the same thing: two models, scored on the same window of rows, each row marked right or wrong for each model. They differ only in what they demand before they say yes.

A table of three rows headed the lab's three gate rules, fixed before the run, titled three ways to decide, on the same rows. Naive: promote if the challenger's accuracy is higher than the champion's on the window, by any amount. Margin: promote if it is higher by at least 1 percentage point. McNemar: promote if it is higher AND, counting only the rows where exactly one model was right, a split that uneven would come up less than 5% of the time by coin flips. Beneath: accuracy: the share of rows a model got right. One point: one hundredth of accuracy.

Naive. Higher wins. If the challenger got 158 of 200 rows right and the champion 157, the challenger is promoted. This is what a pipeline does when someone writes "if new score is greater than old score, deploy". It is a rule I have often seen in pipelines, because it is the first one anyone writes.

Margin. The challenger must be ahead by at least 1 point: on 200 rows, that is at least 2 more rows right. This is the rule the pipelines lesson describes, with a margin set in advance. The idea is sound: demand a little more than zero, so that tiny, lucky leads do not get through. The question this lesson asks is whether a fixed margin is enough, and the answer depends on the size of the window.

McNemar. The challenger must be ahead, and the lead must pass a test that asks: if the two models were equally good, how surprising would a split this uneven be? The test reads only the rows where the two models disagree. If that surprise is large enough, meaning such a split would happen less than 5% of the time between two equal models, the challenger is promoted.

All three rules were fixed in the lab file before it ran. None was tuned after I saw the results.

What McNemar's Test Counts

McNemar's test is named after the psychologist Quinn McNemar, who published it in 1947; this lesson uses its exact form, which counts coin tosses directly. It is built for exactly this situation: the same cases, judged twice, each time right or wrong. Here the cases are rows and the two judges are two models.

Put every row of a window into one of four boxes. Both models right. Both models wrong. Only the challenger right. Only the champion right. The first two boxes say nothing about which model is better: on those rows the two behaved the same. Only the last two boxes can tell them apart, just like the cafe customers who clearly preferred one cup.

A hand-drawn sketch headed sketched: the first 200 evaluation rows, seeds 0 and 1, real counts, titled only the rows where they disagree can tell them apart. A grid of four boxes, columns challenger right and challenger wrong, rows champion right and champion wrong. Both right: 149. Only champion: 8. Only challenger: 9. Both wrong: 34. Beneath the grid: the test reads only these 17 rows. Beneath the sketch: accuracy 0.785 against 0.790: one row apart. The naive rule promoted; the margin rule and McNemar (p 0.500) did not. Measured after the results.

This is a real window: the first 200 evaluation rows, with the trees trained with seed 0 as the champion and seed 1 as the challenger. They were both right on 149 rows and both wrong on 34. On 17 rows they disagreed: the challenger alone was right on 9, the champion alone on 8. So the challenger scored 0.790 and the champion 0.785, one row apart, and the naive rule promoted it.

Now the test. Suppose the two models were equally good. Then each of the 17 disagreeing rows is like a coin toss: as likely to go to one model as the other. The test asks how often 17 fair coin tosses give 9 or more heads. The answer, called the p-value, is 0.500: half the time. Nothing surprising at all, so McNemar said no.

Two panels headed McNemar on two real windows of 200 rows, titled the same question, two different answers. Seeds 0 and 1: 9 to 8, only challenger right to only champion right: p 0.500, a fair coin does this often. Trees and trees + lag: 16 to 4, p 0.006: a fair coin almost never splits 20 flips this unevenly. Beneath: p: the chance of a split at least this uneven if the two models were equal. The gate promotes below 0.05.

The second panel is the same first window for a pair where the challenger really is better: the trees with the previous half-hour's label as an extra input, from lesson 1. There the two disagreed on 20 rows, and the challenger was right on 16 of them. Sixteen or more heads in 20 fair tosses happens with a chance of 0.006, so McNemar said yes.

What the Lab Ran

I wrote the lab's design at the top of its file, gate_lab.py, before it ran. The dated record is in the chapter plan (2026-09-29, "GATE LAB (batch 4) designed before running").

An editorial page in four labelled zones, headed what the lab ran: gate_lab.py, designed before it ran, titled 21 pairs, three rules, three window sizes. The data: Elec2, as in lessons 1 and 2: train on the first 36,249 half-hours, evaluate on the last 9,063. 20 seed pairs: champion: boosted trees with random_state 2k; challenger: the same trees with 2k + 1, for k from 0 to 19. Meant as pairs where neither is better. One better pair: champion: the trees, seed 0 (0.7527); challenger: trees + lag (0.8193), from lesson 1. Windows: every non-overlapping block of 200, 1,000 and 5,000 evaluation rows, in time order: 45, 9 and 1 per pair. Beneath: the true null, McNemar on whole pairs and the window-by-window tables came after the results.

The data is Elec2 again: 45,312 half-hours of the New South Wales electricity market in time order, where each row asks whether the price is UP or DOWN against its average over the last 24 hours. Every model trains on the first 36,249 rows and is evaluated on the last 9,063.

The pairs. Twenty seed pairs: the same boosted trees trained with seeds 0 and 1, 2 and 3, and so on up to 38 and 39. The design called them null pairs, meant as pairs where neither model is better. The next slides show that this was wrong, and why. One better pair: the plain trees as champion against the trees plus lag as challenger, 0.7527 against 0.8193 on the whole evaluation, a gain of 6.7 points.

The windows. A real gate usually sees the most recent rows, so the lab cut the evaluation into contiguous blocks of N rows, meaning rows that sit back to back in time with no gaps, and with no overlap, for N of 200, 1,000 and 5,000. That gives 45 windows of 200 per pair, 9 of 1,000, and just 1 of 5,000. For every pair, rule and N, the lab stored the share of windows in which the challenger was promoted.

A sequence diagram with four columns: the window, champion, challenger, the gate. Headed one gate decision, as the lab ran it, titled both models, the same rows, one rule. Step 1, the window sends the champion N rows, inputs. Step 2, the window sends the challenger the same N rows. Step 3, the champion sends the gate right or wrong, per row. Step 4, the challenger sends the gate right or wrong, per row. Step 5, the gate applies its rule: promote? Beneath: both models were trained once, before any window. The windows move; the models do not.

The Main Run: Seed Pairs Got Through Often

Here is what the lab stored, in results/gate.json.

A two-column table headed the main run, gate.json: share of windows promoted, titled the seed pairs were promoted far more than 5% of the time. Left, pairs and rule; right, N = 200 / 1,000 / 5,000. 20 seed pairs, naive: 0.410 / 0.406 / 0.450. 20 seed pairs, margin: 0.329 / 0.256 / 0.350. 20 seed pairs, McNemar: 0.166 / 0.239 / 0.350. Better pair, naive: 0.867 / 1.000 / 1.000. Better pair, margin: 0.867 / 1.000 / 1.000. Better pair, McNemar: 0.578 / 1.000 / 1.000. Beneath, left: the seed pairs were not equal: on the whole evaluation, -3.2 to +7.2 points. Beneath, right: at N = 5,000 each pair has only one window.

For the seed pairs, the naive rule promoted the challenger in 0.41 to 0.45 of the windows, depending on N. That part was expected. The surprise was McNemar: 0.17 of windows at N = 200, 0.24 at 1,000, and 0.35 at 5,000. If these pairs were truly equal, McNemar should promote in about 5% of windows or fewer. It was promoting three to seven times as often, and more often with bigger windows, which is the wrong direction for a test that is merely noisy.

The better pair looked as it should. Every rule promoted it in every window of 1,000 and 5,000 rows. With 200 rows, naive and margin promoted it in 0.867 of windows and McNemar in 0.578.

Before reading anything into the seed pairs' numbers, I checked the pairs themselves. On the whole evaluation of 9,063 rows, the challenger minus the champion ranged from -3.2 to +7.2 points. Those are not equal models. So the main run's "false promotions" are a mix of two things: gate errors, and real differences between two seeds that the gate correctly noticed. The main run cannot separate the two, and I will not quote its McNemar numbers as false-promotion rates.

Two Seeds Were Two Different Models

This is lesson 2's finding, now showing up at the gate. Lesson 2 found that the default boosted trees hold back a random tenth of the training rows, chosen by the seed, so each seed trains on a different 90% of the data. The models that come out differ, sometimes by several points.

A dot chart headed each seed pair: the real gap on the whole evaluation, against how often naive promoted, titled the gate mostly followed a real difference. Twenty dots on a scale of challenger minus champion, points, all 9,063 rows, from -4 to 8, and share of 200-row windows promoted from 0 to 1, with a dotted vertical line at 0 labelled no difference. Dots to the left of the line sit mostly between 0.15 and 0.5; dots to the right sit mostly between 0.45 and 0.9, and the dot at about +7 sits near 0.87. Beneath: 20 pairs, naive rule, N = 200. Correlation +0.76. 12 of 20 pairs differ by more than 1 point. Measured after the results.

After the results, I plotted each seed pair's real gap on the whole evaluation against how often the naive rule promoted its challenger in 200-row windows. The two go together: the correlation, a number from -1 to 1 that measures how closely two things rise together, is +0.76. Pairs where the challenger really was better got promoted more; pairs where it was worse got promoted less. 12 of the 20 pairs differed by more than 1 point on the whole evaluation.

Then I ran McNemar's test on each pair's whole evaluation, all 9,063 rows, in both directions.

Three panels headed McNemar on all 9,063 rows, each seed pair, titled two seeds are two different models. Challenger better: 7 of 20, p below 0.05 one way. Champion better: 11 of 20, p below 0.05 the other way. Cannot tell: 2 of 20, neither p below 0.05. Beneath: two seeds disagree on 276 to 1,340 of the rows. Lesson 2: a seed changes which training rows the model learns from.

In 7 of the 20 pairs the challenger was clearly better, in 11 the champion was, and only in 2 could the two not be told apart. Two seeds disagreed on 276 to 1,340 of the 9,063 rows. So on this data, a retrain with a new seed is not "the same model again". It is a different model, and most of the time a measurably better or worse one.

That changes what the main run's McNemar numbers mean. With 5,000 rows, each pair had one window covering more than half of the evaluation. McNemar promoted 7 challengers there, and 6 of those 7 were among the 7 that were clearly better on the whole evaluation. That is mostly not a false promotion. It is an honest answer to a question the design got wrong: these were never null pairs.

For a team, the lesson is practical. A scheduled retrain can make the model worse, by several points, with no change anyone made. The gate is the only thing between that model and your users. To measure how often a gate is fooled, I needed pairs that were truly equal.

A True Null: Let a Coin Decide Each Row

I designed a follow-up after I saw the main results, and wrote its design into the lab file's followup mode before it ran. Its results are in results/gate_followup.json.

The idea is a way to build two models that are equal by construction. Take each seed pair. On every evaluation row, toss a coin. On heads, swap the two models' outcomes on that row: the champion gets the challenger's right-or-wrong, and the challenger gets the champion's. On tails, leave the row alone. After the swap, on every single row, each of the two "models" has the same chance of having the better outcome. Neither can be better by design, so any promotion is a false one.

A hand-drawn sketch headed sketched: the first 12 evaluation rows where seeds 0 and 1 disagree, real outcomes and real coins, titled a true null: a coin decides who gets each row. Five rows of 12 boxes. Champion: R seven times, then w five times. Challenger: w seven times, then R five times. Coin: S in the second, fourth, ninth and tenth boxes, blank elsewhere. Champion after: R, w, R, w, R, R, R, w, R, R, w, w. Challenger after: w, R, w, R, w, w, w, R, w, w, R, R. Beneath the boxes: R = right, w = wrong, S = swapped. Beneath the sketch: evaluation rows 13, 14, 16, 29, 31, 32, 37, 74, 120, 124, 134, 136, counting from 0. Where the two agree a swap changes nothing, so only these rows matter. Follow-up, designed after the results.

The sketch shows the first 12 evaluation rows where seeds 0 and 1 disagree, with their real outcomes and the real coin tosses. Where the two models agree, a swap changes nothing. Where they disagree, the coin decides which model gets the right answer. The disagreeing rows are exactly the ones McNemar reads, so this builds what McNemar's test assumes: that each disagreeing row is a fair coin.

The follow-up did this for all 20 seed pairs, with one fixed seed for the coins, and ran the same three rules on the same contiguous windows as the main run.

A bar chart headed follow-up, designed after the results: 20 true-null pairs, contiguous windows, titled with equal models, only McNemar stayed near 5%. Three groups of three bars, naive, margin and McNemar, at window sizes 200, 1,000 and 5,000 rows, on a scale from 0 to 0.6, with a dashed line at 0.05 labelled 5%. At 200: naive about 0.43, margin about 0.30, McNemar just above 0. At 1,000: naive about 0.51, margin about 0.09, McNemar just above 0. At 5,000: naive about 0.45; margin and McNemar show no bar. Beneath: naive 0.43 / 0.51 / 0.45; margin 0.30 / 0.09 / 0.00; McNemar 0.018 / 0.022 / 0.000. Windows: 900, 180, 20.

Why Higher Wins Is a Coin Flip

The naive rule's result has a simple cause. When two models are equally good, which one scores higher on a given window is decided by the luck of the rows in it. It is a coin flip, and a coin flip lands heads about half the time. That is exactly what "higher wins" did: it promoted an equal challenger in 0.43 to 0.51 of contiguous windows, whatever the size.

Three panels headed true null, 4,000 random draws of 200 rows; after the results, titled why higher-wins is a coin flip, not quite 50%. Challenger ahead: 43.4%, the naive rule promotes. Exactly tied: 12.7%, no promotion. Champion ahead: 43.9%, no promotion. Beneath: ties get rarer as windows grow: 12.7%, 5.4%, 2.5% at 200, 1,000 and 5,000 rows.

Why is it a little under half at 200 rows? Because of ties, which I counted on the follow-up's random draws. In 12.7% of 200-row draws the two equal models got exactly the same number of rows right, and a tie is not a promotion. The challenger was ahead in 43.4%, the champion in 43.9%. As windows grow, ties get rarer: 5.4% at 1,000 rows and 2.5% at 5,000. So the naive rule sits near 50% at every size. Bigger windows do not help it at all.

That is the key point about the naive rule. It does not become safer with more data. With 200 rows or 5,000, it promotes an equal model about half the time. A pipeline that retrains every week with this rule would, over a year, swap its model many times for no reason, and each swap is a new model with its own new mistakes that someone may have to debug.

It also means the naive rule cannot tell a small real gain from luck. If the challenger is really a little better, it will be ahead in somewhat more than half of windows, and behind in the rest. The rule gives no warning which case you are in.

The Margin Rule and the Size of the Window

The margin rule is the naive rule with a hurdle: be ahead by at least 1 point. It helped a great deal at 1,000 rows (0.09) and at 5,000 (0 of 20 contiguous windows, 0.02 of random draws), but not much at 200 (0.30). Why the difference?

Because a point is a different thing at each size. On 200 rows, 1 point is 2 rows. On 1,000 rows, it is 10 rows. On 5,000 rows, 50. And the luck of the rows moves the gap between two equal models by an amount that shrinks as the window grows.

A bar chart headed true null, random rows: how far apart two equal models land in one window, titled a fixed 1-point margin shrinks in meaning as N grows. Three bars at window sizes 200, 1,000 and 5,000 rows, on a scale of typical gap (standard deviation), points, from 0 to 2, with a dashed line at 1 labelled the 1-point margin. The bar at 200 reaches about 1.8, above the line; at 1,000 about 0.8, just below it; at 5,000 about 0.4. Beneath: standard deviation 1.8, 0.8, 0.4 points at 200, 1,000 and 5,000 rows. The margin promoted 0.31, 0.11, 0.02 of the random draws.

I measured that spread on the follow-up's random draws of equal models. The standard deviation, a measure of how far values typically land from their average, was 1.83 points at 200 rows, 0.82 at 1,000 and 0.40 at 5,000. So with 200 rows, two equal models typically land almost 2 points apart, and a 1-point hurdle is easily cleared by luck. With 5,000 rows they typically land less than half a point apart, and 1 point is rarely reached.

A fixed margin, then, means something different at every window size. It is not wrong; it is just not self-adjusting. If you set a margin, you have to choose it for the size of your window. A margin sized to stop luck at 200 rows, one that two equal models cross only about one time in twenty, is 1.645 × 1.83 = 3.01 points: 7 rows on 200, since part of a row does not count. I checked it on the report's data after the review: equal models passed it in 0.031 of the 900 contiguous windows, and the 6.7-point gain in 28 of 45 windows; a flat 3 points (6 rows) passed 32 of 45. So a margin that big would also stop most small real gains. That last step is reasoning, not measured: I had no pair with a small gain.

An isometric drawing of three blocks, heights to scale, headed the rows McNemar gets to read: the median over the contiguous windows of 20 true-null pairs; heights to scale, titled it reads a few percent of the window. A thin flat block labelled 8 rows of 200, a small cube labelled 43 rows of 1,000, and a tall block labelled 194.5 rows of 5,000. Beneath: the rows where exactly one model was right, the only rows McNemar reads: about 4% of each window here. Measured after the results.

A Real Gain Needs Enough Rows

A gate has two jobs. Stopping a challenger that is not better is one. Letting through one that is better is the other. The better pair, trees plus lag against the plain trees, is a real gain of 6.7 points on the whole evaluation: the challenger alone was right on 855 rows, the champion alone on 252.

A bar chart headed the better pair, window by window: 45 windows of 200 rows; after the results, titled ahead in 39 of 45 windows, behind in 5. Forty-five bars in time order, on a scale of challenger minus champion from -0.1 to 0.3, growing up or down from a dashed line at 0. Most bars are above the line, the tallest near 0.26 around window 10; a handful dip just below it, the lowest near -0.05 around window 20. Beneath: trees + lag against trees, 0.8193 against 0.7527 overall. Tied in 1. McNemar missed 19 of 45 windows; from 1,000 rows it missed none.

Window by window, with 200 rows the better model was ahead in 39 of the 45 windows, tied in 1 and behind in 5. The gap in one window ranged from -5.0 to +26.5 points. So even a model that is better by 6.7 points overall loses some short windows. The naive and margin rules promoted it in 0.867 of windows, and McNemar in only 0.578: it missed 19 of the 45. In 6 of those the challenger was not ahead at all (5 behind, 1 tied), so no rule could promote. In the other 13 its lead rested on too few disagreeing rows: a median of 18 across the missed windows, against 24 over all 45.

A two-column table headed the better pair, gate.json: share of windows promoted, titled a real gain needs enough rows to be seen. Left, rule; right, N = 200 / 1,000 / 5,000. Naive: 0.867 / 1.000 / 1.000. Margin: 0.867 / 1.000 / 1.000. McNemar: 0.578 / 1.000 / 1.000. Rows where they disagree, median: 24 / 109 / 715. Beneath, left: whole evaluation: only the challenger right on 855 rows, only the champion on 252. Beneath, right: at 5,000 rows there is one window, so 1.000 is one decision.

With 1,000 rows every rule caught it in all 9 windows; the smallest window gain was +3.6 points. With 5,000 rows, every rule caught it too, but that is a single window, so it is one decision, not a rate.

So a careful rule costs something. McNemar made far fewer false promotions than the other two, and in exchange it needed more rows to see a real gain. That is not a flaw to fix by lowering the bar. Every gate makes this trade, and the answer is more rows, not a looser rule. How many rows are enough depends on the size of the gain you care about: here McNemar caught a 6.7-point gain in 0.578 of 200-row windows and in all 9 windows of 1,000 rows; sizes in between were not tried. A 1-point gain would need many more, and I did not measure how many.

Neighbouring Rows: What This Lab Cannot Settle

The follow-up had a second question. A real gate sees the most recent rows, one block after another, and neighbouring half-hours are alike: lesson 1 found the label repeats the previous one 85.3% of the time. McNemar's test assumes each disagreeing row is a separate coin toss. If neighbouring rows go the same way, a block of them could look more uneven than chance allows, and the test would be fooled more often than 5% of the time. People call that inflating the test.

So the follow-up also scored the same true-null pairs on N rows drawn at random from the whole evaluation, 200 draws per pair and size. If McNemar were near 5% on random rows but higher on contiguous windows, the neighbouring rows would be to blame.

A table of four rows headed follow-up: contiguous windows against rows drawn at random, titled what the second comparison can and cannot settle. Contiguous: McNemar 0.018 / 0.022 / 0.000; 900, 180 and 20 windows. Random rows: McNemar 0.026 / 0.041 / 0.056; 4,000 draws each. At 5,000: 13 of 20 pairs ever promoted; the top two hold 78% of the promotions. The limit: the swap also breaks the link between neighbouring rows, so this does not show whether real windows of neighbouring rows inflate the test. Beneath: random draws of 5,000 from 9,063 share most rows: one pair counts almost once.

It came out the other way: contiguous windows promoted less often than random rows. But that does not show that contiguous windows are safe, and I only saw why after the run. The swap tosses a new coin on every row. So after the swap, which model got one row tells you nothing about which model got the next row. The swap removed the very link between neighbours that I wanted to test. So this comparison cannot tell whether real windows of neighbouring rows inflate the test. I leave that question open.

The random-row numbers have a second problem, which I found by reading the draws pair by pair. At 5,000 rows, one random draw takes more than half of the 9,063 evaluation rows. So the 200 draws of one pair share most of their rows, and they mostly give the same answer. McNemar promoted in 13 of the 20 pairs at least once, but two pairs held 78% of all its promotions. In those two pairs, the coins happened to leave the challenger ahead over the whole evaluation. So the 0.056 behaves more like "2 of 20 pairs" than like a rate from 4,000 separate tries.

Checking Again and Again

One more way a gate is fooled has nothing to do with the rule. A team trains a challenger, checks it against the champion, and it fails. A day later, with more rows, they check again. And again. Each check is a new chance for luck to push the challenger over the line, and the gate only needs to pass once.

Two panels headed true null, 20 pairs, checked after every 200 rows; after the results, titled checking again and again finds a winner that is not there. One check, at the end: 2 of 20, McNemar promoted, on the first 9,000 rows. 45 checks as rows arrive: 6 of 20, McNemar promoted at least once. Beneath: the naive rule: 13 of 20 at the end, 18 of 20 at some check. 20 pairs: a count, not a rate.

I measured this on the 20 true-null pairs. For each pair I checked the gate after every 200 new rows, each time on all the rows so far: 45 checks, up to 9,000 rows. With one check at the end, McNemar promoted 2 of the 20 equal challengers. Checking 45 times as the rows arrived, it promoted 6 of the 20 at some point. The naive rule went from 13 to 18 of 20. With 20 pairs these are counts, not rates, but the direction is the one the reasoning predicts: every extra look is another chance to be fooled.

The fix is to decide in advance when you will look, and to look once, or to use a method built for repeated looks, which demands a stricter bar at each look. A later lesson in this chapter measures this same effect on live traffic, where checking after every batch of requests is exactly what a team wants to do.

What I Would Do Instead

Here is what the lab points to, in the order a team meets it.

Score both models on the same rows. A paired comparison is only possible when both models answer the same rows. Keep a fixed evaluation set, or a fresh window, and score the champion again on it rather than reusing an old score from different rows.

Write the rule down before anyone sees a score. The rule, the threshold and the window size all belong in the pipeline's code or configuration, fixed in advance. Changing the rule after seeing that the challenger narrowly failed is a form of the winner's curse that lesson 3 measured: a choice made after seeing the scores flatters itself.

Use a paired test, not a raw comparison. McNemar's test for right-or-wrong answers counts only the rows where the two disagree, and it held near 5% on equal models here, where "higher wins" promoted about half of them. It needs nothing beyond counting and a few lines of arithmetic. Note what the swap test can and cannot show: under the swap, McNemar holding near 5% is what the maths requires, because the swap builds exactly the coin-flip assumption the test rests on. The lab's new evidence is about the other two rules. And even a test that holds at 5% is fooled about one time in twenty: a weekly retrain between equal models gives up to about 2.6 false promotions a year (52 × 0.05, arithmetic).

Choose the window size for the gain you care about. A small window cannot see a small gain, and a careful rule on a small window will say no to real improvements. Here McNemar caught a 6.7-point gain in 0.578 of 200-row windows and in all 9 windows of 1,000 rows; sizes in between were not tried. Decide what gain is worth promoting, then use enough rows to see it.

Check once, or correct for checking often. Every extra look at the same comparison is another chance for luck. Here, 45 looks turned 2 false promotions into 6.

Log the decision. Record the rule, the rows, the four counts and the p-value with every promotion. When someone asks later why a model went live, that record is the answer, and it makes a bad promotion something you can find and undo.

Try It Yourself

This script is the lab made small. It downloads the same data, trains the champion (the trees with seed 0) and two challengers (the trees with seed 1, like a scheduled retrain, and the trees plus lag), cuts the evaluation into windows of 1,000 rows, and prints each rule's decision in every window. It does not need a GPU.

A real screenshot of VS Code with gate_demo.py open, showing the docstring that says what the script is and how to run it, the imports, the lines that load Elec2, turn the label into 1 for UP and 0 for DOWN and add the previous label as an extra input, the cut at 80%, the right function that trains one model and marks each evaluation row right or wrong, and the p_value function for McNemar's test. Beneath: copy it from the box on the slide.

Before you run this lab. You need Python 3 and two libraries: pip install scikit-learn pandas. scikit-learn holds the model and the download; pandas holds the table. The first run downloads Elec2 from OpenML (under 1 MB compressed), so it needs an internet connection once; scikit-learn keeps a copy in a folder in your home directory (scikit_learn_data) for later runs. McNemar's test needs no extra library: it is a sum of counts from Python's own math.comb. I ran it with scikit-learn 1.9.1 on a Mac. scikit-learn runs the same way on Windows and Linux, but I have not checked the numbers there, and another version of scikit-learn may give different decimals, which is why the first line printed is the version.

"""The promotion gate: a champion, two challengers, three gate rules.

Lesson 4 of 'The ML & AI Lifecycle', made small. It needs Python 3 with scikit-learn
and pandas (pip install scikit-learn pandas). The first run downloads Elec2 from
OpenML (under 1 MB) and keeps a copy for later runs.
    python gate_demo.py

Author: Roni Das
Created: 2026-09-29
"""
from math import comb

import sklearn
from sklearn.datasets import fetch_openml
from sklearn.ensemble import HistGradientBoostingClassifier

# 45,312 half-hours in time order. The label: is the NSW price UP or DOWN
# against the average of the last 24 hours?
data = fetch_openml("electricity", version=1, as_frame=True, parser="auto").frame
X = data.drop(columns="class").astype(float)
y = (data["class"].astype(str) == "UP").astype(int).to_numpy()
X_lag = X.copy()
X_lag["prev_label"] = [y[0]] + list(y[:-1])   # the label of the half-hour before

cut = int(0.8 * len(y))            # train on the first 80%, evaluate on the rest
print(f"scikit-learn {sklearn.__version__}, {len(y) - cut:,} evaluation rows")


def right(inputs, seed):
    # train one model; True on each evaluation row it gets right
    model = HistGradientBoostingClassifier(random_state=seed)
    model.fit(inputs[:cut], y[:cut])
    return model.predict(inputs[cut:]) == y[cut:]


def p_value(b, c):
    # McNemar, exact, one-sided: chance of b or more heads in b + c coin flips
    n = b + c
    return 1.0 if n == 0 else sum(comb(n, i) for i in range(b, n + 1)) / 2 ** n


champion = right(X, 0)
challengers = {"trees, seed 1 (a retrain)": right(X, 1),
               "trees + lag, seed 0": right(X_lag, 0)}

for name, challenger in challengers.items():
    print(f"\nchallenger: {name}")
    print(f"whole evaluation: champion {champion.mean():.4f}, "
          f"challenger {challenger.mean():.4f}")
    print("window  champ  chall    b    c      p  naive  margin  McNemar")
    total = {"naive": 0, "margin": 0, "McNemar": 0}
    for w, start in enumerate(range(0, len(champion) - 999, 1000), 1):
        cm, ch = champion[start:start + 1000], challenger[start:start + 1000]
        b = int((ch & ~cm).sum())          # rows only the challenger got right
        c = int((cm & ~ch).sum())          # rows only the champion got right
        gap = ch.mean() - cm.mean()
        p = p_value(b, c)
        said = {"naive": gap > 0, "margin": gap >= 0.01,
                "McNemar": gap > 0 and p < 0.05}
        for rule in said:
            total[rule] += said[rule]
        yes = ["yes" if said[r] else "no" for r in said]
        print(f"{w:>6}  {cm.mean():.3f}  {ch.mean():.3f}  {b:>3}  {c:>3}  {p:.3f}"
              f"  {yes[0]:>5}  {yes[1]:>6}  {yes[2]:>7}")
    print(f"promoted: naive {total['naive']} of {w}, margin {total['margin']} of {w},"
          f" McNemar {total['McNemar']} of {w}")

The Lab Report

A real terminal recording headed python gate_report.py, titled every table in this lesson, from the stored files and the data. It opens with 345 checks against the refits: all agree, then prints seven numbered sections: 1, the main run, with the seed pairs at 0.410, 0.406 and 0.450 for naive and the better pair at 0.867, 1.000 and 1.000; 2, the seed pairs are different models: challenger better in 7 of 20, champion better in 11, cannot be told apart in 2; 3, the true null, contiguous windows and random rows; 4, what the rates hide: ties, the standard deviation of the gap, 1 point in rows, the disagreeing rows in random draws and in contiguous windows (8, 43, 194.5), and a margin sized for 200 rows, 3.01 points, which the better pair passed in 28 of 45 windows; 5, the better pair window by window, with McNemar's 19 misses split into 6 where the challenger was not ahead and 13 where it was; 6, one window worked; 7, checking again and again: McNemar 6 of 20 at some check against 2 at the end, then two lines of arithmetic: 7 of 8 disagreeing rows for p 0.035, and 2.6 false promotions a year for 52 weekly gates. Beneath: the lab's own report. It fits the same models again and stops unless every stored number matches.

The report lives in scripts/labs/lifecycle/gate_report.py. It reads the lab's stored files, results/gate.json and results/gate_followup.json, lessons 1 and 2's stored files, and the Elec2 data from scikit-learn's local copy. The lab stored promotion rates, not each model's right and wrong rows, so the report trains the lab's 41 models again with the same data, split and settings. It stops unless every stored number comes back exactly: each model's accuracy on the whole evaluation, every pair's rate for every rule and window size, and the follow-up's true-null rates, which it rebuilds with the follow-up's own coin seed and draw seed. It makes 345 checks in all, and they all agree. It changes nothing in the lab's files.

Its json mode writes every number to results/gt-report.json, which the figures read. The demo mode checks the student script's stored run, and the box mode writes the playground below and checks that it gives the lab's rates.

Run the Gate Yourself

This box has no model in it. It holds the real right-or-wrong outcome of seven models on all 9,063 evaluation rows: the trees with seeds 0 to 5 and the trees plus lag. Each row's seven outcomes are packed into two hexadecimal digits (counting in sixteens, with the digits 0 to 9 and a to f), and the right function unpacks one model's column. The report checked every rate it prints against gate.json. It runs in your browser.

As it is, the box prints the promotion rates of three pairs at each window size, exactly the lab's: seeds 0 and 1, seeds 4 and 5 (a pair where the challenger was 7.2 points better on the whole evaluation) and the trees plus lag.

Now build a true null yourself: a, b = swapped(right(0), right(1)) and then rates(a, b). The box tosses its coins with Python's own random numbers, not the lab's, so it lands near the follow-up's rates, not on them; try a few values of seed. Then try peek(a, b), which checks McNemar after every 200 rows and returns the first row count at which it promoted, or None, and compare peek(a, b, rule='naive').

The Code, Part by Part

Loading. fetch_openml("electricity", version=1) downloads Elec2 once and reads the local copy after that. The label column holds the words UP and DOWN, turned into 1 and 0. X_lag is a copy of the inputs with one extra column, the previous half-hour's label, as in lesson 1.

The cut. cut = int(0.8 * len(y)) is 36,249. Every model learns from the rows before it and is evaluated on the 9,063 rows after it.

right. Trains one boosted-trees model with the given seed and returns, for every evaluation row, True if the model got it right. This one list per model is all a gate needs: the champion's list and the challenger's list, row by row.

p_value. McNemar's test, exact and one-sided. b is the number of rows only the challenger got right and c the number only the champion got right. It adds up the ways b + c fair coin tosses could give b or more heads, using math.comb, and divides by all the ways. No statistics library is needed.

The windows. range(0, len(champion) - 999, 1000) gives the start of each 1,000-row window, nine of them. For each window, marks the rows where only the challenger was right, and the rows where only the champion was.

How to Set Up a Promotion Gate

A hand-sketched column of six boxes joined by arrows, headed setting up a promotion gate, titled decide the rule before you see the scores. 1, score both models on the same rows. 2, write the rule down before any score. 3, use a paired test: count the disagreements. 4, size the window for the gain you care about. 5, check once, or correct for checking often. 6, log the decision, the rows and the counts. Beneath: here, with 200 rows, the naive rule promoted equal models in 0.43 of windows and McNemar in 0.018.

Score both models on the same rows. Keep the champion's predictions for the evaluation rows, or score it again, so that every row has two outcomes side by side.

Write the rule down before any score. Put the test, the 0.05 line and the window size in the pipeline's configuration, and review changes to it like changes to code.

Use a paired test. For right-or-wrong answers, McNemar's test. Count the four boxes, keep the two that disagree, and compute the p-value.

Size the window. Decide the smallest gain worth a promotion, then check, with a pair of models you know differ by about that much, that your window catches it most of the time. If it does not, collect more rows before deciding.

Check once, or correct. Decide when the gate runs. If it must run again and again as rows arrive, use a method built for repeated looks rather than the same 0.05 line each time.

Log the decision. The rule, the rows, the four counts, the p-value and the outcome, next to the model's run receipt from lesson 2, the record of code, seed, data, settings and versions that lets a run be repeated.

Which Rule, When?

A two-column table headed grounded in this lesson's numbers, titled which rule, when? Left, a paired test fits when: both models can be scored on the same rows; one right-or-wrong answer per row; you must tell a small gain from luck (with enough rows); you must defend each promotion later. Right, it is not enough when: the two models never see the same rows, as in a live split; the score is not right or wrong per row, such as an average error; you check again after every batch of new rows; rows come in long runs of alike rows: untested here. Beneath, left: pair the rows. Count the disagreements. Beneath, right: then ask whether the rows are independent.

Use McNemar's test when both models answer the same rows, right or wrong. That is the usual offline gate: a fixed evaluation set, or the latest window of labelled rows, scored by both models. Here it held at or below about 5% on equal models. Under the swap that is what the maths requires, since the swap builds the test's own assumption; the lab's new evidence is how badly the other two rules did.

Use it when the gain you want is small. The naive rule cannot tell a small gain from luck, and a fixed margin mostly measures the window size. A paired test adapts to the rows that matter.

Use it when you will need to explain a promotion. The four counts and a p-value are a record anyone can check.

Do not use it when the two models never see the same rows. In a live test where each user is sent to one model or the other, there are no paired rows; that needs an unpaired comparison, which compares two separate groups of rows instead of the same rows twice, and a later lesson in this chapter deals with live traffic.

Do not use it as it stands when the score is not right or wrong. For an average error or a ranking score, a paired test for numbers is needed instead; McNemar's test only counts disagreements.

Do not rely on it alone when the rows come in runs. This lab did not settle whether neighbouring rows fool the test. If your rows are strongly alike in time, spread the window over a longer period or check the test on a pair you know to be equal, as the follow-up did.

Do not use the naive rule for anything that matters. Here it promoted an equal model in about half of all windows, at every size.

What These Numbers Can and Cannot Tell You

A two-column page headed read before you trust these numbers, titled what these runs are, and what they are not. They are: one dataset: Elec2, 1996 to 1998; one model family, one better pair; 20 seed pairs, 20 true nulls; the rules came first; the true null came after; the swap is one way to build equal models. They are not: not a rate for other data; not a test of other scores; 5,000-row results: few windows; not tuned to the results; not changed after it ran; not proof for clustered rows.

One dataset, one model family. Everything here is boosted trees on Elec2. The rates depend on how often two models disagree, which was 3.0% to 14.8% of rows for these pairs. With models that disagree more or less, the numbers would move. The better pair is one gain of one size, 6.7 points.

The main run's pairs were not null. The design called the seed pairs null; the results showed they were not. I have kept their numbers but not called them false-promotion rates. The true null is a follow-up, designed after the main results and written down before it ran.

The swap is one way to build equal models. It makes the two models equal row by row, and in doing so it also breaks any link between neighbouring rows. Real equal models could make errors in runs, and this lab cannot say what that does to the test.

Few windows at 5,000 rows. Each pair has one 5,000-row window, so those numbers are 20 decisions, and the random-row draws of one pair share most of their rows. Read the 5,000-row numbers as counts, not rates.

After the results. McNemar on each pair's whole evaluation, the ties, the standard deviation of the gap, the rows each window disagrees on, the better pair window by window, the worked windows and the repeated checks were all measured after I saw the results. The repeated checks cover 20 pairs: a count, not a rate.

No test of the rates themselves. The design described rates and declared no significance test on them. The p-values in this lesson belong to the gate's own decisions, not to the comparison of rules.

What to Do Next

A hand-drawn list headed before your next promotion, titled five questions for your gate. Same rows?: were both models scored on exactly the same rows? Rule first?: was the rule written down before anyone saw a score? Paired?: does the rule count the rows where the two disagree? Enough rows?: is the window big enough to see the gain you care about? How often?: how many times was the gate checked before it passed? Beneath: here: naive promoted equal models in 0.43 of 200-row windows.

Find the rule your pipeline uses to promote a model, and ask the five questions on the card. If the rule is "the new score is higher", replace it. If it has a margin, find out how many rows it is applied to and whether the margin means anything at that size.

Then run a small version of this lab on your own data. Take your current model and a retrain with a different seed, score both on your evaluation rows, and count the rows where they disagree. That one number tells you how many rows your gate is really deciding on. Then build a true null by swapping outcomes at random, as the follow-up did, and see how often your rule promotes it.

The next lesson planned for this chapter moves the gate into production. A canary release sends a small share of live traffic to the new model; a shadow release runs the new model on live traffic without using its answers. That lesson measures how much traffic it takes to see that a new model is worse.

A closing card headed to keep, titled higher is not better until enough rows say so. In large type: 0.43 against 0.018. Beneath: two equal models, 200-row windows: how often higher-wins and McNemar promoted the challenger. Then: score both on the same rows, count the rows where they disagree, fix the rule and the window first, and check once. Last: one dataset, 20 pairs: a way to check, not a law.

Knowledge Check

Knowledge Check

4 questions - Score 80% to pass

Q1

Two models that are equal by construction are compared on 200-row windows, and the one with the higher accuracy is promoted. How often was the challenger promoted here?

Q2

Which rows does McNemar's test use to decide?

Q3

With equal models, the 1-point margin promoted the challenger in 0.30 of 200-row windows but 0.00 of 5,000-row windows. Why?

Q4

In the main run, McNemar promoted the second seed of a pair in up to 0.35 of windows, far above 5%. What explained most of that?

Two words about what the test is not. It is one-sided: it only asks whether the challenger is better, never whether it is worse, because a gate only promotes in one direction. And 0.05 is a line chosen in advance, not a law of nature. It says: I accept being fooled about one time in twenty when the two models are really equal.

In every window, both models answer exactly the same rows. McNemar's test needs that: it must know, row by row, which model was right. Each seed gave one model, trained once. The design reported the share of windows promoted and planned no significance test on those shares; a significance test is a calculation of how likely a difference is to come from chance alone.

With two models that were equal by construction, here is how often each rule promoted the challenger, at N = 200, 1,000 and 5,000:

RuleContiguous windowsRows drawn at random
naive0.43 / 0.51 / 0.450.43 / 0.49 / 0.53
margin0.30 / 0.09 / 0.000.31 / 0.11 / 0.02
McNemar0.018 / 0.022 / 0.0000.026 / 0.041 / 0.056

The naive rule promoted an equal challenger in roughly half of all windows, at every size. The margin rule depended heavily on the size: 0.30 of 200-row windows, but 0.09 of 1,000-row windows. McNemar stayed at or below about 5% throughout. The last column, rows drawn at random, is the second half of the follow-up, and it has a limit of its own that a later slide explains.

McNemar adjusts by itself, because it counts the rows that matter. In the follow-up's contiguous windows, the two models disagreed on a median of only 8 rows in a 200-row window, 43 in 1,000 and 194.5 in 5,000. With 8 disagreeing rows, a split of 6 to 2 is not rare enough for the test; it takes 7 to 1. That is plain arithmetic of coin tosses, not a lab measurement: 7 or more heads in 8 tosses comes up with a chance of 0.035, 6 or more with 0.145. With so few rows the test cannot land exactly on 5%, and it stays below it, which is why its rate in contiguous 200-row windows was 0.018 rather than 0.05.

This is a real run in VS Code's terminal (python gate_demo.py).

A real screenshot of VS Code's terminal after running python gate_demo.py. It prints scikit-learn 1.9.1, 9,063 evaluation rows. Then, for the challenger trees, seed 1 (a retrain): whole evaluation: champion 0.7527, challenger 0.7464; nine window lines with the two accuracies, b, c, p and three yes or no decisions; window 4, 0.821 against 0.822 with b 16 and c 15, is promoted by naive only; window 5, 0.800 against 0.819 with b 31 and c 12 and p 0.003, is promoted by all three; promoted: naive 3 of 9, margin 1 of 9, McNemar 1 of 9. Then, for trees + lag, seed 0: champion 0.7527, challenger 0.8193; nine windows, every one yes for every rule; promoted: naive 9 of 9, margin 9 of 9, McNemar 9 of 9.

When I ran it, the two whole-evaluation accuracies of each pair and the number of windows each rule promoted matched the lab's stored gate.json, and every one of the 18 window lines matched the report's refit. The report's demo mode checks all of it. Look at window 4 of the retrain: the challenger was right on 822 rows and the champion on 821, one row, and the naive rule promoted it. Window 5 is the retrain's one McNemar promotion, and it is a real result: on those 1,000 rows, the seed 1 model was clearly better. On the whole evaluation it was worse, 0.7464 against 0.7527, which is exactly the lesson of the seed pairs.

What came before the run, in gate_lab.py: the data, the 20 seed pairs and the better pair, the three rules, the window sizes and the stored rates. What came after I saw the main results: the follow-up's true null and random rows, designed after the results and written down before it ran. What came after all the results, in the report: McNemar on each pair's whole evaluation, the disagreeing rows, the ties, the standard deviation of the gap, the per-pair split of the random-row rates, the better pair window by window, the worked windows and the repeated checks.

Four brand cards headed the tools, with their logos, titled what ran where. scikit-learn: the 41 tree models, the data download. pandas: the table of half-hours. NumPy: the windows, the swaps, the draws. Python: McNemar's test: math.comb, nothing else.

ch & ~cm
cm & ~ch

The three rules. gap > 0 is naive, gap >= 0.01 is the 1-point margin, and gap > 0 and p < 0.05 is McNemar. These are the lab's own rules, written the same way.

In the lab file, gate_lab.py does the same for 20 seed pairs and the better pair, at 200, 1,000 and 5,000 rows, and stores the share of windows each rule promoted in results/gate.json.