Ml Lifecycle

Picking the Best of Many: How Much of the Winner's Lead Survives on New Months

0 of 23 complete

0%

Contents

Back|Ml LifecyclePicking the Best of Many: How Much of the Winner's Lead Survives on New Months
1/23
63 min left
Prerequisites
Same Code, Different Model: What Changes Between Two Identical Training Runsrequired
Related Topics
The Score That Lied: A Random Split Against a Split in TimeWhy Production BreaksTomorrow Is Different: A Model Frozen on 2011, Scored Through 2012Why Production BreaksTrain/Serve Skew: One Input Computed Two WaysWhy Production BreaksFailing Silently: Which Checks Catch a Broken Input Before the Answers ArriveWhy Production BreaksWatching Inputs Before the Answers Arrive: Drift Measures Against the Real ErrorWhy Production Breaks
1 of 23

The Tryout on a Windy Day

Imagine a running club that wants to sign one new runner for the autumn season. Two hundred people come to a one-day tryout in spring. Each runs one race, and the club signs whoever is fastest. It sounds fair. But that day there was a strong wind, and it blew harder at some moments than at others. The fastest runner may simply have had the wind at her back. The more people who run, the more likely it is that the fastest time is partly wind.

There is a second problem, and it is quieter. The tryout was on a short, flat track in spring. The season is in autumn, on long hilly courses. A runner who is great on a flat sprint may be ordinary on a hill. So the club can be wrong twice: once because the winning time was partly luck, and again because the tryout tested a different kind of race from the one that matters.

An illustration of a young man in a hooded top and dark trousers, standing with his hands in his pockets, next to text. Headed the tryout on a windy day, titled the more settings I tried, the worse the winner did later. Beside him: 200 random settings of the same boosted trees; keep the one with the best validation score, then check it on later months. One setting at random: validation 0.703, test 0.690. Best of 200: validation 0.718, test 0.664. The lazy rule from lesson 1 on the same test: 0.848. Last: trying more raised the winner's validation score and lowered its test score.

This lesson does the same thing to a program that learns from examples. I tried 200 versions of one program and kept the one that looked best. Then I checked it on later months that had played no part in the choice. The winner looked better than a random pick on the months I chose with, and worse on the months that came after. The rest of the lesson measures how much of that was the wind, and how much was the wrong track.

Where This Lesson Starts

Two words first. Accuracy is the share of guesses that were right, and a point is one hundredth of accuracy, so going from 0.70 to 0.73 is a gain of 3 points. This chapter follows a model through its life. Start with the dumbest model showed that on this data a baseline, a simple rule that a model must beat, beat every trained model on the future. The rule, called persistence, has no learning at all: it says whatever the last half-hour was. Same code, different model trained one model 20 times with the same code and got 20 different models, because a default setting set aside a random tenth of the training rows. The seed alone moved accuracy by 7.3 points.

That lesson ended with a question. If a team tries many versions and keeps the best, how much of the winner's lead is real? That is what this lesson measures.

A flowchart headed inside step 4 of the life of a model, titled try many settings, keep one, then judge it once. Five boxes joined by arrows: train each setting on the first 60%; score each on validation, the next 20%; keep the best on validation; score that one on the test, the last 20%; compare with the baseline. A dotted arrow from the last box back to the first is labelled NO: never tune again after the test. Beneath: the test is looked at once, at the end. Going back to tune after seeing it turns the test into a second validation set.

The pipelines lesson, training pipelines and orchestration, has a slide called "Experiment Tracking: Every Run Logs What It Did". It explains, in words, how a tracking store records every run so that a team can sort the runs by a score and find the best one. I will not repeat it. This lesson measures what happens next: when you sort 200 runs by a score and pick the top one, what that top score is worth.

Eight Words for This Lesson

A hand-drawn list headed eight words for this lesson, titled choosing the best of many runs. Setting: a choice made before training, such as how fast the trees learn; also called a hyperparameter. Tuning: trying many settings and keeping the one that scores best. Validation set: rows held back from training, used only to choose between settings. Test set: rows held back from everything, looked at once, at the very end. K: how many settings were tried before one was kept. Winner's curse: the best of many scores is partly luck, so it falls on new data. Rank correlation: do two scores put the settings in the same order? 1 yes, 0 no relation. Drift: the data changing over time, so later months are not like earlier ones. Beneath: a score used to choose a setting cannot also be the score you report.

A model is a program that learned from examples. Before it learns, someone has to make some choices that the learning itself does not make: how big each step of learning is, how complicated the model may get, how long it trains. Each of these is a setting. In machine learning they are usually called hyperparameters, to tell them apart from the numbers the model learns by itself. Tuning means trying many settings and keeping the one that scores best.

To choose fairly, you need rows the model did not learn from. The validation set is a block of rows held back from training and used only to compare settings. The test set is a second block, held back from everything, including the choosing, and looked at once at the end. K is how many settings were tried before one was kept.

The winner's curse is the name for a simple fact: when you pick the best of many noisy scores, the best one is usually lucky, so it scores lower when you measure it again. A rank correlation asks whether two scores put things in the same order: 1 means the same order, 0 means no relation at all, and a negative number means a reversed order. The kind used here is Spearman's rank correlation, which compares the two orders, not the raw scores. Drift means the data changing over time.

One word needs care, because lesson 2 used it differently. There, the validation slice was a random tenth of the training rows that early stopping set aside inside one training run. Here, the validation set is a separate block of later months, outside training, used to compare whole settings. They share a name and not much else.

Three Piles of Data, Three Jobs

Lesson 1 cut the data into two parts, training and test. Tuning needs three, because it adds a job: choosing. Each pile gets exactly one job, and the jobs must not mix.

An editorial page in four labelled zones, headed the data: Elec2 again, cut into three parts in time, titled 45,312 half-hours: learn, choose, judge. Train, the first 60%: 27,187 rows, 1996-05-07 to 1997-11-24; every setting learns from these only; lessons 1 and 2 trained on the first 80%. Validation, the next 20%: 9,062 rows, 1997-11-24 to 1998-06-01; used only to choose the setting. Test, the last 20%: 9,063 rows, 1998-06-01 to 1998-12-06; the same test rows as lessons 1 and 2; looked at once. The question in every row: is the NSW price UP or DOWN against its average over the last 24 hours? Beneath: persistence, the lazy rule of lesson 1: validation 0.8695, test 0.8484. Dates rebuilt from row numbers.

Training rows are what each setting learns from. Validation rows are for choosing: every setting is scored on them, and the best is kept. Test rows are for judging the one you kept. The data is Elec2 again, 45,312 half-hours of the New South Wales electricity market in time order, and each row asks whether the price is UP or DOWN against its own average over the last 24 hours. Because the rows are in time order, the three parts are cut in time: training is the oldest 60%, validation the next 20%, and test the newest 20%, the same test rows as the two lessons before.

Keep one thing in mind all through the lesson. Every one of the 200 settings learns from the first 60% only, so its scores sit below those in lessons 1 and 2, which trained on the first 80%. The default trees with early stopping off, the same model as lesson 2's, scored 0.7520 on test when trained on 80% there, and 0.7077 when trained on 60% here.

Why must the test be touched only once? Because the moment you use a score to make a choice, that score stops being a fair estimate. If I look at the test, go back, change a setting, and look again, I am choosing on the test, and the test has become a second validation set. Its number then carries the same optimism as the validation number, and I have no clean number left to report.

Why not just choose on the test and skip validation? For the same reason. The score you choose on is the score that flatters you. You need one set to choose on and a different set to find out what the choice is worth.

What the Lab Ran

I wrote the lab's design at the top of its file, select_lab.py, before it ran. The dated record is in the chapter plan (2026-09-29, "SELECT LAB (batch 3) designed before running").

A sequence diagram with four columns: the lab, the dice, the trees, the scorer. Headed what the lab ran, once per setting, 200 times, titled only the setting changes from run to run. Step 1, the lab asks the dice to draw a setting. Step 2, the dice send back six random values. Step 3, the lab tells the trees to learn from the first 60%. Step 4, the trees send the scorer their guesses for the next 20%. Step 5, the trees send their guesses for the last 20%. Step 6, the scorer computes validation and test. Beneath: the dice use seed 0; the trees use random_state 0 with early stopping off, so a setting gives one model. Then 1,000 draws of K settings, seed 1.

The model is scikit-learn's HistGradientBoostingClassifier, the boosted trees from lessons 1 and 2: many small trees of yes-or-no questions, built one after another, each correcting the mistakes of the ones before. The 200 settings were drawn at random from fixed ranges, written down before the run. Early stopping, a rule that stops adding trees once a held-out slice of the training rows stops improving, was switched off. The trees used random_state=0, scikit-learn's name for the seed of their random choices. Two of the ranges are log scale, which means each tenfold step is equally likely: a learning rate between 0.01 and 0.03 is as likely as one between 0.1 and 0.3. Lesson 2 found that with early stopping off, its trees came out the same whatever the seed, and here the seed is fixed anyway. So each setting gives exactly one model, and the only thing that varies from run to run is the setting.

A table of six rows headed what the dice could choose, fixed before the run, titled six settings, each drawn at random. learning_rate: 0.01 to 0.3, even on a log scale: how big a step each new tree takes. max_leaf_nodes: 8 to 64: how many end points one tree may have. max_depth: none in 30% of draws, else 3 to 12: how many questions deep. min_samples_leaf: 5 to 200: the fewest rows an end point may hold. l2_regularization: 0.0001 to 10, log scale: a brake on extreme answers. max_iter: 50 to 400: how many trees in all. Beneath: early stopping off, so the setting is the only thing that varies.

The Winner Got Worse the More I Tried

Here is the headline, from the lab's stored file, results/select.json.

A line chart headed the winner of K random settings, mean of 1,000 draws, from select.json, titled the winner's validation rose; its test fell. Two lines over K = 1, 5, 20, 50, 200 on a scale from 0.65 to 0.73, and a dashed line at 0.692 labelled median test of all 200. Validation, solid, rises from about 0.703 to about 0.718. Test, dashed, starts just below the median line and falls to about 0.664. Beneath: validation 0.703, 0.711, 0.714, 0.716, 0.718. Test 0.690, 0.689, 0.678, 0.670, 0.664. Gap 0.013 at K=1, 0.054 at K=200. Persistence on the test: 0.848, off this scale.

With K = 1, no choosing at all, the kept setting scored 0.7034 on validation and 0.6904 on test, a gap of 0.0130. With K = 200, the best of all of them, it scored 0.7175 on validation and 0.6639 on test, a gap of 0.0536. In between, the pattern is steady: the more settings I tried, the higher the winner's validation score and the lower its test score. From K = 1 to K = 200, validation rose by 1.4 points and test fell by 2.7 points.

Look at where the test line ends. The median test score of all 200 settings, the middle one if you line them up, was 0.6922. The winner of 200 scored 0.6639 on test: it came 173rd of the 200. Choosing did not just fail to help here. It picked a setting that was worse on the future than a setting picked with no care at all. With K = 20 the winner fell below that median in 71% of the draws, and with K = 50 in 92%.

And the lazy rule from lesson 1, persistence, which says whatever the last half-hour was, scored 0.8484 on the same test rows. The best test score of any of the 200 settings was 0.7502. No setting came close.

Validation Said Almost Nothing About the Test

Choosing on validation can only work if a setting that is good on validation tends to be good on test. So I put each setting's two scores side by side.

A dot chart headed all 200 settings, one dot each: validation score against test score, titled validation said almost nothing about the test: rank correlation 0.03. Two hundred dots on a scale of validation accuracy, November 1997 to June 1998, from 0.68 to 0.72, and test accuracy from 0.60 to 0.76. The dots form a loose cloud with no upward slope; the few dots at the far right, near 0.715 on validation, sit between about 0.65 and 0.69 on test. A dotted vertical line near 0.7175 is labelled best on validation. Beneath: validation ran 0.684 to 0.718; test 0.618 to 0.750. The best on validation scored 0.664 on test, rank 173 of 200. Among the 50 best on validation, the rank correlation was -0.45: measured after the results.

The rank correlation between validation and test, over all 200, was 0.03: almost no relation. Knowing a setting's validation score told me close to nothing about its test score. Worse, among the 50 settings that did best on validation, the rank correlation was -0.45, a clear reversed order: in this top group, the better a setting looked on validation, the worse it tended to do on test. That second number I measured after I saw the results.

Two more things stand out in the stored scores. The settings were close on validation, 0.6841 to 0.7175, a range of 3.3 points. On test they were spread four times as wide, 0.6182 to 0.7502. So the validation months could barely tell the settings apart, while the test months separated them strongly, and in a different order.

The setting that was best on test, number 143, scored 0.7502 there. Its validation score was 0.7013, below the validation median of 0.7039. Nobody choosing on validation would ever have kept it.

The Winner's Curse, and Why It Happens

Here is the general idea, before I split this lab's result into its parts. Think of every score as two pieces added together: how good the setting really is, and luck. On any finite set of rows, some settings happen to get lucky rows and some get unlucky ones.

When you pick the highest of many scores, you are more likely to pick one whose luck was good. On new rows the luck is new, so the winner's score usually falls back. The more settings you try, the more extreme the luckiest one, and the bigger the fall. That is the winner's curse, and it happens even when nothing about the world changes.

A hand-drawn sketch headed sketched: the student script's 20 settings, real scores, titled first on validation, eighth on test. Two rows of 20 boxes. The top row, val, is numbered 1 to 20 left to right: each column is one setting, placed by its validation rank. The bottom row, test, holds each setting's rank on test: 8, 1, 15, 5, 6, 20, 14, 11, 13, 16, 17, 4, 12, 3, 9, 10, 7, 18, 19, 2. Beneath the boxes: the best on validation came 8th on test; the worst on validation came 2nd on test. Beneath the sketch: the winner: validation 0.7110, test 0.6989. The runner-up: 0.7100 and 0.7358.

The sketch is the student script's 20 settings, the lab's first 20, with real scores. The winner on validation came 8th of 20 on test. The runner-up, just 0.001 behind it on validation, was the best of the 20 on test. And the setting that was worst on validation came 2nd on test. With 20 settings this is one draw, not a rule. The lab's 1,000 draws of 20 are the rule: on average, the winner of 20 scored 0.6784 on test, below the 0.6922 median.

Two panels headed one setting picked at random, against the best of 200, titled choosing raised the score I could see and lowered the one I could not. K = 1, no choosing: 0.703 / 0.690, validation / test; gap 0.013. K = 200, best of all: 0.718 / 0.664, validation / test; gap 0.054. Beneath: from 1 to 200: validation up 1.4 points, test down 2.7. The best of all 200 on validation came 173rd of 200 on test.

But notice something the plain curse does not explain. When validation and test measure the same thing, luck on its own should make the winner's test score fall back toward the average, not past it. A lucky winner is still, on average, at least as good as a random pick. Here the winner of 200 fell well below the median. Something more than luck is at work, and the next slides find it.

Validation and Test Are Different Seasons

The validation rows are November 1997 to the first morning of June 1998. The test rows are June to December 1998. They are different months, and on this data that matters. I compared the two parts after I saw the results.

A two-column table headed the validation months against the test months, measured after the results, titled validation and test are different seasons. Left, the measure; right, validation; test. Dates: Nov 1997 to Jun 1998; Jun to Dec 1998. Share of rows UP: 0.389; 0.451. Share of rows that flip: 0.131; 0.152. Mean NSW price, scaled: 0.0399; 0.0671. Persistence accuracy: 0.8695; 0.8484. Mean of the 200 settings: 0.7035; 0.6901. Beneath, left: one random setting loses this much between the two parts: the gap at K=1. Beneath, right: mean gap 0.0134 with no choosing at all.

The two parts differ in plain ways. In the validation months the price was UP on 38.9% of rows; in the test months, 45.1%. The average scaled price was 0.0399 in validation and 0.0671 in test, and 0.0608 over the training rows here. Lesson 1 gave 0.0556 for its training rows because it averaged the first 80%, which includes the calm, cheap validation months. The label changed from one half-hour to the next, a flip, on 13.1% of validation rows and 15.2% of test rows. Even the lazy rule noticed: persistence scored 0.8695 on validation and 0.8484 on test.

So some of the gap is there before any choosing. The average setting scored 0.7035 on validation and 0.6901 on test, a gap of 0.0134, almost exactly the 0.0130 gap at K = 1. That part is not the winner's curse at all. It is drift: the test months were simply harder for these trees than the validation months.

A line chart headed accuracy month by month over the validation and test months; after the results, titled from July, the validation winner fell behind. Two lines over fourteen months from November 1997 to December 1998 on a scale from 0.4 to 0.9, with a dotted vertical line labelled test starts between May and June. Mean of 200 settings, solid, and the validation winner, dashed. Up to June the dashed winner sits slightly above the solid mean in every month; both dip together in July to about 0.59; from September the dashed line falls well below the solid one, to about 0.63 in November and 0.62 in December. Beneath: worst test month for the winner against the mean setting: Nov 1998, 0.628 against 0.715. Only June holds rows of both parts; Nov 1997 and Dec 1998 are partial months.

The month-by-month numbers show more. In every validation month, the validation winner scored at or above the average setting, as you would expect from the setting chosen on those months. In June, the first test month, it was still ahead, 0.698 against 0.682. From July it fell behind, most of all in October, 0.675 against 0.725 for the average setting, and November, 0.628 against 0.715. July was hard for every setting (0.591 on average), the same July that lesson 1's trees broke down in.

A Control: The Same Months on Both Sides

Two causes are mixed in the main result: selection luck, and the change of months between validation and test. The stored scores alone could not separate them, so I added a follow-up to the lab. I designed it after I saw the main results, and I wrote its design into the lab file's followup mode before it ran. Its results are in results/select_followup.json.

The idea is a control: the same experiment with one thing changed. I trained the same 200 settings again on the same training rows (every validation and test score came back identical to the stored ones, to twelve decimal places). Then I made a control split by days: I pooled the validation and test months together and assigned each whole day, at random, to one side or the other: half the days became a control validation set, the other half a control test set. Now both sets come from the same months. If the main gap were selection luck, it should appear here too. I did this with 20 different random splits.

A line chart headed follow-up, designed after the results: the winner's test score, two ways of splitting, titled with the same months on both sides, the winner's test rose. Two lines over K = 1, 5, 20, 50, 200 on a scale from 0.64 to 0.74. Different months (main), solid, falls from about 0.690 to about 0.664. Same months (control), dashed, rises from about 0.697 to about 0.727. Beneath: gap at K=1 and K=200: main 0.0130 and 0.0536; same months -0.0004 and -0.0022, the mean of 20 random splits by whole days. This first control pools the test months, where settings really differ more, so it understates luck.

With the same months on both sides, the rank correlation between control validation and control test was 0.89 on average (0.71 to 0.96 across the 20 splits), against 0.03 in the main split. Choosing worked: the winner's control test score rose with K, from 0.6971 at K = 1 to 0.7265 at K = 200, almost exactly as fast as its control validation score rose. The gap stayed near zero at every K: -0.0004 at K = 1 and -0.0022 at K = 200, on average over the 20 splits.

Three panels headed where the gap at K=200 comes from, measured on the stored scores, titled most of the growth came when the months changed. No choosing (K=1): 1.3 points: later months are harder. Luck alone: 0.9 points of growth, 1 to 200, validation days split in two. Choosing, new months: 4.1 points of growth, 1 to 200. Beneath: total gap at K=200: 5.4 points. Luck alone is a second follow-up: choose on half the validation days, check on the other half. About a fifth of the growth.

What the Validation Months Rewarded

If the two parts reward different settings, it should show in which settings did well on each. So I measured, after the results, the rank correlation of each of the six settings with the validation score and with the test score, over all 200.

A two-column table headed rank correlation of each setting with the score, all 200; after the results, titled the two seasons rewarded different settings. Left, the setting; right, with validation; with test. Learning rate: -0.08; +0.51. Min samples leaf: +0.16; -0.31. Max leaf nodes: +0.26; +0.06. Max iter: -0.13; +0.13. L2 regularization: +0.06; -0.08. Max depth (147 with a limit): +0.49; +0.21. Beneath, left: 1 = a higher value always scored higher; 0 = no relation. Beneath, right: 200 settings, one run each: a pattern, not a proof.

The clearest line is the learning rate, how big a step each new tree takes. On validation it made almost no difference (-0.08). On test, faster learners did clearly better (+0.51). min_samples_leaf, the fewest rows an end point of a tree may hold, went the other way: slightly good on validation (+0.16), bad on test (-0.31). The top five settings on validation all had small learning rates (0.012 to 0.059) and large leaves (123 to 197 rows). That is a description of these 200 settings, not a proof.

Why would the test months reward faster, finer trees? One possible reason, and it is a guess I have not tested: the test months had higher prices and more flips than the validation months, and the training period, so a model that adapts more sharply to its inputs may cope better with prices it rarely saw. The validation months had low, calm prices, where a slow, smooth model loses little. I measured the price levels and flip rates; I did not measure the reason.

How Noisy Is One Score? The Link to Lesson 2

Lesson 2 found that the seed alone moved one setting's accuracy by 7.3 points. That is a lot of noise. How much noise is there in one validation score here?

Three panels headed how much one validation score moves, by resampling its days; after the results, titled the lead held up across resampled days; which setting came first did not. Winner's lead: 1.4 points over the middle setting. Resampled days: 0.4 to 2.5, the lead in 95% of 1,000 resamples. Wins again: 28% of resamples; 44 different winners. Beneath: standard deviations: all 200 settings on validation 0.6 points; lesson 2's 20 seeds of one setting 1.7. Here early stopping was off, so the seed did not move them.

I measured it with a bootstrap, after the results: make 1,000 new samples of the 190 validation days, each drawn at random with repeats allowed, and score the settings on each sample. Whole days are drawn, not single half-hours, because neighbouring half-hours are alike. A standard deviation measures how far values usually sit from their average. Across the samples, the winner's validation score had a standard deviation of 0.0140: most samples put it within about 1.4 points of its real value.

Its lead over the middle setting was 1.37 points, and that lead held up across resampled days: in 95% of the samples it stayed between +0.43 and +2.46. But which setting came first did not hold up. A second follow-up picked the best of all 200 again on each of 1,000 resampled sets of validation days: the original winner won again in only 28% of them, and 44 different settings won at least once. So the lead of one chosen pair was steady, while the choice itself was fragile, and neither said anything about the test months.

Now the link to lesson 2, comparing like with like. Here I switched early stopping off, so each setting gave one model and the seed changed nothing. The 200 settings' validation scores had a standard deviation of 0.62 points. In lesson 2, with early stopping at its default, the 20 seeds of one setting had a standard deviation of 1.72 points. If I had left the default on here, each setting's score would have carried seed noise almost three times as large as the usual difference between settings, and picking the top score would mostly have picked the luckiest seed. That is a reasoned expectation from two measured numbers; this lab did not run it.

Other Ways to Choose

The follow-up also tried two other ways of choosing, both designed after the main results.

Two windows. Each setting was also trained on the first 40% of rows and scored on the next 20% (May to November 1997), an earlier window. Its selection score was the average of that earlier score and its normal validation score. This is a small version of a forward scheme: judge each setting on more than one later window, each time training only on what came before, the way the model will really be used.

A line chart headed follow-up, designed after the results: the winner's test score by how it was chosen, titled two windows kept the winner near or above the median. Two lines over K = 1, 5, 20, 50, 200 on a scale from 0.64 to 0.72, and a dashed line at 0.692 labelled median test. One window (main), solid, falls from about 0.690 to about 0.664. Two windows, dashed, rises to about 0.704 at K = 5 and 20, stays near 0.702 at K = 50, and falls back to about 0.691 at K = 200, just under the median line. Beneath: two windows: the mean of the score on rows 40 to 60% (trained on the first 40%) and on the validation rows. Winner's test at K=200: 0.691, against 0.664 with one window.

It helped, partly. With two windows the winner's test score was 0.7037 at K = 5 and 0.7038 at K = 20, above the 0.6922 median, against 0.6892 and 0.6784 with one window. At K = 200 it fell back to 0.6906, about the median. The earlier window happened to agree with the test months better than the validation months did: its rank correlation with test was 0.62, and with validation -0.11. That is luck of this calendar as much as a rule; with other data the earlier window could be the one that misleads.

The reverse. Choosing on the test months and checking on the validation months, from the stored scores, gave the mirror image: the winner's validation score fell from 0.7034 at K = 1 to 0.6980 at K = 20, below the validation median of 0.7039. So the mismatch runs both ways. It is not that the test months are special; it is that these two seasons reward different settings.

An isometric drawing of six blocks, heights to scale above 0.5, headed accuracy on the same test rows, titled within these ranges, trained on 60%, no setting came near the lazy rule. From left to right: persistence 0.848, the tallest; best on test 0.750; default, not tuned 0.708; median setting 0.692; random setting 0.690; best of 200 0.664, the shortest. Beneath: all six trained on the first 60%. Best on test is unknowable in practice. The random setting and the best of 200 are means of 1,000 draws.

One more comparison matters most. The best of the 200 settings on the test itself is a number you could only find by choosing on the test, which breaks the rule of this lesson. Even that setting stayed far below persistence. Within these ranges, trained on 60%, tuning could not have closed the gap. And the default setting, not tuned at all and trained on the same 60%, scored 0.7077 on test: above a random setting and above the average winner at every K.

What I Would Do Instead

Here is what the lab points to, in the order a team meets it.

Hold out a final test set, and touch it once. In this lab that is what made the damage visible at all. The validation score of the best of 200 was 0.7175, the best number in the whole lab. Only the untouched test showed 0.6639.

Report the winner's test score, not its validation score. The validation score of a winner is the score you chose with, so it is always too high: it was picked because it was high. Here the report would say 0.664, not 0.718.

Keep count of how many settings you tried. K belongs next to the score. A winner of 5 and a winner of 200 do not mean the same thing: here their validation scores differed by only 0.7 points, and their test scores by 2.5. An experiment tracker records every run, as the pipelines lesson explains; the count is simply the number of rows in it.

Prefer a smaller or simpler search. On this data, when you choose on one window of changing data, a winner of 5 did about as well on test as a random pick (0.6892 against 0.6904), and every larger search did worse. A default setting that you did not tune has no winner's curse at all: here it scored 0.7077 on test, above the average winner at every K.

Choose on more than one window. For data in time order, a forward scheme scores each setting on several later windows, each trained only on what came before. In general, nested validation puts the whole search inside an outer loop: an inner split chooses, an outer split judges the choice, and the outer score includes the cost of choosing. Here two windows kept the winner near or above the median; I did not test more windows or a full nested scheme.

Why a forward or nested scheme costs more. Each extra window means training every setting again on a different stretch of the past, so two windows cost twice the training of one, and a full nested scheme costs the whole search once for every outer split. For trees that train in a few seconds, like these, that is nothing. For a model that trains for a day, it is a real budget, and a smaller search with more windows is often a better use of it than a larger search with one. The extra training buys two things: a choice made on more than one season, and a final score that already counts the effect of choosing, instead of one that hides it.

Compare the winner with the baseline, on the test. Here persistence scored 0.8484 and the tuned winner 0.6639. A search can only choose between the settings you give it; it cannot tell you that a rule with no learning is better.

Try It Yourself

This script is the lab made small. It downloads the same data, draws the lab's first 20 settings with the same seed, trains each on the first 60%, and prints each setting's validation and test score. Then it prints the winner on validation, the median test score of the 20, and persistence on the same test. It does not need a GPU.

A real screenshot of VS Code with select_demo.py open, showing the docstring that says what the script is and how to run it, the imports, the lines that load Elec2 and turn the label into 1 for UP and 0 for DOWN, the two cuts at 60% and 80%, the seeded random generator, the log_uniform and draw functions that pick one random setting; the loop over 20 settings is further down. Beneath: copy it from the box on the slide.

Before you run this lab. You need Python 3 and three libraries: pip install scikit-learn pandas scipy. scikit-learn holds the model, the scores and the download; pandas holds the table; SciPy is the library the lab's report uses for the rank correlations, and scikit-learn needs it anyway. The first run downloads Elec2 from OpenML (under 1 MB compressed), so it needs an internet connection once; scikit-learn keeps a copy in a folder in your home directory (scikit_learn_data) for later runs. I ran it with scikit-learn 1.9.1 on a Mac. scikit-learn runs the same way on Windows and Linux, but I have not checked the numbers there, and another version of scikit-learn may give different decimals.

"""Picking the best of many: 20 random settings, one winner, then the test.

Lesson 3 of 'The ML & AI Lifecycle', made small. It needs Python 3 with scikit-learn,
pandas and SciPy (pip install scikit-learn pandas scipy). The first run downloads
Elec2 from OpenML (under 1 MB) and keeps a copy for later runs.
    python select_demo.py

Author: Roni Das
Created: 2026-09-29
"""
import numpy as np
from sklearn.datasets import fetch_openml
from sklearn.ensemble import HistGradientBoostingClassifier
from sklearn.metrics import accuracy_score

# 45,312 half-hours in time order. The label: is the NSW price UP or DOWN
# against the average of the last 24 hours?
data = fetch_openml("electricity", version=1, as_frame=True, parser="auto").frame
X = data.drop(columns="class").astype(float)
y = (data["class"].astype(str) == "UP").astype(int).to_numpy()

n = len(y)
a, b = int(0.6 * n), int(0.8 * n)  # train, then validation, then test, in time
rng = np.random.default_rng(0)     # the lab's seed: the lab's first 20 settings


def log_uniform(low, high):
    # a random number between low and high, spread evenly on a log scale
    return float(np.exp(rng.uniform(np.log(low), np.log(high))))


def draw():
    # one random setting: how fast the trees learn, how big they grow, how many
    s = {"learning_rate": log_uniform(0.01, 0.3),
         "max_leaf_nodes": int(rng.integers(8, 65))}
    s["max_depth"] = None if rng.random() < 0.3 else int(rng.integers(3, 13))
    s["min_samples_leaf"] = int(rng.integers(5, 201))
    s["l2_regularization"] = log_uniform(1e-4, 10)
    s["max_iter"] = int(rng.integers(50, 401))
    return s


scores = []
for k in range(20):
    model = HistGradientBoostingClassifier(early_stopping=False, random_state=0,
                                           **draw())
    model.fit(X[:a], y[:a])                               # learn: first 60%
    val = accuracy_score(y[a:b], model.predict(X[a:b]))   # choose: next 20%
    test = accuracy_score(y[b:], model.predict(X[b:]))    # judge: last 20%
    scores.append((val, test))
    print(f"setting {k:>2}  validation {val:.4f}  test {test:.4f}")

win = max(range(20), key=lambda k: scores[k][0])   # the best on validation
print(f"winner on validation: setting {win}")
print(f"  its validation {scores[win][0]:.4f}, its test {scores[win][1]:.4f}")
print(f"median test of all 20: {np.median([t for _, t in scores]):.4f}")
print(f"persistence on the test: {(y[b - 1:-1] == y[b:]).mean():.4f}")

The Lab Report

A real terminal recording headed python select_report.py, titled every table in this lesson, from the stored files. It opens with 453 checks: all agree, then prints seven numbered sections: the split in time; the 200 settings, with rank correlation +0.028; the winner of K settings, from 0.7034 and 0.6904 at K = 1 to 0.7175 and 0.6639 at K = 200; the same-months follow-up, rank correlation +0.891; the table by month and each setting's correlations; the noise in one score, sd 0.0140; two windows, the reverse, and persistence 0.8484. Beneath: the lab's own report. It trains nothing and stops unless every stored guess gives the stored score.

The report lives in scripts/labs/lifecycle/select_report.py. It reads the lab's stored file, results/select.json, the follow-up's results/select_followup.json, lessons 1 and 2's stored files, and the Elec2 labels from scikit-learn's local copy. It trains nothing. The follow-up stored every setting's guesses on the validation and test rows, so the report recomputes every score from those guesses and the labels, and stops unless each one equals the stored score to twelve decimal places. It also replays the lab's 1,000 draws for every K with the same seed and stops unless it gets the stored numbers. It makes 453 checks in all, and they all agree. It changes nothing in the lab's files.

Its json mode writes every number to results/wc-report.json, which the figures read. The demo mode checks the student script's stored run, and the box mode writes the playground below and checks it.

What came before the run, in select_lab.py: the split, the ranges of the six settings, the 200 draws, the 1,000 selection draws per K, the rank correlation and the persistence line. What came after I saw the results: the same-months control, the two-window score and the reverse (the follow-up), and in the report the rank correlation of the top 50, the table by month, the correlations of each setting, the bootstrap, and the comparison with lesson 2.

Play the Tuning Game Yourself

This box has no model in it. It holds the real validation and test score of all 200 settings, and the same 200 models scored on the two halves of the same months, from the follow-up's first random split. It runs in your browser.

As it is, the box prints the rank correlation of validation and test, +0.028, the median test score, 0.6922, and persistence, 0.8484, then plays the tuning game 1,000 times for each K. The rank correlation and the K = 200 line are exactly the lab's, and the report's box mode checks that. The other lines use Python's own random draws instead of the lab's, so they land near the lab's numbers, not on them: within 0.0011 in my check.

Try curse(20, same_months=True) to play the same game when validation and test come from the same months. These are the first control's scores, split 1, which pools the test months and so understates luck. This split starts at a gap of about -0.03 at every K, because its two halves differ in difficulty, so compare the change from curse(1, same_months=True) to curse(20, same_months=True), not the gap itself. Try curse(5) against curse(200), and rank_corr(SAME_VAL, SAME_TEST). Then try a different seed, curse(20, seed=7), and see how much one set of 1,000 draws moves.

The Code, Part by Part

Loading. fetch_openml("electricity", version=1) downloads Elec2 once and reads the local copy after that. The label column holds the words UP and DOWN, turned into 1 and 0. The eight inputs become plain numbers with astype(float), as in the lab.

The two cuts. a is 60% of the rows, 27,187, and b is 80%, 36,249. Training is everything before a, validation is from a to b, and test is from b to the end. Nothing is shuffled.

The random settings. rng = np.random.default_rng(0) is the lab's random generator with the lab's seed, so the script draws exactly the lab's first 20 settings. log_uniform picks a number evenly on a log scale, so a value between 0.01 and 0.03 is as likely as one between 0.1 and 0.3. draw asks for the six values in the lab's order; the order matters, because each call takes the next numbers from the generator.

The loop. Each setting builds a fresh HistGradientBoostingClassifier with early_stopping=False and random_state=0, learns from the first 60%, and is scored on validation and on test. Both scores are printed, one line per setting.

How to Tune Without Fooling Yourself

A hand-sketched column of six boxes joined by arrows, headed tuning without fooling yourself, titled from the first cut to the report, in order. 1, cut train, validation, test by time. 2, lock the test away before tuning. 3, decide how many settings to try. 4, choose on validation, over several windows. 5, score the winner on the test, once. 6, report K, the test score, the baseline. Beneath: here step 6 said: best of 200, test 0.664; persistence 0.848.

Cut the data by time, into three parts. If rows arrive in time order, train on the oldest, choose on the next, and judge on the newest, the way the model will be used.

Lock the test away before you start. Decide which rows are the test before you look at any score, and do not look at the test again until the choice is final.

Decide how many settings to try, and write it down. There is no safe number; here a winner of 5 already did no better on test than a random pick. Every extra setting is another chance for a lucky or ill-suited winner.

Choose on more than one window. For data in time order, score each setting on two or more later windows and average them. Check that the windows agree: if the rank correlation between them is near zero or below, as it was here between the earlier window and validation (-0.11), the choice means little. The playground below lets you check the orders yourself.

Score the winner on the test, once. That is the number to report.

Report K, the winner's test score and the baseline together. "Best of 200, test 0.664, persistence 0.848" tells a reader far more than "validation 0.718".

When a Wide Search Pays, and When It Does Not

A two-column table headed grounded in this lesson's numbers, titled when a wide search pays, and when it does not. Search more when: separate choosing windows agree on the order of settings; the validation set is large next to the gaps you chase; the data does not change much over time; you keep a test set that was never used to choose. Search less, or not at all, when: validation says little about the test: rank correlation 0.03 here; the settings differ by little: 3.3 points from worst to best on validation here; the months you choose on are not the months you serve; a lazy rule already wins: 0.848 against 0.750 at best. Beneath, left: every number here is one dataset, one model family. Beneath, right: here, more tries found settings that suited the wrong months.

Use a wide search when the choosing rows are like the serving rows. In the same-months control, validation and test agreed (rank correlation 0.89) and the winner's test score rose with K, from 0.697 to 0.727. There, more tries really helped.

Use it when the settings differ by more than the noise. A setting that is ten points better will almost always win, however noisy the scores are. A setting that is half a point better will often lose to luck.

Use it when you keep a clean test set. Then even a bad search is caught before it ships, as it was here.

Do not search widely when the data drifts and you choose on one window. Here the best of 200 came 173rd of 200 on the future.

Do not search widely when the settings barely differ. Here all 200 spanned 3.3 points on validation. A search over settings that close is mostly a search over noise and over the quirks of the validation months.

Check the baseline before a wide search. If a rule with no learning is far ahead, as persistence was here, a better setting will not close the gap; a better framing or better inputs might.

What These Numbers Can and Cannot Tell You

A two-column page headed read before you trust these numbers, titled what these runs are, and what they are not. They are: one dataset: Elec2, 1996 to 1998; one model family, 200 settings; one run per setting, seed 0; the design came first; control and two windows: a follow-up; months and correlations: after. They are not: not a law for other data; not a comparison of search methods; no significance test; not refitted on train plus validation; day halves are not independent; why the seasons differ: a guess.

One dataset, one model family. Everything here is boosted trees on Elec2, with one split in time. Here luck alone was about a fifth of the growth, 0.9 of 4.1 points. On data that does not drift, luck may be most of it. I have not tested other data.

Random search only. The lab drew settings at random. Smarter search methods, which pick the next setting based on the last ones, were not tested, and this is not a comparison of them.

The winner was not retrained. In practice many teams retrain the winning setting on training plus validation rows before testing it. The lab scored the model trained on the first 60%, as its design said. A retrained winner might score differently.

Design first, follow-up after. The lab's design was written before it ran. The same-months control, the two-window score and the reverse are a follow-up, designed after I saw the main results. The top-50 correlation, the month table, the setting correlations and the bootstrap were measured after the results. The second follow-up, the luck control within the validation months, the resampled winners and the untuned default, was designed after the first follow-up, prompted by a review. The controls' day halves come from neighbouring days, so they are more alike than truly new data and may understate luck, and the first control also pooled the test months' real differences into both halves.

Guesses stay guesses. Why the test months rewarded faster learning rates is my guess, and I have not tested it. The effect of leaving early stopping on is a reasoned expectation from lesson 2's numbers, not something this lab ran.

What to Do Next

A hand-drawn list headed before you believe a tuned score, titled five questions for your own search. How many?: how many settings, seeds or models were tried before this one? Which set?: was the reported score measured on the rows used to choose? Same time?: do the choosing rows come from the months the model will serve? How noisy?: how much does one score move with the seed or the days? Baseline?: does the winner beat the lazy rule on the test? Beneath: here: best of 200 on validation 0.718, on test 0.664, persistence 0.848.

Find the last model your team tuned and ask the five questions on the card. The first two are usually answered by the experiment tracker: count the runs, and check which rows the reported score came from. If the reported score is the validation score of the best run, you do not yet know how good the model is. Score it once on rows nobody chose with.

Then check the third question with a small version of this lab: score your top few settings on two different windows of your validation data, and see if they come out in the same order. If they do not, the choice between them is not telling you much, and a wider search will not fix that.

The next lesson planned for this chapter moves from choosing a model to promoting one: the gate a pipeline uses to decide whether a newly trained model replaces the one in use, and how often a model that is no better gets through.

A closing card headed to keep, titled the best on validation is not the best on the future. In large type: 0.718 → 0.664. Beneath: the best of 200 settings on validation, then on the test. One setting picked at random: 0.703 → 0.690. Then: keep a test set you never choose on, count what you tried, and report the winner's test score next to the baseline. Last: one dataset, one model family: a way to check, not a law.

Knowledge Check

Knowledge Check

4 questions - Score 80% to pass

Q1

A team tries 200 settings and keeps the one with the best validation score. Which number should they report?

Q2

In the same-months control, validation and test came from the same months. What happened to the winner's test score as K grew?

Q3

The gap at K = 1, with no choosing at all, was 0.013. What does that part of the gap come from?

Q4

The best of all 200 settings on the test itself scored 0.750. Persistence scored 0.848. What does that say?

The selection. Each of the 200 was trained on the training rows and scored on validation and on test. Then the lab played the tuning game 1,000 times for each K of 1, 5, 20, 50 and 200: pick K settings at random, keep the one with the best validation score, and write down its validation score and its test score. The gap is the winner's validation score minus its test score. One run per setting; no significance test was declared.

A review of this lesson found a flaw in that control, and it was right. The test months spread the settings far more than the validation months did: the standard deviation of the 200 test scores was 2.45 points, against 0.62 on validation. Pooling those months into both halves gave the control a large, real difference between settings to find, which makes choosing look safer than it is. So the first control understates luck.

A second follow-up, designed after the first and prompted by that review, measures luck fairly. It uses the validation months only: their days are split at random into two halves, the setting is chosen on one half and checked on the other, 20 times over. Both halves are the same season, and the settings differ there only as much as they really do on validation. The halves agreed only loosely (rank correlation 0.47 on average, 0.23 to 0.68). The gap grew from -0.0055 at K = 1 to +0.0031 at K = 200: a growth of 0.86 points from luck alone.

So here is the decomposition. Of the 5.4-point gap at K = 200, 1.3 points are there with no choosing at all: the test months are harder. The other 4.1 points appear as K grows. Ordinary selection luck, measured within the validation months, accounts for about 0.9 points of them, roughly a fifth. Most of the growth came from choosing on months that were not like the months the model then served: the settings that suited the validation months were, more often than not, the wrong settings for the test months.

Two limits. Neighbouring days are alike, so halves made of mixed days are more alike than two truly independent samples, and both controls may understate luck. And the first control mixed the test months' real differences in with luck, which is why its growth came out at -0.2 points and the fair one at 0.9. For any one split of the second control the gap moves too; only the average is the measurement.

This is a real run in VS Code's terminal (python select_demo.py).

A real screenshot of VS Code's terminal after running python select_demo.py. It prints 20 lines, one per setting, with its validation and test accuracy, from setting 0, validation 0.7110, test 0.6989, to setting 19, validation 0.6968, test 0.7037. Then: winner on validation: setting 0; its validation 0.7110, its test 0.6989; median test of all 20: 0.6848; persistence on the test: 0.8484.

When I ran it, all 20 settings matched the lab's first 20 stored scores in select.json to every printed decimal, and so did the winner, the median and persistence; the longest printed line was 42 characters. The report's demo mode checks all of it. In this one draw of 20, the winner came 8th of 20 on test, above the median of these 20, so this particular draw was kinder than the average draw of 20 in the lab. That is the point of the 1,000 draws: one draw can go either way.

Four brand cards headed the tools, with their logos, titled what ran where. scikit-learn: the 200 trees, the scores, the data download. SciPy: the rank correlations. NumPy: the random settings and the 1,000 draws. Python: the report, the demo, the box.

The winner. max(range(20), key=...) picks the setting with the highest validation score. The script then prints that winner's two scores, the median test score of all 20, and persistence: y[b - 1:-1] is the labels shifted by one row, so each test row gets the label of the half-hour before it.

In the lab file, select_lab.py does the same for 200 settings, then plays the tuning game 1,000 times for each K and stores the results in results/select.json.