Ml Lifecycle

Same Code, Different Model: What Changes Between Two Identical Training Runs

0 of 22 complete

0%

Contents

Back|Ml LifecycleSame Code, Different Model: What Changes Between Two Identical Training Runs
1/22
67 min left
Prerequisites
Start With the Dumbest Model: The Baselines a Model Must Beat, on the Futurerequired
Related Topics
The Score That Lied: A Random Split Against a Split in TimeWhy Production BreaksTomorrow Is Different: A Model Frozen on 2011, Scored Through 2012Why Production BreaksTrain/Serve Skew: One Input Computed Two WaysWhy Production BreaksFailing Silently: Which Checks Catch a Broken Input Before the Answers ArriveWhy Production BreaksWatching Inputs Before the Answers Arrive: Drift Measures Against the Real ErrorWhy Production Breaks
1 of 22

The Same Shopping List, Twice

Imagine you give a friend a shopping list and a recipe, and ask them to cook the same dinner on Monday and again on Tuesday. Same shop, same list, same recipe. You would expect the same dinner. But the shop has a rule you never read: at the door, one bag in every ten is taken aside at random for a quality check, and it never comes home. On Monday the missing bag held the onions. On Tuesday it held half the rice. Your friend followed the recipe exactly both times, and the two dinners taste different.

An illustration of a young woman standing, holding a shopping bag on one shoulder, beside the text: the same code, the same data, the same default settings. Only the seed changed. 20 seeds gave 20 different models: accuracy from 0.7179 to 0.7906 on the future. That spread, 7.3 points, is most of the 9.6-point gap between these trees and the lazy rule of lesson 1. Last: a default setting quietly set a random tenth of the rows aside.

Nobody in that story made a mistake. Nobody changed the list. The difference came from a rule that was switched on by default, and that picked something at random every time. If Tuesday's dinner is worse, you cannot find out why by reading the recipe, because the recipe did not change.

This lesson is about the same thing in a computer program that learns from examples. I trained one program 20 times with the same code on the same data and got 20 different programs out. Their scores on the future ran from 0.7179 to 0.7906. The reason was one setting I never wrote, left at its default value, which set aside a random tenth of the examples every time. The rest of the lesson finds that setting, measures what it does, and shows what to write down so that a run can be repeated exactly, and why, even then, one run is not enough to say that one program is better than another.

Where This Lesson Starts

The lesson before this one, start with the dumbest model, trained boosted trees on electricity prices and scored them against simple rules that learn nothing. The trees scored 0.7527 on the future. A rule that just repeats the previous half-hour's answer scored 0.8484. Every model in that lesson ran once with seed 0, except a follow-up that ran the trees with the previous label as an extra input (trees plus lag) with 20 seeds. That lesson ended with a question: if I train the same model twice, on the same data, with the same code, do I get the same model?

The pipelines lesson, training pipelines and orchestration, has a slide called "Reproducibility: Freeze All Three Inputs". It explains, in words, that a run can only be replayed if its data, its code and its environment were frozen and recorded. I will not repeat it here. This lesson measures it, and adds a fourth input that slide does not name: the random choices a training run makes, including the ones hidden inside default settings that nobody chose.

In the life of a model, this is still the training step. It matters for every step after it. If a shipped model misbehaves, you need to rebuild it to debug it. If an auditor asks how a decision was made, you need to show the exact model. If a new model is worse, you need to go back to the old one. And if you want to know whether a change helped, you need to compare two runs that differ only in that change. All four need a run that can be repeated.

Eight Words for This Lesson

A hand-drawn list headed eight words for this lesson, titled repeating a training run exactly. Run: one training of a model: code, data and settings in, a model out. Reproducible: run it again and get the same model: the same guess on every row. Seed: a starting number for the random choices; same seed, same choices. random_state: scikit-learn's name for the seed of one model or split. Default: the value a setting takes when you do not write it. Early stopping: stop adding trees when a held-out slice stops improving. Validation slice: training rows held out to check the model while it trains. Pin: fix a value and write it down, so the next run uses the same one. Beneath: a run you cannot repeat is a run you cannot debug.

A run is one training of a model: code, data and settings go in, a trained model comes out. A run is reproducible if you can run it again and get the same model, meaning the same guess on every row, not just a similar score.

Many training methods make random choices on purpose: which rows to look at, in which order, which to hold back. A computer makes these choices with a random number generator, which is not truly random. It produces a fixed sequence of numbers from a starting number called the seed. The same seed gives the same sequence, and so the same choices. In scikit-learn, the Python library used in this chapter, the seed of a model or a split is a setting called random_state.

A default is the value a setting takes when you do not write it. Early stopping is a rule for boosted trees: keep adding trees only while a held-out slice of the training rows keeps getting better, and stop when it stops improving. That held-out slice is the validation slice. To pin something is to fix its value and write it down, so the next run uses the same value.

Why a Run Must Repeat

Before the numbers, here is why a team should care, in the order the lifecycle meets it.

Debugging. A model in production starts making strange guesses. The first thing anyone does is rebuild it and look inside. If the rebuild gives a different model, you are no longer looking at the model your users met. You are looking at a different one that was trained the same way.

Audits. Some models make decisions that people can challenge: a loan, a price, an alert. When someone asks why, you need to show the exact model that decided, and show that anyone following your record would get the same one.

Rollback. A new model goes live and does worse, so you go back to the old one. If the old one's file was kept, that is easy. If only its code was kept, you must train it again, and you will only get the old model back if the run repeats.

Comparing models. This is the one people forget. A team changes one thing, say a setting or a new input column, trains again, and sees the score go up. Did the change help? Only if the two runs differed in that change and nothing else. If the training itself moves the score by several points from run to run, a single before-and-after pair tells you nothing. This lesson measures how much it moves here.

What the Lab Ran

I wrote the design at the top of the lab file, repro_lab.py, before I ran it. The dated record is the chapter plan, CHAPTER-PLAN.md (2026-09-29, "REPRO LAB (batch 2) designed before running").

An editorial page headed what the lab ran: repro_lab.py, designed before it ran, titled one dataset, one split, one model, five ways to vary it. Four labelled zones. The data: Elec2, as in lesson 1: is the NSW electricity price UP or DOWN against its last 24 hours? Train on the first 36,249 half-hours, test on the last 9,063. The model: boosted trees (HistGradientBoostingClassifier), every setting left at its default except random_state. Five sources, 20 runs each: seed 0 to 19; row order shuffled 20 ways; seed with early stopping off; the same run 20 times; logistic regression with seeds 0 to 19. What was stored: each run's accuracy on the future, how many different models came out, the largest share of test rows two runs disagreed on, and each run's number of trees. Beneath: everything past these five sources, in this lesson, was measured after I saw the results.

The data is Elec2 again, as in lesson 1: 45,312 half-hours of the New South Wales electricity market from 1996 to 1998, in time order. Each row asks: is the price UP or DOWN compared with its average over the last 24 hours? The split is forward in time: train on the first 36,249 rows, test on the last 9,063, which are the future from the model's point of view. Accuracy is the share of test rows where the guess was right.

The model is scikit-learn's HistGradientBoostingClassifier, which I call boosted trees. It builds up to 100 small decision trees, one after another, each one correcting the mistakes of the ones before. Every setting was left at its default except random_state.

Five sources of variation, each run 20 times. seed: random_state from 0 to 19, rows in their original order. row_order: random_state fixed at 0, training rows shuffled 20 different ways. : seeds 0 to 19 again, with early stopping switched off. : seed 0, rows in order, the same run 20 times; it should never change. : a simpler model, , which adds up a weight for each input and turns the total into a chance of UP; it ran with seeds 0 to 19, and its default method has no random steps, so it was expected not to change.

Twenty Seeds, Twenty Models

Here is the headline, straight from the lab's stored file, results/repro.json.

A bar chart headed boosted trees, default settings, 20 seeds: accuracy on the future, titled the same code gave 0.7179 to 0.7906. Twenty bars, one per seed from 0 to 19, on a scale from 0.6 to 0.9, all between about 0.72 and 0.79. Two dashed lines cross the chart: persistence, 0.8484, above every bar; and early stopping off, 0.7520, through the middle of the bars. Beneath: lowest seed 4 at 0.7179, highest seed 2 at 0.7906; median 0.7522. Seed 0 is lesson 1's model, 0.7527.

With the seed as the only change, the 20 runs scored from 0.7179 (seed 4) to 0.7906 (seed 2). That is a spread of 0.0727, or 7.3 percentage points, which I will just call points. The median, the middle value when the 20 scores are put in order, was 0.7522, and the mean 0.7530. Every one of the 20 runs gave a different model. Seed 0 is lesson 1's model, 0.7527, so lesson 1 happened to land in the middle.

The scores are only part of it. Two of these 20 models disagreed on up to 19.7% of the test rows: for about one test row in five, the answer depended on which seed was used, a number nobody chose for any reason.

A two-column table headed the lab's five sources, 20 runs each, from repro.json, titled two sources moved the model; three did not. Left, what changed between runs; right, accuracy; different models. The seed, 0 to 19: 0.7179 to 0.7906; 20. The order of the training rows: 0.7268 to 0.7655; 20. The seed, early stopping off: 0.7520 every run; 1. Nothing, the same run 20 times: 0.7527 every run; 1. Logistic regression, seeds 0 to 19: 0.6491 every run; 1. Beneath: largest share of test rows two seeds disagree on: 19.7%. The repeat and logistic runs never differed.

The other four sources tell you where to look. Shuffling the training rows with the seed held at 0 also gave 20 different models, from 0.7268 to 0.7655. Running the same seed on the same rows 20 times gave one model, 0.7527 every time, so the computer itself is not adding noise here. Logistic regression gave one model, 0.6491, whatever its seed. And the boosted trees with early stopping switched off gave one model, 0.7520, for every one of the 20 seeds. So the seed only mattered while early stopping was on.

How Big Is 7.3 Points?

A spread only means something next to the decisions it could change. Lesson 1 compared the trees with simple rules, and its follow-up also ran the trees plus lag with 20 seeds. Here are both sets of 20 seeds, next to logistic regression and persistence, on one chart.

A dot chart headed the seed spread next to lesson 1's gaps, forward accuracy, titled two models, 20 seeds each, against persistence. Four positions on a scale from 0.6 to 0.9, the axis labelled rule or model; both tree models: 20 seeds each. Logistic, one dot near 0.65; trees, a column of 20 dots from about 0.72 to 0.79; trees + lag, a column of 20 dots from about 0.79 to 0.84; persistence, one dot near 0.85. Beneath: trees 0.7179 to 0.7906; trees + lag 0.7884 to 0.8406 (lesson 1's follow-up). The ranges overlap, but lag beats plain in 396 of 400 pairs, by 5.6 points on average. No seed of either beats persistence (0.8484).

The gap in lesson 1 between the trees (0.7527) and persistence, the rule that repeats the previous half-hour (0.8484), was 9.6 points. The spread from the seed alone is 7.3 points, three quarters of that gap. Even so, every one of the 20 seeds stayed below persistence and above logistic regression, so this ranking does not depend on the seed.

The trees plus lag show what a closer call looks like. With seed 0, lesson 1 measured a gain of 6.7 points over the plain trees (0.8193 against 0.7527). Its follow-up ran the trees plus lag with 20 seeds: they scored from 0.7884 to 0.8406, with a mean of 0.8091, and none of the 20 beat persistence. The two ranges overlap: the best plain seed, 0.7906, is above the worst lag seed, 0.7884. So a single run of each could, with unlucky seeds, have put the plain trees ahead. Paired every way, 20 lag seeds against 20 plain seeds, the lag model won 396 of the 400 pairs, and its mean gain was 5.6 points, not the 6.7 that seed 0 showed. The gain is real on this data; its size, measured once, was a little flattering.

Where the Seed Gets In

Why does the seed matter only while early stopping is on? I read the scikit-learn source code installed on my machine (version 1.9.1) to find out, then measured what I read. Everything on this slide was added after I saw the results.

The documentation for early_stopping says: "If 'auto', early stopping is enabled if the sample size is larger than 10000 or if X_val and y_val are passed to fit." The default is 'auto'. In the code, the line is self.do_early_stopping_ = n_samples > 10_000. Our training set has 36,249 rows, so early stopping was on, and nobody asked for it.

The documentation for validation_fraction, default 0.1, says: "Proportion (or absolute size) of training data to set aside as validation data for early stopping." And for random_state: "Pseudo-random number generator to control the subsampling in the binning process, and the train/validation data split if early stopping is enabled." In the code, fit draws a number from random_state and passes it to scikit-learn's own train_test_split, with stratify=y. Stratified means the slice keeps the same share of UP and DOWN as the training rows. Two other uses of the seed did not apply here. The subsampling for binning only starts above 200,000 rows (subsample=int(2e5) in the binning code). And the same seed also drives feature subsampling, choosing a random subset of the inputs at each split, but only when the setting is below 1; its default is 1, so that was off too.

Row Order, and Why It Only Mattered Through the Slice

The row_order source kept the seed at 0 and shuffled the training rows instead. It also gave 20 different models. Is the order of the rows a second, separate source of randomness?

Two panels headed row order, with and without early stopping, titled order mattered only through the slice. Shuffled rows, default settings: 20 models, 0.7268 to 0.7655, seed fixed at 0. Shuffled rows, early stopping off: 1 model, 20 of 20 identical to the unshuffled one. Beneath: the second panel was run after the results. The same seed on shuffled rows picks different rows for the slice; in 5 of 5 runs checked, that slice explains the model exactly.

Here it was not. The split that makes the slice picks positions: "the rows at these places in the list". Shuffle the list and the same seed picks different rows. To test that, the report trained the same 20 shuffles with early stopping off, which the lab had not done. All 20 made exactly the same guesses as the unshuffled model. Then, for the first five shuffles, it rebuilt the slice from the shuffled list and trained a no-early-stopping model on the rows left behind: all five matched their default runs exactly. So, for these trees, the order of the rows mattered only because it changed which rows the slice took. That is measured on these trees; other methods, such as ones that learn from rows in small batches, can depend on the order directly.

A table of four rows headed what the 20 validation slices took out of training, after the results, titled the slices barely overlap. One slice: 3,625 rows, a tenth of 36,249: 111 to 181 from each full month, Jun 1996 to May 1998. Two slices: share 318 to 405 rows, mean 363: about a tenth of a slice, as pure chance predicts. 20 slices: 31,888 rows held out at least once; 4,361 never. The labels: UP share of every slice 0.4179, the same as training: the split is stratified, it keeps the mix of UP and DOWN. Beneath: each seed trains on a different 90% of the same rows.

How different are the slices? Each takes rows from all through the training period, and two slices share only about 363 rows, a tenth of each slice, which is what two random draws would share. Across 20 seeds most training rows were left out at least once. Because the split is stratified, every slice had the same share of UP rows as training, so the seeds did not differ in how many UP rows they saw, only in which ones.

Why would a different tenth of the rows change the model so much? What follows is general reasoning about how boosted trees work, not something this lab measured. Each tree asks yes-or-no questions such as "is the NSW price above a certain value?", and it picks those values, the cut points, from the training rows it is given. Take away a different tenth and some cut points move a little. The next tree is built to fix the mistakes of the trees before it, so a small change in the first tree changes what the second one learns, and so on for 100 trees. The final models can then disagree on the rows that sit near a cut point, and the next slide shows where they did.

Where the Twenty Models Disagree

The lab stored accuracies, not guesses. To see which test rows the seeds disagree on, the report trained the 20 seed models again and stopped unless every accuracy matched repro.json to twelve decimal places. They all matched. Everything on this slide was added after the results.

A bar chart headed the 9,063 test half-hours, by how many of the 20 seed models said UP, titled the 20 models agree on 6,508 rows and split on 2,555. Twenty-one bars from 0 to 20 on a scale up to 4,000 rows: a tall bar at 0 and a tall bar at 20, with low bars in between. Beneath: all 20 said DOWN on 3,004 rows and UP on 3,504; 7 to 13 said UP on 362. Where all 20 agree, they are right on 5,361 of 6,508.

All 20 models gave the same answer on 6,508 of the 9,063 test rows, and they were right on 5,361 of those. All 20 said DOWN on 3,004 rows and UP on 3,504. On the other 2,555 rows, at least one model said something different. On only 362 of them were the models close to evenly divided, with 7 to 13 of the 20 saying UP; on the rest, 6 models or fewer went against the others.

Two panels headed test rows where the 20 seed models split, measured after the results, titled they disagree most where the label changes. Flips (1,374 rows): 44.1%, label differs from the row before. Steady rows (7,689): 25.3%, label same as the row before. Beneath: a flip is where lesson 1's persistence rule is always wrong, so it is where a model has to do better than persistence.

A flip is a row whose label is not the same as the row before it. Lesson 1 found that persistence is wrong on exactly those rows, so the flips are the only rows where a model can gain on persistence. The seeds split on 44.1% of the flips and on 25.3% of the steady rows. The rows that matter most for beating persistence are the rows where the seed matters most.

A bar chart headed share of test rows where the 20 seed models split, by time of day, titled night rows split most. Eight bars for blocks of 3 hours starting at 00:00, 03:00, 06:00, 09:00, 12:00, 15:00, 18:00 and 21:00, on a scale up to 0.5. Beneath: 00:00 44%, 03:00 25%, 06:00 31%, 09:00 22%, 12:00 23%, 15:00 23%, 18:00 21%, 21:00 38%. Each block holds about 1,128 rows.

By time of day, the seeds split most on the rows from midnight to 03:00 (44% of rows) and from 21:00 to midnight (38%), and least from 09:00 to 21:00, between 21% and 23%. I do not know why. I checked the two obvious reasons in the report and neither fits. It is not the flips: the block with the most flips, 06:00 to 09:00 (25.2% of its rows), splits less than midnight to 03:00, which has 19.2%. And it is not how much the price moves: 03:00 to 06:00 has the smallest average change from one half-hour to the next, yet only 25% of its rows split.

July Again

Lesson 1 found that the trees broke down in July 1998: seed 0 said UP on about nine rows in ten when fewer than half were UP. Here is the same test, month by month, for all 20 seeds.

A two-column table headed forward test by month, 20 seed models; dates rebuilt from row numbers, titled July moved most: from 0.533 to 0.729. Left, headed month (rows): seeds split, the month with its rows and the share of rows where the seeds split; right, headed accuracy range; said UP, the lowest to highest accuracy and the lowest to highest share of rows called UP. Jun 1998 (1,431): 17%; 0.703 to 0.741; 0.39 to 0.51. Jul 1998 (1,488): 39%; 0.533 to 0.729; 0.57 to 0.92. Aug 1998 (1,488): 33%; 0.707 to 0.851; 0.60 to 0.79. Sep 1998 (1,440): 23%; 0.765 to 0.828; 0.47 to 0.65. Oct 1998 (1,488): 27%; 0.762 to 0.841; 0.31 to 0.55. Nov 1998 (1,440): 28%; 0.713 to 0.772; 0.17 to 0.40. Dec 1998 (288): 42%; 0.701 to 0.889; 0.38 to 0.76. Beneath: December holds only 6 days. In July the real UP share was 0.450.

July is where the seed mattered most. Across the 20 seeds, July's accuracy ran from 0.533, only a little better than guessing at random, to 0.729. Seed 0, lesson 1's model, scored 0.553 in July, the 7th lowest of the 20, so another seed would have told a milder story about that month. But the failure itself is not the seed: every seed called UP on more July rows than were UP, 0.571 of rows or more against 0.450. December varied a lot too, but it holds only six days.

An isometric drawing of seven blocks in a row, one per test month of 1998, heights to scale, headed the gap between the best and worst seed in each test month, titled the seed mattered most in July and December. Jun 3.8, Jul 19.6, Aug 14.4, Sep 6.3, Oct 7.9, Nov 5.9, Dec 18.8. Beneath: points of accuracy between the best and the worst of the 20 seeds, month by month, 1998. December holds only 6 days; July is the month lesson 1's trees broke down in.

What lesson 1 measured about July: its average price was well above the training average, 0.0726 against 0.0556. Those numbers are on the dataset's own scale: Elec2 stores every price as a number between 0 and 1, not in dollars. One possible reason for the seed mattering most there, and it is a guess: when prices are unusual, the trees are guessing beyond what they learned, and small differences in the training rows push those guesses in different directions. In June, the first month of the test and the closest to training, the best and worst seed were only 3.8 points apart.

A table of three rows headed the best seed against the worst, row by row, measured after the results, titled seed 2 and seed 4 disagree on 1,553 rows. Where they differ: seed 2 right on 1,106, seed 4 right on 447. Net rows for the better seed, by month: Jun -32, Jul +154, Aug +214, Sep +89, Oct +112, Nov +79, Dec +43. Accuracy: 0.7906 against 0.7179. Beneath: the better seed lost June and won every later month. The gap is not one bad stretch.

Why One Run Is Not Enough

Now the lifecycle question. A team trains a model twice, changes one thing in between, and compares the two scores. How often would the seed alone fool them?

A bar chart headed all 190 pairs of the 20 seeds: the accuracy gap between the two, titled 119 of 190 pairs differ by more than 1 point. Four bars on a scale up to 100 pairs: 1 or less, 1 to 3, 3 to 5, over 5. Beneath: 1 or less: 71; 1 to 3: 74; 3 to 5: 38; over 5: 7. Median gap 1.4 points. Every pair: the same code, two seeds.

There are 190 ways to pick two of the 20 seeds, and every pair is the same code on the same data. The median gap between the two scores was 1.4 points, and 7 pairs were more than 5 points apart. Now suppose a team sees a 2-point gain after a change, measured with one run on each side. Here, 70 of the 190 pairs of seeds were more than 2 points apart with nothing changed at all. A 2-point gain from one run each is well inside what the seed alone produced.

Here is the same problem as a real decision: should early stopping be on or off for this model?

Two panels headed is early stopping better here? one run each, then twenty, titled one run could have said yes or no. Seeds where early stopping on scored higher: 10 of 20, than early stopping off, 0.7520. Seeds where it scored lower: 10 of 20, than early stopping off, 0.7520. Beneath: mean with it on: 0.7530. With seed 2 you would say keep it by 3.9 points; with seed 4, turn it off by 3.4. It never stopped early: this is 90% of the rows against 100%.

With early stopping off, every seed gives 0.7520. With it on, the seed decides. If you had run seed 2, early stopping would look 3.9 points better and you would keep it. If you had run seed 4, it would look 3.4 points worse and you would turn it off. Over all 20 seeds it came out higher 10 times and lower 10 times, with a mean of 0.7530 against 0.7520. Be clear about what this compares: early stopping never stopped here, so the comparison is really training on 90% of the rows against training on 100%, not the benefit of stopping early. On this data that choice made no clear difference on average and made the result depend on luck. One run each could have told you either story with confidence.

A two-column table headed a later lesson in this chapter, gate.json: pairs of seeds on the whole test, titled two seeds are two different models. Left, what the gate lab compared; right, challenger minus champion. 20 pairs of seeds, same code and data: 20 different gaps. The lowest gap: -3.2 points. The highest gap: +7.2 points. Beneath: a retrain with a new seed is a new model. A gate must not mistake that for progress.

What to Pin, and How to Log It

The pipelines lesson named three inputs to freeze: data, code and environment. This lab adds detail to each, and a fourth, the seed.

The code. Commit the training code before the run and record the commit. A git commit is a saved version of the code with a short name, such as 45ab474c for the commit that last changed repro_lab.py. Without it, nobody can tell whether two runs used the same code.

The seed. Set random_state on every model and every split, and write it down. Here, with the code, data, order and versions unchanged, seed 0 gave one model in 20 runs.

The data snapshot and its order. Save the exact rows, and a fingerprint: a short code computed from the bytes of the data (here a SHA-256 hash), which changes if a single value changes. Keep the order of the rows, or sort by a fixed key before training. Here the order changed the model through the slice.

Every setting, defaults included. The setting that moved this model was one I never wrote. Log the full list of settings the model actually used; in scikit-learn, model.get_params() returns them all.

The library versions. A default can change between versions of a library, and so can the way a split draws its rows. Write down the version of Python and of every library. The report does it with sklearn.__version__ and the like.

A table of five rows headed the run receipt this lab's report writes, from the installed libraries, titled what to write down so a run can be repeated. Code: scripts/labs/lifecycle/repro_lab.py at git commit 45ab474c. Versions: Python 3.13.15, scikit-learn 1.9.1, NumPy 2.5.3, pandas 3.0.6, SciPy 1.18.1, Darwin arm64. Data: OpenML 151 version 1, 45,312 rows, fingerprint 5a3683554b30c1cc. Split and order: first 80% of rows in file order to train, last 20% to test. Settings: random_state 0, early_stopping auto, validation_fraction 0.1, max_iter 100. Beneath: write the defaults down too. They are settings you did not choose.

Try It Yourself

This script is the lab made small. It downloads the same data, trains the boosted trees with seeds 0, 1 and 2, first with the default settings and then with early stopping off, and prints each accuracy and the number of trees built. It does not need a GPU.

A real screenshot of VS Code with repro_demo.py open, showing the docstring that says what the script is and how to run it, the imports, the lines that load Elec2 and turn the label into 1 for UP and 0 for DOWN, the cut at 80%, the version line, and the two loops over early stopping and seeds 0 to 2. Beneath: copy it from the box on the slide.

Before you run this lab. You need Python 3 and two libraries: pip install scikit-learn pandas. scikit-learn holds the model, the score and the download; pandas holds the table. The first run downloads Elec2 from OpenML (under 1 MB compressed), so it needs an internet connection once; scikit-learn keeps a copy in a folder in your home directory (scikit_learn_data) for later runs. I ran it with scikit-learn 1.9.1 on a Mac. scikit-learn runs the same way on Windows and Linux, but I have not checked the numbers there. This lesson is about exactly that: another version of scikit-learn, or another machine, may give different decimals, and the first line the script prints is the version, so you can compare.

"""Same code, different model: the same boosted trees trained with three seeds.

Lesson 2 of 'The ML & AI Lifecycle', made small. It needs Python 3 with scikit-learn
and pandas (pip install scikit-learn pandas). The first run downloads Elec2 from
OpenML (under 1 MB) and keeps a copy for later runs.
    python repro_demo.py

Author: Roni Das
Created: 2026-09-29
"""
import sklearn
from sklearn.datasets import fetch_openml
from sklearn.ensemble import HistGradientBoostingClassifier
from sklearn.metrics import accuracy_score

# 45,312 half-hours in time order. The label: is the NSW price UP or DOWN
# against the average of the last 24 hours?
data = fetch_openml("electricity", version=1, as_frame=True, parser="auto").frame
X = data.drop(columns="class").astype(float)
y = (data["class"].astype(str) == "UP").astype(int).to_numpy()

cut = int(0.8 * len(y))            # train on the first 80% in time
print(f"scikit-learn {sklearn.__version__}, {cut:,} training rows")

for early_stopping in ("auto", False):
    for seed in (0, 1, 2):
        # "auto" is the default: above 10,000 rows it turns early stopping on,
        # and early stopping holds out a random 10% of the training rows.
        model = HistGradientBoostingClassifier(
            random_state=seed, early_stopping=early_stopping)
        model.fit(X[:cut], y[:cut])
        acc = accuracy_score(y[cut:], model.predict(X[cut:]))
        mode = "on " if early_stopping == "auto" else "off"
        print(f"seed {seed}  early stopping {mode}  accuracy {acc:.4f}"
              f"  trees {model.n_iter_}")

The Lab Report

A real terminal recording headed python repro_report.py, titled every table in this lesson, from the stored files and the data. It opens with 128 checks against the data and refits: all agree, then prints six numbered sections: 1, the five sources, 20 runs each, with seed 0.7179 to 0.7906 and seed_no_es 0.7520 for every run; 2, where the seed gets in, with 20 of 20 seeds equal to early_stopping=False on the kept rows; 3, what the validation slices take out, 111 to 181 rows from each full month; 4, where the 20 models disagree, by month, by three-hour block and for the best and worst seed; 5, one run against one run, including the trees plus lag over 20 seeds; 6, the run receipt. Beneath: the lab's own report. It fits the same models again and stops unless every score matches.

The report lives in scripts/labs/lifecycle/repro_report.py. It reads the lab's stored file, results/repro.json, lesson 1's results/baseline.json and results/baseline_followup.json, the gate lab's results/gate.json, and the Elec2 data from scikit-learn's local copy. The lab stored accuracies, not guesses, so the report trains the lab's models again and compares each accuracy with the stored one to twelve decimal places: seeds 0 to 19 with the default settings, where it also checks that each built 100 trees; seeds 0 to 19 with early stopping off; the 20 shuffles; the repeat, trained twice, which must give the same guesses both times; and logistic regression with seeds 0 and 7, which must give the same guesses. It changes nothing in repro.json.

It also runs the checks I added after the results: it rebuilds each seed's validation slice from scikit-learn's own split and trains a no-early-stopping model on the rest (20 of 20 identical), trains the 20 shuffles with early stopping off (20 of 20 identical to the unshuffled model), rebuilds the slice for 5 of the 20 shuffles (5 of 5 identical), and checks that the 10 gate-lab pairs within seeds 0 to 19 use this lab's seeds. Then it prints the tables in this lesson and the run receipt. It makes 128 checks in all, and they all agree.

Its mode writes every number to , which the figures read. The mode checks the student script's stored run, and the mode writes the playground below and checks that it gives the lab's 20 accuracies.

Compare the Seeds Yourself

This box has no model in it. It holds the 9,063 forward test labels and the guesses of all 20 seed models for every row, recomputed by the report and checked against the lab's 20 accuracies. Each guess is a 0 or a 1, so the 20 guesses for one row can be written as a single 20-digit binary number, one digit (a bit) per seed. To keep the box small, that number is written in hexadecimal, counting in sixteens with the digits 0 to 9 and a to f, which takes five characters instead of twenty. The guesses function unpacks it for you. It runs in your browser.

As it is, the box prints each seed's accuracy, from 0.7179 to 0.7906 with a spread of 0.0727, which the report's box mode checks against repro.json, and then compares seed 2 with seed 4: they differ on 1,553 rows, and seed 2 is right on 1,106 of them.

Try by_month(2, 4) to see the best and worst seed month by month, and votes() to see how many of the 20 models said UP on each row. Try compare(0, 1) for two neighbouring seeds. Then pick two seeds and pretend they were two versions of a model, a "before" and an "after". How big a gain would you have believed?

The Code, Part by Part

Loading. fetch_openml("electricity", version=1) downloads Elec2 once and reads the local copy after that. The label column holds the words UP and DOWN, turned into 1 and 0. The eight inputs become plain numbers with astype(float), as in the lab.

The version line. sklearn.__version__ is printed first, because the result depends on it. If yours is not 1.9.1, the accuracies may differ.

The cut. cut = int(0.8 * len(y)) is 36,249. The model learns from X[:cut] and is scored on everything after. Nothing is shuffled.

The two loops. The outer loop sets early_stopping to "auto", the default, and then to False. The inner loop sets random_state to 0, 1 and 2. Each pass builds a fresh model, so nothing carries over between runs.

What "auto" does. With more than 10,000 training rows, "auto" turns early stopping on, and early stopping holds out a random 10% of the training rows, chosen by random_state. With False, all 36,249 rows are used and has nothing to choose here.

How to Make a Training Run Repeatable

A hand-sketched column of eight boxes joined by arrows, headed making a training run repeatable, titled pin these, in this order. 1, commit the code; note the commit. 2, set random_state on every model and split. 3, save the data snapshot and its fingerprint. 4, keep the row order, or sort by a key. 5, log every setting, defaults included. 6, record the library versions. 7, rerun once: same guesses, row for row? 8, to compare two models, add seeds. Beneath: steps 1 to 6 fix what goes in, step 7 proves it, step 8 compares. Pinning makes one run repeatable; it does not make one run enough.

The order matters: first fix everything that goes into the run, then prove the run repeats, and only then use it to compare models.

Commit the code. Train from a committed version of the code, never from unsaved changes, and record the commit.

Set random_state everywhere. On every model and every split, including splits you did not write, such as the one inside early stopping. If a library has a global random generator, set that too.

Save the data, not a query. A query against a live table can return different rows tomorrow. Save the exact rows, or a versioned snapshot, with a fingerprint.

Keep the order. Sort by a fixed key, such as time and then an id, before training. A database or a file system does not promise the same order twice.

Log every setting. Log get_params(), not just the settings you typed. Read the documentation of every default that uses a seed, a sample or a split.

Record the versions. Python, every library, and the machine type. A lockfile, a file listing the exact version of every library, does this for a whole project.

Rerun once, and compare guesses. Train again from the receipt and check that every guess matches. A matching score is not enough.

Then, to compare two models, add seeds. Run the change with several seeds, for example five or ten, and look at the spread before you believe a gain.

Pin One Seed, or Run Many?

A two-column table headed grounded in this lesson's numbers, titled pin one seed, or run many? Pin everything when you: debug a model that misbehaves; must rebuild a shipped model for an audit or a rollback; compare two changes and need like with like; want the run to repeat: here, 20 of 20 repeats were identical. Run several seeds when you: decide whether a change helps; report how good a model is; could be fooled by luck: seeds here spanned 7.3 points; pick a setting like early stopping on or off. Beneath: pinning answers: can I get this model back? Many seeds answer: is this model really better?

Pin everything when you need a model back. Debugging, audits and rollbacks all need the exact model, and a pinned run gives it: here 20 runs of seed 0 gave one model every time.

Pin everything when you compare two changes on one seed. Same seed, same rows, same order, same versions, so that the change you made is the only difference. That is a fair comparison of two models, for that seed.

Run several seeds when you decide whether a change helps. One pinned seed can still be a lucky or an unlucky one: here seed 2 would have made early stopping look good and seed 4 bad. Report the spread, not one number.

Run several seeds before you report a score to anyone. The score of one seed is one draw. The median and the range of several is the honest summary. Here that is a median of 0.7522 and a range of 0.7179 to 0.7906.

Do not pin a seed and then pick the best one. Trying many seeds and shipping the highest scorer on the test turns the seed into a setting you tuned on the test. A later lesson in this chapter measures how much of such a winner's lead is luck.

Do not expect pinning to survive a change of version or machine. A new library version can change a default or the way a split draws rows. Pinning holds within one recorded environment.

Removing the randomness is not always the answer. Here, turning early stopping off removed the spread and scored about the same as the average. With other data, early stopping can really help. Measure it over several seeds before you decide.

What These Numbers Can and Cannot Tell You

A two-column page headed read before you trust these numbers, titled what these runs are, and what they are not. They are: one dataset: Elec2, 1996 to 1998; one model family, default settings; 20 runs per source; the five sources came first; the slice checks came after; one machine, one set of versions. They are not: not the size of the spread elsewhere; other models vary in other ways; no significance test; not chosen after the results; their reasons are partly guesses; other versions may differ.

One dataset, one model, default settings. The size of the spread, 7.3 points, belongs to these trees on this data and this split. With more training rows, or a test more like the training period, the slice would probably matter less; I did not test that. Neural networks and other methods have other sources of randomness, such as the starting weights and the order of batches, which this lab does not cover.

20 runs per source, no significance test. The design described the runs and declared no test. The correlation of -0.52 between calling UP and accuracy is over 20 points and is a pattern, not a proof.

The design came first; the explanations came after. The five sources and the stored measures were written in repro_lab.py before it ran. Reading the scikit-learn source, rebuilding the slices, the shuffled runs with early stopping off, and every table of where the seeds disagree were added after I saw the results. The slice was rebuilt for all 20 seeds but for only 5 of the 20 shuffles.

What is measured and what is guessed. Measured: the default model makes the same guesses as a no-early-stopping model trained on the rows the slice left, for all 20 seeds; row order mattered only through the slice, for these trees. Guessed: why July split most. Not explained: why night rows split most; the two reasons I checked did not fit.

One machine, one set of versions. Everything ran on one Mac with scikit-learn 1.9.1. I have not checked whether another version or another operating system gives the same models. The lesson's advice is to record that, because I cannot promise it.

The dates are rebuilt, as in lesson 1: months and hours come from row numbers, checked against the weekday column.

What to Do Next

A hand-drawn list headed before you trust a comparison of two models, titled five questions. Repeat?: if I run it again, do I get the same guesses, row for row? Defaults?: which default setting uses randomness I never set? Logged?: are the code, seed, data, order, settings and versions written down? Spread?: how far apart are runs of the same code with other seeds? Bigger?: is the difference I see larger than that spread? Beneath: here: one piece of code, 0.7179 to 0.7906 by seed alone.

Take the last model your team trained and try to rebuild it from what was written down. If you cannot, find what is missing: usually the code version, the data snapshot, the seed of some split, or the library version. Then train it five times with five seeds and write down the spread. Put that number next to the last improvement anyone claimed for the model. If the improvement is smaller than the spread, measure it again with several seeds before you believe it.

It takes an afternoon. The receipt it produces is the start of everything later in this chapter: a pipeline can only promote, roll back and compare models whose runs it can repeat.

The next lesson planned for this chapter takes the other road out of this lesson: if a team tries many settings or many seeds and keeps the best, how much of the winner's lead is luck, and how much of it survives on fresh data?

A closing card headed to keep, titled same code is not the same model. In large type: 0.7179 to 0.7906. Beneath: boosted trees on the future, 20 seeds, every other setting the default. With early stopping off: 0.7520 every time. Then: pin the code, the seed, the data, the order, the settings and the versions to repeat a run. Run several seeds to compare two models. Last: one dataset, 20 runs per source: a way to check, not a law.

Knowledge Check

Knowledge Check

4 questions - Score 80% to pass

Q1

The boosted trees scored 0.7179 to 0.7906 over 20 seeds, but 0.7520 for every seed with early stopping off. What made the seeds differ?

Q2

A teammate changes one setting, trains once before and once after, and sees accuracy rise by 2 points. What does this lesson suggest?

Q3

Which record lets someone rebuild the exact model later?

Q4

The 20 seed models agreed on 6,508 of 9,063 test rows. Among these four, where did they disagree most?

seed_no_es
repeat
logistic
logistic regression

A sequence diagram with four columns: the lab, the model, the future, the scorer. Headed one run of the lab, repeated 20 times per source, titled only one thing changes from run to run. Step 1, the lab makes a new model with seed k. Step 2, the lab tells the model to learn from the first 80%. Step 3, the lab takes the last 20%. Step 4, the future sends the model its inputs only. Step 5, the model sends its guesses to the scorer. Step 6, the future sends the real labels to the scorer. Step 7, the scorer computes accuracy. Beneath: k runs from 0 to 19. The repeat and row_order runs keep seed 0; row_order shuffles the training rows instead. The model never sees the test labels; only the scorer does.

For every source the lab stored each run's accuracy, the number of different models that came out (two models count as the same if they make the same guess on every test row), the largest share of test rows on which any two runs disagreed, and how many trees each run built. The design declared no significance test, a calculation of how likely a difference this size would be by chance alone: 20 runs per source are described, not tested.

max_features

A flowchart headed inside fit(), scikit-learn 1.9.1, read from its source, titled where random_state reaches the trees. fit: 36,249 training rows leads to a question, early_stopping auto: more than 10,000 rows? If yes: random_state picks a random 10%, stratified: same UP share, then train on the other 90%. If early_stopping=False: train on all 100%. Both lead to 100 trees, then stop. Beneath: measured after the results: in all 20 seed runs the model makes the same guess on every test row as a no-early-stopping model trained on that 90%. The slice decides the model.

So each seed held out 3,625 of the 36,249 training rows, a tenth, and trained on the other 32,624. Then I checked the surprising part. The lab stored each run's number of trees, and every one of the 20 runs built all 100. Early stopping never stopped. The report also read each run's score on the held-out slice after every tree: in all 20 runs the best score came at the very last tree, so the model was still improving when it hit the default limit of 100 trees.

A hand-drawn sketch headed sketched: the first 40 training rows, real slices, titled each seed sets aside different rows. Three rows of 40 narrow boxes, labelled seed 0, seed 1 and seed 2, with four or five boxes in each row marked x, in different places. Beneath the rows: x = held out for early stopping. Beneath the sketch: rows 0 to 39, counting from 0. Held out: seed 0, rows 11, 17, 18, 23, 36; seed 1, rows 7, 10, 19, 22; seed 2, rows 12, 19, 21, 28. Each seed sets aside 3,625 of 36,249 rows.

That makes the test simple. If early stopping changed nothing about when training stopped, then the default model should make the same guesses as a model with early stopping off, trained on the 90% of rows that the slice left behind. The report rebuilt each seed's slice from scikit-learn's own split with the same seed and trained that model. For all 20 seeds the two made the same guess on every one of the 9,063 test rows.

Two panels headed added after the results: rebuilt from scikit-learn's own split, titled early stopping never stopped, but it still took its slice. Runs that used all 100 trees: 20 of 20, best score on the held-out slice came at the last tree in all 20. Matches the 90% model: 20 of 20, early stopping off, trained on the kept 90%: the same guess on every test row. Beneath: each run trained on 32,624 rows, not 36,249. Early stopping changed nothing about when it stopped, only which rows it learned from.

So the whole 7.3-point spread comes from one thing: which tenth of the training rows each seed left out. That is measured. The seed did not make the trees random in any other way here.

The best seed, 2, and the worst, 4, disagreed on 1,553 rows. Seed 2 was right on 1,106 of them and seed 4 on 447: a net 659 rows, which is exactly the 7.3 points between them. Seed 2 lost June by 32 rows and won every month after it. So the better seed was not lucky in one week; it was better nearly all the way through this test.

A dot chart headed each seed model: share of test rows it called UP, against its accuracy, titled seeds that said UP more often tended to score lower. Twenty dots on a scale of share called UP from 0.40 to 0.62 and accuracy from 0.70 to 0.80, with a dotted vertical line at the real UP share 0.451. Every dot is to the right of the line. Two of the three most accurate seeds sit at the left of the group, near 0.47 and 0.49, and the three that called UP most, near 0.55, 0.58 and 0.60, are among the lower ones; the middle cluster, between 0.49 and 0.52, shows no clear trend of its own. Beneath: 20 seeds, one dot each. Called UP on 0.469 to 0.601 of rows; correlation with accuracy -0.52. Measured after the results; 20 points, a pattern, not a proof.

One more measured pattern: every seed called UP on more rows than were UP (0.469 to 0.601 of rows, against 0.451), and the seeds that called UP more often tended to score lower, with a correlation of -0.52 over the 20 seeds. A correlation runs from -1 to 1; -0.52 means a clear but loose downward trend. With 20 points it is a pattern, not a proof.

A later lesson in this chapter looks at the gate a pipeline uses to decide whether a newly trained model replaces the live one. Its lab paired seeds of these same trees, 20 pairs, and on the whole test the gaps between two seeds ran from -3.2 to +7.2 points. So a retrain with a new seed is a genuinely different model, sometimes better, sometimes worse, with nothing else changed. That lesson measures how often a gate is fooled by it.

How many seeds are enough? There is no fixed number, and I did not measure one. What this lab gives is a way to find out on your own data: run the same code with five or ten seeds, look at how far apart the scores are, and treat any difference smaller than that as unproven until it holds across seeds. Report the median and the range, not the best run. It costs five or ten times the training time, which for a model that trains in seconds, like these trees, is nothing, and for a model that trains for days is a real decision.

There are two ways out, and they answer different questions. Pin everything and compare like with like: use the same code apart from the change, the same seed, the same data, the same order and the same settings for both runs, so the only difference is the change you made. That makes the comparison fair for that one seed. Run several seeds on each side and compare the spreads: if the change is real, it should hold across seeds, as the trees plus lag did in 396 of 400 pairs. Pinning tells you whether you can get a model back. Many seeds tell you whether a model is really better.

Put all of it in one record per run, a run receipt, stored next to the model file, with the score the run produced so that a rebuild can be checked against it. A small dictionary written to a JSON file is enough to start; an experiment tracker such as MLflow does the same job for many runs and many people, as the pipelines lesson explains. The report's receipt function builds this one. The receipt is what lets someone else, or you in six months, get the same model back.

Then check it. Run the training twice with the receipt's values and compare the guesses row for row, not just the score. Two different models can have nearly the same score: here seed 0 with early stopping on scored 0.7527 and the no-early-stopping model scored 0.7520, a gap you could easily miss, yet the two disagree on 235 of the 9,063 test rows.

This is a real run in VS Code's terminal (python repro_demo.py).

A real screenshot of VS Code's terminal after running python repro_demo.py. It prints seven lines: scikit-learn 1.9.1, 36,249 training rows; then seed 0, early stopping on, accuracy 0.7527, trees 100; seed 1, on, 0.7464, trees 100; seed 2, on, 0.7906, trees 100; seed 0, off, 0.7520, trees 100; seed 1, off, 0.7520, trees 100; seed 2, off, 0.7520, trees 100.

When I ran it, all six accuracies matched the lab's stored runs in repro.json to every printed decimal. The three seeds with the default settings give three models; the same three seeds with early stopping off give one. Every run built 100 trees. If your numbers differ, compare the version line first.

json
results/rn-report.json
demo
box

What came before the run, in repro_lab.py: the data, the split, the model, the five sources and the stored measures. What came after I saw the results: reading the scikit-learn source, rebuilding the slices, the no-early-stopping runs on the kept rows and on shuffled rows, where the seeds disagree by month, by time of day and at flips, the best-against-worst table, the pairs of seeds, the early-stopping comparison, the comparison with the trees plus lag and the receipt.

Four brand cards headed the tools, with their logos, titled what ran where. scikit-learn: the trees, their hidden split, the data download. pandas: the table of half-hours. NumPy: the shuffles, the votes, the fingerprint. Python: the lab, the report, the demo, the box.

random_state

The printout. Each line shows the seed, whether early stopping was on, the accuracy on the future and model.n_iter_, the number of trees built. That last number is how you can see that early stopping never stopped: it is 100 every time, the default limit.

In the lab file, repro_lab.py does the same for seeds 0 to 19, adds the shuffled rows, the repeated run and logistic regression, and stores every accuracy in results/repro.json.