Ml Lifecycle

Lineage and Rollback: Rebuilding Last Month's Model Exactly, and What Breaks When You Cannot

0 of 22 complete

0%

Contents

Back|Ml LifecycleLineage and Rollback: Rebuilding Last Month's Model Exactly, and What Breaks When You Cannot
1/22
63 min left
Prerequisites
Stale Pieces in a Pipeline: A Cached Input, an Old Scaler, and What Each One Hidesrequired
Related Topics
The Score That Lied: A Random Split Against a Split in TimeWhy Production BreaksTomorrow Is Different: A Model Frozen on 2011, Scored Through 2012Why Production BreaksTrain/Serve Skew: One Input Computed Two WaysWhy Production BreaksFailing Silently: Which Checks Catch a Broken Input Before the Answers ArriveWhy Production BreaksWatching Inputs Before the Answers Arrive: Drift Measures Against the Real ErrorWhy Production Breaks
1 of 22

The Bread Everyone Liked

Imagine a small bakery. Last month the baker made a new kind of bread, and the customers loved it. This month she tried to improve it, and people started to complain. So she decides to go back: she will bake last month's bread again, exactly as it was, until she understands what went wrong.

An illustration of a woman in a long skirt and cardigan, holding a notebook to her chest, next to text. Headed the bread everyone liked, titled can you bake last month's bread again, exactly? Beside her: one model, trained on 30,000 half-hours with seed 7, rebuilt five ways, as if a month later. With every note kept: the same answer on 1.000 of the 6,000 rows it served. The saved file loaded back: 1.000. With the seed not written down: 0.920 to 0.970 over 19 guesses, never exact, and the data check stayed quiet. Last: write every ingredient down, and keep a copy of the loaf.

She opens her notebook. It says: flour, water, salt, starter, bake at 230 degrees. It does not say which bag of flour she used, and the mill has since changed its wheat. It does not say how long the dough rested, because that day she simply waited until it looked right. She bakes the bread from the notebook, and it is close. It is not the same. And she cannot tell which missing note made the difference.

There were two ways she could have made this easy. She could have written down everything, including the small things that seemed not to matter. Or she could have frozen one loaf from last month, so that she could taste the real thing whenever she wanted.

This lesson is about the same problem in a program that learns from examples. I trained a model, wrote down how I made it, and then tried to get it back in five different ways, as if a month had passed. Some ways gave back the real thing. One gave back something that looked almost the same, and nothing in my notes warned me that it was not.

Where This Lesson Starts

This chapter follows one model through its life, on the same electricity data. Start with the dumbest model set the baselines. Same code, different model showed that two training runs with the same code can give different models. The lessons after it judged new models, put them live, retrained them, and, in stale pieces in a pipeline, looked at old parts left inside a working system.

A flowchart headed the morning something breaks, titled two ways back to last month's model. A box, the new model starts giving bad answers, leads to a box, go back to last month's model. That box splits in two. Left: load the copy that was saved, leading to: the same answers, if the copy was kept. Right: train it again from the notes, leading to: the same answers, only if the notes were complete. Beneath: this lesson measures both roads on the electricity data, and what breaks on the second when a note is missing.

This lesson starts on a bad morning. A new model is live and something is wrong with it. The safe move is to go back to the model that worked before, and then look for the cause calmly. There are two roads back. The first is to load a saved copy of the old model. The second is to train the old model again from the notes that were kept about it.

The first road needs a copy. The second needs complete notes. Most teams think they have one or the other. The lab in this lesson checks both, and then takes notes away one at a time to see what happens.

The course's lesson on training pipelines and orchestration explains in words why a pipeline should record what it did. I will not repeat that. This lesson measures what the record buys you on the day you need it.

Ten Words for This Lesson

A hand-drawn list headed ten words for this lesson, titled what it takes to get a model back. Lineage: where a model came from: which rows, which settings, which seed, which library. Record: the written note of that lineage, saved next to the model. Seed: the number that fixes the random choices inside training. Hash: a short fingerprint of exact numbers; change one number and it changes. Artifact: a saved file a step made: here, the trained model itself. Snapshot: the saved model file, kept so it can be loaded back. Registry: the shelf where model files and their records are kept by version. Rollback: putting an older model back into service. Rebuild: training the older model again from its record. Backfill: correcting old stored data after the fact. Beneath: a rebuild is only as exact as its record.

A model's lineage is where it came from: which rows of data it learned from, which settings, which seed, and which versions of the software. The written note of that lineage, saved next to the model, is its lineage record, or just the record.

Training a model involves random choices, such as which rows to hold back for checking. A computer makes those choices from a fixed list of numbers that starts from one number, the seed. The same seed gives the same choices every time. A different seed gives different choices, and often a slightly different model. Lesson 2 of this chapter measured how different.

A hash is a short fingerprint of some exact numbers. The one used here, SHA-256, turns any amount of data into 64 letters and digits. If even one number in the data changes, the fingerprint changes completely. So if you store the fingerprint of your training data today, you can check later whether the data is still exactly the same, without keeping a second copy of it.

An artifact is a file that a step of a pipeline saved. The trained model is one. A saved copy of the trained model, kept so it can be loaded back later, is a snapshot. A is the place where a team keeps these files and their records, by version. Rollback means putting an older model back into service. Rebuild means training the older model again from its record. Backfill means correcting old data after it was stored, for example when late information shows that some old labels were wrong.

Two words from earlier lessons: a is the right answer for a row, here whether the electricity price went UP or DOWN against its average over the last 24 hours, and is the share of rows a model got right. One new measure: is the share of rows on which two models give the same answer. It needs no right answers at all.

What a Lineage Record Holds

Before the lab could test a rebuild, it needed a model to rebuild and a record of how it was made. I wrote both into the design before the run. Here is the record, with the real values the lab stored.

An editorial page in six labelled zones, headed the lineage record the lab wrote, real values, titled six notes, saved next to the model. Rows: 0 to 30,000, written as numbers, not as "the last 30,000 rows". Data fingerprint: SHA-256 of those rows' inputs and labels: 7f391cd43a8a... (64 characters in all). Settings and seed: the model's default settings, and the seed: 7. Libraries: scikit-learn 1.9.1, numpy 2.5.3. Answers fingerprint: SHA-256 of its 6,000 answers on the month it served: b978599166b1... The saved file: the trained model, pickled (turned into bytes): 369,066 bytes.

Rows. The record says which rows the model learned from, as plain numbers: rows 0 to 30,000 of the data. This seems too obvious to write down, and the lab shows later why it is not.

Data fingerprint. A SHA-256 hash of the inputs and labels of exactly those rows. The row numbers say which rows; the fingerprint says what was in them.

Settings and seed. The model uses scikit-learn's default settings, plus one written value: the seed, 7. The lab's record stored only the seed; the defaults were implied by the library version. A real record should write them out, because a default can change.

Libraries. The versions of scikit-learn and numpy that trained it. A model trained with one version of a library may train differently with another, so the version is part of the lineage.

Answers fingerprint. After training, the model answered the next 6,000 rows, the period it "served" in this story. The record stores a hash of all 6,000 answers. This field is not in every team's records, and it turns out to matter most.

The saved file. The trained model itself, saved with Python's pickle, which turns an object in memory into a file of bytes; a model saved this way is called pickled. Here the file was 369,066 bytes, about a third of a megabyte. One warning that belongs next to every pickle: loading a pickle file can run code hidden inside it, so load model files only from a registry you control and never from an untrusted source. scikit-learn's documentation points to the skops format, which avoids pickle, and to ONNX as safer choices.

What the Lab Ran

I wrote the lab's design at the top of its file, lineage_lab.py, before it ran. The dated entry is in the chapter plan (2026-09-29, "LINEAGE LAB (batch 8) designed before running").

The data is Elec2 again: 45,312 half-hours of the New South Wales electricity market, May 1996 to December 1998, in time order. The model is the boosted trees from lesson 1, many small trees of yes-or-no questions built one after another, with scikit-learn's default settings and seed 7. It learned from rows 0 to 30,000 and then "served" rows 30,000 to 36,000: the lab stored its answers on those rows as the reference that every rebuild is compared with.

A two-column list headed the five rebuilds, fixed before the run, titled one record, five ways to use it a month later. Full record: same rows, same settings, same seed, as written down. Snapshot: no training: the saved model file loaded back. Seed missing: same rows and settings, seed not written down: 19 guesses, 0 to 19 without 7. Data backfilled: the same record, but 1% of the old labels (282) corrected upstream after the model was built. Rows shifted: rows written as "the last 30,000", read a month later: now rows 6,000 to 36,000. Beneath: scored by agreement: the share of the 6,000 served rows where the rebuild gives the original's answer. "The served month" is the lab's name: in the calendar it is 125 days, 22 January to 26 May 1998.

A note on a word before going on. The lab's design calls these 6,000 rows "the month it served", and I keep that name because it is short. But 6,000 half-hours is 125 days in the calendar, from 22 January to 26 May 1998, so the "month" is really about four months. Nothing in the results depends on the length; it is only a name.

The five rebuilds were fixed in the design. Full record trains again with everything the record says. Snapshot trains nothing: it loads the saved file back. Seed missing pretends the seed was never written down, so it tries the other nineteen seeds from 0 to 19, one at a time (the design said 20; excluding the original's seed leaves 19). Data backfilled uses the full record, but the stored history has changed underneath it: 1% of the labels in rows 0 to 30,000 were flipped, as a late correction from upstream, the team or system that supplies the data, would do. Rows shifted pretends the rows were written down as "the last 30,000 rows as of that day" instead of as numbers, and the rebuild happens 6,000 rows later, so the same words now mean rows 6,000 to 36,000.

Five Rebuilds of One Model

Here is what the lab stored in results/lineage.json.

A two-column table headed the main run, lineage.json, titled two ways gave the model back; three did not. Left, rebuild; right, agreement with the original. Full record: 1.0000, the same answers. Snapshot: 1.0000, the same answers. Data backfilled: 0.9583. Rows shifted: 0.7385. Seed missing, 19 guesses: 0.9197 to 0.9702, 0 exact. Beneath, left: served month: rows 30,000 to 36,000. Beneath, right: one run each; 19 for the guesses.

The full record worked. Trained again with the same rows, settings and seed, the model gave the same answer on all 6,000 served rows, and the hash of its answers matched the stored hash exactly. After the results, I also pickled the rebuilt model and compared the pickled bytes with the original's, in memory: they were identical, 369,066 bytes each. On the same machine, with the same library versions, the record was enough to get the exact model back.

The snapshot worked too, as it must: loading the saved file gives back the same object, and it gave the same 6,000 answers.

The other three did not. With the labels corrected upstream, the rebuild agreed on 0.9583 of the rows. With the rows shifted by a month, only 0.7385: the rebuild disagreed with the original on more than a quarter of the rows it was meant to reproduce. And with the seed missing, not one of the 19 guesses gave the original model. They agreed on 0.9197 to 0.9702 of rows, with a middle value of 0.9568.

A chart headed lineage.json: agreement with the original on the served month, titled every guessed seed landed short of an exact rebuild. On a scale from 0.70 to 1.00: full record and snapshot bars reach 1.00; a column of 19 dots for the seed guesses sits between about 0.92 and 0.97; backfilled reaches about 0.96; shifted about 0.74. Beneath: 1.0000, 1.0000, guesses 0.9197 to 0.9702 (median 0.9568), 0.9583, 0.7385. The axis starts at 0.70.

Look at where the seed guesses sit on the chart. They are close to the top, closer than the backfilled rebuild in some cases. A model that agrees with the original on 97% of rows looks like a success. It is not the original model, and on a day when you are trying to understand what the original did, even the closest guess gave a different answer on 179 of the 6,000 rows: 179 rows of wrong evidence.

The Seed Nobody Wrote Down

Which of the three failures would you notice? The lab asked whether the record would warn, and here the seed is different from the other two.

Two panels headed seed not written down: what the record's data check says, titled the quiet failure. Data fingerprint: matches: same rows, same labels: the check has nothing to say. Answers: 0.957: median agreement over 19 guesses; 0 exact. Beneath: every seed gave a model trained the right way on the right rows. It just was not last month's model.

For the backfilled and shifted rebuilds, the data fingerprint fired: the data being fed to the rebuild was not the data in the record, and the record said so. For the seed guesses, the data fingerprint matched. The rows were the right rows. The settings were the right settings. Everything the team could check about the inputs to training was correct, and the model that came out was still a different model. From the inputs side, this failure is completely quiet.

After the results, I laid out which field of the record would fire for each rebuild.

A grid headed which field of the record fires for each rebuild; after the results, titled only the answers fingerprint sees a missing seed. Two columns: data fingerprint and answers fingerprint. Full record: quiet, quiet. Snapshot: quiet, quiet. Data backfilled: fires, fires. Rows shifted: fires, fires. Seed missing: quiet, and fires on 19 of 19. The cells that fire are outlined. Beneath: the answers fingerprint compares the rebuild's answers on the served month with the stored ones. It says the rebuild is different. It cannot say why.

The one field that caught every guessed seed was the answers fingerprint: run the rebuilt model on the stored month, hash its answers, and compare with the hash in the record. All 19 guesses failed that check. So I need to correct the simple version of this lesson's headline. A missing seed is silent to a record that only describes the inputs. It is not silent to a record that also stores the model's answers on a fixed set of rows. What that check cannot do is tell you what is missing. It says "this is not the model", and then you are back to the baker with her notebook, wondering which note was lost.

Why does the seed matter at all here? Lesson 2 measured the answer for this model and data: by default, scikit-learn's boosted trees set aside a random 10% of the training rows to decide when to stop, and the seed decides which 10%. With that setting switched off, every seed in lesson 2 gave the same model. I did not repeat that test in this lab, so for these 30,000 rows it is the likely reason rather than a measurement.

Close in Accuracy, Not the Same Model

Agreement measures whether a rebuild is the same model. Accuracy measures whether it is a good one. They answer different questions, and after the main results I designed a follow-up to measure accuracy too. I wrote its design into the lab file's followup mode before it ran, and its results are in results/lineage_followup.json. It scored every rebuild against the real labels, on the served month and on the next 6,000 rows, 36,000 to 42,000, which no rebuild had trained on. Like the served "month", this next one is really 125 days, up to 28 September 1998.

A scatter chart headed follow-up, designed after the results: each guessed seed, agreement against accuracy, titled close in accuracy, still not the same model. Horizontal axis: agreement with the original's answers, 0.91 to 1.00. Vertical axis: accuracy on the served month, 0.68 to 0.72, with a dashed line marked original near 0.707. Nineteen dots for the guessed seeds spread between agreement about 0.92 and 0.97 and accuracy about 0.69 and 0.71; one dot for the full record sits on the original line at agreement 1.00. Beneath: guesses scored 0.6912 to 0.7100 on the served month, at most 1.55 points from the original's 0.7067; next month 0.5760 to 0.6063, up to 2.63 points away.

The original scored 0.7067 on the served month. The seed guesses scored from 0.6912 to 0.7100, so the furthest was 1.55 points away, where a point is one hundredth of accuracy. On the next month the original scored 0.6023 and the guesses 0.5760 to 0.6063, up to 2.63 points away. Three guesses were more accurate than the original on the served month, and two on the next month.

So in accuracy the guesses were close, and in some cases better. If all you need is "a model about as good as last month's", a guessed seed would do. But a rollback in an incident usually needs more than that. You want the answers customers actually received, so you can compare them with the new model's answers, find which rows changed, and explain what happened. For that, a model that differs on 3 to 8% of the rows is the wrong evidence. And there is no dot on this chart that tells you which of the guesses is closest, because without the stored answers you would not know where the original sits.

Where the Guesses Disagree

If the guesses disagree with the original on 3 to 8% of rows, which rows are those? I measured this after the results, using the original model's own confidence: for every served row, its estimated probability of UP, a number from 0 to 1 that the model turns into an answer by saying UP above 0.5.

Two panels headed where the guessed seeds disagree with the original; after the results, titled on the rows the original was unsure about. Borderline, 504 rows: 87.5%: at least one guess disagrees; the original's chance of UP was 0.4 to 0.6. The other 5,496 rows: 11.6%: at least one guess disagrees. Beneath: 61.4% of all disagreements came from the borderline rows, which are 8.4% of the month.

I called a row borderline when the original's probability of UP was between 0.4 and 0.6, a line I chose after seeing the results. 504 of the 6,000 served rows, 8.4%, were borderline. On 87.5% of them, at least one of the 19 guesses gave a different answer. On the other 5,496 rows, only 11.6% had any disagreement. Of all the single disagreements, one guess on one row, 61.4% came from the borderline rows.

A bar chart headed served rows by how many of the 19 guesses disagree with the original; after the results, titled most rows agree every time; a thin edge flips. Six bars on a scale from 0 to 5,000 rows: 0 guesses disagree, a tall bar near 4,900; 1, about 330; 2 to 4, about 350; 5 to 9, about 210; 10 to 14, about 120; 15 to 19, under 100. Beneath: 4,919, 330, 350, 213, 121, 67 rows. 7 rows: all 19 guesses disagree.

On 4,919 rows, every guess gave the original's answer. On 1,081 rows at least one guess did not, and on 7 rows all 19 disagreed. The picture is of a model with a firm middle and a thin, soft edge. Change the seed and the edge moves; the middle stays.

This matters for an incident. The rows that change from seed to seed are the rows where the model was least sure. Counted one guess at a time and averaged over the 19, a single guessed seed changed about one borderline answer in three (3,192 of 504 x 19 = 9,576) against about one in fifty elsewhere (2,003 of 5,496 x 19 = 104,424). So a rebuild with a guessed seed gets the firm rows right and changes a third of the unsure ones. In my opinion, not something I measured, the unsure rows are often the ones an investigation most needs.

When the Data Moved Underneath

The two other failures come from the data, not from the record's settings. Both are common in real systems.

An isometric drawing of three blocks, heights to scale, headed rows shifted: the old range and the new one, titled "the last 30,000 rows" moved by a month. Left: a short block, 6,000 rows, only in the old range. Middle: a tall block, 24,000 rows, in both. Right: a short block, 6,000 rows, only in the new: the served month. Beneath: the rebuild trained on the month it was meant to be judged on. It changed 1,569 of 6,000 served answers.

Rows shifted. Many pipelines describe their training data with a rule instead of a list: "the last 30,000 rows", "the last 90 days". The rule is convenient, because the same pipeline keeps working as new data arrives. But the rule is not a record. Read a month later, it points at different rows. Here the new range shared 24,000 rows with the old one, dropped the oldest 6,000, and added 6,000 new ones, and the new ones were the served month itself. The rebuild changed 1,569 of the 6,000 served answers. The data fingerprint fired, as it should, because the rows were different.

Two panels headed data backfilled: 1% of the old labels corrected; after the results, titled a small correction, a loud fingerprint. Labels corrected: 282: of 30,000: 121 UP to DOWN, 161 DOWN to UP. Answers changed: 250: of the 6,000 served; the data fingerprint fired. Beneath: the record did its job here: it said the data was no longer the data the model learned from.

Data backfilled. Stored data is not always frozen. An upstream team finds that some old records were wrong and corrects them in place. Here 282 labels changed, 121 from UP to DOWN and 161 from DOWN to UP, under 1% of 30,000. The rebuild changed 250 of the 6,000 served answers, and the data fingerprint fired.

In both cases the record did its job. It could not stop the data from changing, but it said, clearly, that the data was not the data the model had learned from. That turns a mysterious difference into a known one. The fix is also clear: keep a copy of the exact training rows, or keep a stored version of the data that nobody edits in place, so that the rows named in the record can still be read as they were.

A Wrong Rebuild Can Look Better

The follow-up's accuracy numbers had one more surprise, and I think it is the most useful result in this lesson.

A bar chart headed follow-up, designed after the results: accuracy against the real labels, titled the wrong rebuilds scored higher. Three pairs of bars, served month and next month, on a scale from 0.5 to 1.0. Original: about 0.71 and 0.60. Backfilled: about 0.71 and 0.62. Shifted: about 0.92 and 0.71. Beneath: served month: 0.7067, 0.7103, 0.9182. Next month: 0.6023, 0.6192, 0.7128. The axis starts at 0.50.

The rows-shifted rebuild scored 0.9182 on the served month, where the original scored 0.7067. That is not because it is a better model. It trained on those very rows, so it had seen the answers. On the next month, which neither had seen, it scored 0.7128 against the original's 0.6023. That gap is more believable: the shifted model learned from data 6,000 rows closer to the next month, which is consistent with lesson 6, where retraining on newer rows helped, largely through the date input; I did not test that here. The backfilled rebuild also scored a little higher than the original, 0.7103 and 0.6192. I did not study why, and one run cannot tell it apart from chance.

Two panels headed rows shifted, served month; follow-up, after the results, titled a better score is not the model that served. Original: 0.7067: what customers actually got that month. Rows shifted: 0.9182: it had trained on those very rows. Beneath: investigating an incident with the shifted rebuild would study a model that never ran. Agreement was 0.7385.

Here is why this matters. Imagine the team rebuilds last month's model with the "last 30,000 rows" rule, checks its accuracy on last month's rows, and sees 0.92. It looks excellent. Someone concludes that last month's model was very good and the new model broke something serious. But the model that really served last month scored 0.71 on those rows. The whole investigation would start from a false picture, and every number on the rebuilt model would look better than the truth.

The lesson I take from this: never judge a rebuild by its accuracy. Judge it by whether it is the same model, and the only direct test of that is agreement with the answers the original actually gave.

Rolling Back: Load the File

If rebuilding is fragile, the obvious alternative is not to rebuild at all. Keep the model file, and on the bad morning, load it.

A sequence diagram with four columns: engineer, registry, serving, check. Headed a rollback from a registry, titled load, check, switch: no training at all. Step 1, the engineer asks the registry for last month's version. Step 2, the registry sends back the model file and record. Step 3, the engineer sends the check the answers on the stored month. Step 4, the check replies: fingerprint matches. Step 5, the engineer tells serving: serve this file. Step 6, serving replies: old model live. Beneath: step 3 is the check that caught every guessed seed: 19 of 19. Here the loaded file matched.

A rollback from a registry has three moves. Fetch last month's file and its record by version. Load the file, run it on the stored month, and compare the hash of its answers with the one in the record. If they match, serve it. No training happens at all, so nothing random can change, and no data needs to be read except the stored month for the check. Because loading a pickle can run code, the registry must be one you control: never load a model file from an untrusted source.

The follow-up measured two things about this road, both designed after the main results and written down before they ran.

Two panels headed follow-up, designed after the results: thread count and time, titled thread count did not change the rebuild; loading was fast. 1 or 2 threads: 1.0000: agreement of the full-record rebuild, both times; the same answers: yes. Median time: 0.0075 s: load the file and predict; a refit took 0.66 s. Beneath: times: 5 repeats on one Mac while other jobs ran, so rough. A new library version was not tested.

A different number of threads. A thread is one of several lines of work a program can run at once on different processor cores. A different machine often has a different number of cores, so I ran the full-record rebuild in a fresh process limited to 1 thread, and again to 2, where the main run used all 10 cores of this Mac. Both gave the same 6,000 answers, with the answer hash matching. For this model on this machine, the thread count did not change the result. That is one honest piece of "a different machine"; it is not the whole of it.

Time. Loading the file and answering the 6,000 rows took a median of 0.0075 seconds over 5 repeats. Training again and answering took a median of 0.66 seconds. Other jobs were running on the same Mac at the time, with a load average of 5.8 on 10 cores (the number of jobs using or waiting for a processor core, averaged over the last minute), so these times are rough. The size of the gap is the point, not the decimals: loading was about 88 times faster here, and for a large model that trains for days, the difference is days.

What I Did Not Measure: Library Versions

The record stores the versions of scikit-learn and numpy for a reason. A new version of a library can change the default value of a setting, the order of an internal calculation, or the way a saved file is read. Any of these can change a rebuilt model, or stop an old model file from loading at all. scikit-learn's own documentation says there is no supported way to load a model trained with a different version: it might load, but that is "entirely unsupported and inadvisable". A file that loads with only a warning can still behave differently, and it can also simply be the wrong file, which is why the answers check matters on the loading road too.

I did not test this. Testing it honestly would mean installing a second version of scikit-learn next to the first, and the Python environment on this machine is shared with other work, so I left it alone. Everything in this lesson was measured with scikit-learn 1.9.1 and numpy 2.5.3, and I cannot tell you how large the effect of a version change would be on this model. It could be nothing, or it could be a model that no longer loads.

What I can say from the design is this. A version change is exactly the kind of difference the data fingerprint cannot see, because the data has not changed. Like the missing seed, it would pass every check on the inputs. The answers fingerprint would still catch it, if the rebuilt or reloaded model answered the stored month differently. That is one more reason to store the answers, and one more reason to keep the exact environment that trained a model, for example as a list of pinned library versions, next to its file.

Try It Yourself

This script is the lab made small. It downloads the same data, trains last month's model with seed 7, writes a lineage record with a data fingerprint and an answers fingerprint, and saves the model to bytes. Then it tries the five ways back: the full record, the saved file, backfilled labels, shifted rows, and the 19 guessed seeds. For each it prints the agreement, whether the data fingerprint still matches, and whether the answers are the same. It does not need a GPU.

A real screenshot of VS Code with lineage_demo.py open, showing the docstring that says what the script is and how to run it, the imports of hashlib, pickle, numpy, scikit-learn and its boosted trees, the lines that load Elec2 and turn the label into 1 for UP and 0 for DOWN, the served rows 30,000 to 36,000, the sha function that fingerprints arrays, and the start of the lineage record; the five rebuilds are further down. Beneath: copy it from the box on the slide.

Before you run this lab. You need Python 3 and two libraries: pip install scikit-learn pandas. scikit-learn holds the model and the download, and brings NumPy with it; pandas holds the table; hashlib and pickle come with Python. The first run downloads Elec2 from OpenML, a free public website of datasets for machine learning (under 1 MB compressed), so it needs an internet connection once; scikit-learn keeps a copy in a folder in your home directory (scikit_learn_data) for later runs. It trains 23 small models, so it is quick; I did not time it carefully. I ran it with scikit-learn 1.9.1. scikit-learn runs the same way on Windows and Linux, but I have not checked the numbers there, and another version of scikit-learn may give different models, which is this lesson's point; that is why the first line printed is the version.

"""Lineage and rollback: rebuild last month's model, exactly.

Lesson 8 of 'The ML & AI Lifecycle', made small. It needs Python 3 with
scikit-learn and pandas (pip install scikit-learn pandas). The first run
downloads Elec2 from OpenML (under 1 MB) and keeps a copy for later runs.
    python lineage_demo.py

Author: Roni Das
Created: 2026-09-29
"""
import hashlib
import pickle

import numpy as np
import sklearn
from sklearn.datasets import fetch_openml
from sklearn.ensemble import HistGradientBoostingClassifier as Trees

# 45,312 half-hours in time order. The label: is the NSW price UP (1)
# or DOWN (0) against the average of the last 24 hours?
data = fetch_openml("electricity", version=1, as_frame=True,
                    parser="auto").frame
X = data.drop(columns="class").astype(float)
y = (data["class"].astype(str) == "UP").astype(int).to_numpy()
MONTH = slice(30000, 36000)  # the month the model served


def sha(*arrays):
    # a SHA-256 fingerprint of the exact numbers in the arrays
    h = hashlib.sha256()
    for a in arrays:
        h.update(np.ascontiguousarray(a).tobytes())
    return h.hexdigest()


def data_hash(lo, hi, labels):
    return sha(X.iloc[lo:hi].to_numpy(), labels[lo:hi])


# Last month's model, and its lineage record.
model = Trees(random_state=7).fit(X.iloc[0:30000], y[0:30000])
served = model.predict(X.iloc[MONTH])
record = {"rows": (0, 30000), "seed": 7,
          "data": data_hash(0, 30000, y), "answers": sha(served),
          "sklearn": sklearn.__version__}
blob = pickle.dumps(model)  # the saved model file
print(f"scikit-learn {record['sklearn']}")
print(f"data hash {record['data'][:12]}")
print(f"answers hash {record['answers'][:12]}")
print(f"model file {len(blob):,} bytes")


def rebuild(name, lo, hi, labels, seed):
    m = Trees(random_state=seed).fit(X.iloc[lo:hi], labels[lo:hi])
    show(name, m.predict(X.iloc[MONTH]), data_hash(lo, hi, labels))


def show(name, answers, dhash):
    agree = (answers == served).mean()
    data = "ok" if dhash == record["data"] else "CHANGED"
    same = "same" if sha(answers) == record["answers"] else "differ"
    print(f"{name:<13} agree {agree:.4f}  data {data:<7}"
          f"  answers {same}")


# 1. Everything in the record, used as written.
rebuild("full record", 0, 30000, y, seed=7)

# 2. No rebuild: load the saved file back.
show("snapshot", pickle.loads(blob).predict(X.iloc[MONTH]),
     record["data"])

# 3. Upstream corrected 1% of the old labels after the model was built.
flip = np.random.default_rng(8).random(30000) < 0.01
y_new = y.copy()
y_new[:30000][flip] = 1 - y_new[:30000][flip]
rebuild("backfilled", 0, 30000, y_new, seed=7)

# 4. The rows were written down as "the last 30,000", a month later.
rebuild("rows shifted", 6000, 36000, y, seed=7)

# 5. The seed was never written down: try the other 19 seeds 0 to 19.
agree = []
for seed in [s for s in range(20) if s != 7]:
    m = Trees(random_state=seed).fit(X.iloc[0:30000], y[0:30000])
    agree.append((m.predict(X.iloc[MONTH]) == served).mean())
exact = sum(a == 1.0 for a in agree)
print("seed not recorded: 19 guesses")
print(f"  agree {min(agree):.4f} to {max(agree):.4f}, exact {exact}")

The Lab Report

A real terminal recording headed python lineage_report.py, titled every table in this lesson, from the stored files and the data. It opens with last month's model, trees, seed 7, rows 0 to 30,000, serving rows 30,000 to 36,000, and 55 checks against the refits: all agree. Then six numbered sections: 1, the main run: full record and snapshot 1.0000, data backfilled 0.9583, rows shifted 0.7385, 19 guesses from 0.9197 to 0.9702; 2, which field of the record fires, with the seed guesses quiet on the data hash and 19 of 19 on the answers hash; 3, accuracy on the served and next month; 4, where the seed guesses disagree, with 504 borderline rows; 5, the two data changes and the pickled bytes, the same as the original's; 6, the thread counts and the time. Beneath: the lab's own report. It fits the original and every rebuild again and stops unless every stored number comes back.

The report lives in scripts/labs/lifecycle/lineage_report.py. It reads the lab's stored files, results/lineage.json and results/lineage_followup.json, and the Elec2 data from scikit-learn's local copy. The lab stored agreements and fingerprints, not each answer, so the report trains the original model and every rebuild again with the same rows, labels, settings and seeds, and stops unless every stored number and fingerprint comes back exactly: the record's two fingerprints, the file size, each rebuild's agreement and warnings, the seed guesses' range, middle value and exact count, and every accuracy in the follow-up. It makes 55 checks in all, and they all agree. It changes nothing in the lab's files.

Its json mode writes every number to results/rv-report.json, which the figures read. The demo mode checks the student script's stored run, and the box mode writes the playground below and checks that it gives the lab's numbers.

What came before the run, in lineage_lab.py: the model, the record's fields, the five rebuilds and what to report. What came after I saw the main results: the follow-up's accuracy on two months, the thread counts and the times, designed after the results and written down before it ran. What came after all the results, in the report: which record field fires, where the seed guesses disagree and the borderline line at 0.4 and 0.6, the flipped labels by class, the shared rows, the pickled bytes, and the calendar dates.

Check a Rebuild Yourself

This box has no model in it. It holds, for every one of the 6,000 served rows, the original model's answer, the answers of the 19 guessed seeds and of the backfilled and shifted rebuilds, the real label, and whether the row was borderline, as six hexadecimal digits per row (counting in sixteens, with the digits 0 to 9 and a to f). It also holds the answers fingerprint from the lab's record. The report checked that every agreement, accuracy and fingerprint it gives matches the lab's files. It runs in your browser.

As it is, the box checks that the original's answers hash to the fingerprint stored in the record, which prints True, and then prints the backfilled rebuild, the shifted rebuild and the first three guessed seeds with their agreement, accuracy and whether their answers match the record. None of them does.

Then try agreement(s) and accuracy(s) for every seed in GUESSED_SEEDS, and look for the one closest to the original. Try disagree_on_border(3) to see how much of one guess's disagreement sits on borderline rows. And change one answer by hand, for example ORIGINAL[0] = 1 - ORIGINAL[0], then run answers_hash('original') again: one changed answer out of 6,000 gives a completely different fingerprint.

The Code, Part by Part

Loading. fetch_openml("electricity", version=1) downloads Elec2 once and reads the local copy after that. The label column holds the words UP and DOWN, turned into 1 and 0. MONTH is the slice of rows 30,000 to 36,000, the period the model served.

sha and data_hash. sha feeds the exact bytes of one or more arrays into SHA-256 and returns the 64-character fingerprint. np.ascontiguousarray makes sure the numbers are laid out in memory in one fixed order before hashing, so the same table always gives the same bytes. data_hash fingerprints the inputs and labels of a range of rows.

The record. The model is trained on rows 0 to 30,000 with random_state=7, which is scikit-learn's name for the seed. Its answers on the served rows are kept in served. The record is a plain dictionary: the rows, the seed, the data fingerprint, the answers fingerprint and the scikit-learn version. pickle.dumps(model) turns the trained model into bytes, the same bytes a model file would hold.

rebuild and show. rebuild trains a new model on a range of rows with a given set of labels and seed, and passes its answers to . prints the agreement with , whether the data fingerprint still matches the record, and whether the answers fingerprint does.

How to Make a Model You Can Get Back

A hand-sketched column of six boxes joined by arrows, headed making a model you can get back, titled write it down, fingerprint it, keep the file. 1, write the rows as numbers, never as "the last N". 2, write the seed and every setting. 3, fingerprint the training data. 4, fingerprint the answers on a stored month. 5, save the model file with its record, by version. 6, on rollback, load the file and check the answers. Beneath: here the model file was 369,066 bytes: cheap to keep for every version.

Write the rows as numbers. Record the exact row numbers, dates or file versions a model learned from, never a rule like "the last 30,000 rows" or "the last 90 days". A rule is fine for choosing the data; the record must hold what the rule chose.

Write the seed and every setting. Set the seed yourself, write it down, and write down every setting, including the ones left at their defaults, together with the library versions. A default can change between versions. The lab's record stored only the seed; the defaults were implied by the library version. Write them out.

Fingerprint the training data. Store a hash of the exact rows. It costs 64 characters and tells you at once whether the data under a rebuild is still the data the model learned from.

Fingerprint the answers on a stored month. Keep a fixed set of rows, run the model on them, and store the hash of its answers. This is the only check in this lab that saw every kind of failure, including the missing seed.

Save the model file with its record, by version. Here the file was about a third of a megabyte. Keeping every version that ever served is almost always cheaper than one bad day without it.

On rollback, load the file and check the answers. Load it only from a registry you control, since loading a pickle can run code; run it on the stored month, compare the answers fingerprint, and only then serve it.

Load the Saved File, or Rebuild It?

A two-column table headed grounded in this lesson's numbers, titled load the saved file, or rebuild it? Left, load the saved file when: you need last month's model back, now; you need exactly its answers: 1.0000 here; the file and its record were kept by version; you are investigating what customers got. Right, rebuild from the record when: the file is lost, or cannot be loaded any more; you must retrain on corrected data, on purpose; the record holds rows, seed, settings and versions; you can check the answers fingerprint afterwards. Beneath, left: a load took 0.0075 s here. Beneath, right: a missing seed gave 0.920 to 0.970.

Load the saved file when you need last month's model back quickly. It is the fastest road and, here, an exact one.

Load the saved file when you need exactly the answers customers got. An investigation compares old answers with new ones. Only the real model gives the real old answers.

Load it when the file and its record were kept together by version. Then the record can confirm, with the answers fingerprint, that the file is the right one.

Rebuild when the file is lost, or will not load any more. A library upgrade can make an old file unreadable. Then the record is the only road back, and it must be complete.

Rebuild when you mean to change something. If upstream corrected the labels, training again on the corrected data is the right thing to do on purpose. Just do not call the result last month's model. The data fingerprint will tell you it is not.

Do not rebuild from a record that has no seed or no row numbers, and expect the old model. Here the missing seed gave 0.920 to 0.970 agreement, and the missing row numbers gave 0.739.

Do not judge any rebuild by its accuracy alone. The shifted rebuild scored 0.9182 on rows where the real model scored 0.7067.

What These Numbers Can and Cannot Tell You

A two-column page headed read before you trust these numbers, titled what these runs are, and what they are not. They are: one dataset, one model type; one run each, 19 seed guesses; faults I built on purpose; the design came first; accuracy, threads, time: after; one Mac, one library version; times while other jobs ran. They are not: not a rate for other data; not a law about seeds; not a bug found in the wild; not changed after the run; chosen knowing the results; versions not tested; rough, not a benchmark.

One dataset, one model type, one run each. Everything here is one electricity market and one kind of model, boosted trees with default settings. How far a missing seed moves a model depends on how much randomness the training uses. For a model with no randomness in its training, such as the logistic regression that lesson 2 trained 20 times into one single model, a missing seed would change nothing. For a large neural network, it can change a great deal. The numbers here are for this model.

Faults I built on purpose. The missing seed, the backfill and the shifted rows are faults I made in order to measure them. They are the kind of thing that happens in real systems, but a real backfill might change many more rows, or only a few.

The follow-up and the extra measures came after. Accuracy on two months, the thread counts and the times were designed after I saw the main results and written down before they ran. Which record field fires, the borderline line at 0.4 and 0.6, and the pickled bytes were measured after all the results, with rules I chose knowing them.

One machine, one library version. The thread test is one small piece of "a different machine". The library version was not tested at all. The times were measured while other jobs ran, so they are rough; only the size of the gap between loading and training should be read from them.

What to Do Next

A hand-drawn list headed before your next incident, titled five questions for every model in service. Which rows?: are they written as numbers, or as "the last month"? Which seed?: is the seed written down, with every setting? Same data?: is there a fingerprint of the training data to compare? Same answers?: are the answers on a stored month fingerprinted? The file?: is last month's model file kept, next to its record? Beneath: here a missing seed was invisible to the data check and seen by the answers check 19 times in 19.

Pick one model your team has in service, and ask the five questions on the card. The first three are about the record; the last two are about what lets you check it. If you can only fix one thing this week, fix the last: find last month's model file and make sure it is kept, with its record, where someone on call can find it. Then add the answers fingerprint. It is a few lines of code, like the sha(served) line in the demo, and it is the one check that caught every failure in this lab.

Then try a rehearsal. Before any incident, take last month's model and do a rollback on purpose: load the file, check the answers fingerprint, and time it. A rollback that has never been practised usually fails on the day it is needed, for a reason nobody expected.

The chapter plan's next lessons measure what happens when the right answers arrive late, and then put the whole lifecycle together as one loop.

A closing card headed to keep, titled keep the file, and fingerprint the answers. In large type: 1.000 from the file; 0.920 to 0.970 without the seed. Beneath: loading last month's saved model gave back every answer. Rebuilding it gave every answer back only with the full record; with the seed missing, none of 19 guesses was exact, and the data check stayed quiet. Then: changed data made the fingerprint fire. The shifted rebuild scored 0.9182 against the original's 0.7067: a better score is not the model that served. Last: one dataset, one model type, one run each: a way to check, not a law.

Knowledge Check

Knowledge Check

4 questions - Score 80% to pass

Q1

A rebuild with the seed not written down agreed with the original on 0.920 to 0.970 of served rows. Which part of the record noticed that it was a different model?

Q2

Why was loading the saved model file the safer road back to last month's model here?

Q3

The rows-shifted rebuild scored 0.9182 on the served month, against the original's 0.7067. Why is it the wrong model to study an incident with?

Q4

On which served rows did the guessed seeds disagree with the original most often?

label
accuracy
agreement

A hand-drawn sketch headed sketched: the data fingerprint, first 12 characters, real values, titled the same rows give the same fingerprint; one change gives a new one. Four rows, each a box with an arrow to a fingerprint box. Rows 0 to 30,000, as recorded: 7f391cd43a8a. The same rows, a month later: 7f391cd43a8a. 282 old labels corrected upstream: 00e4a5f1f6db, marked as different. "The last 30,000 rows", a month later: 3f77e94a3f82, marked as different. Beneath the boxes: compare the fingerprint in the record with the one you compute now. Beneath the sketch: 282 changed labels out of 30,000 rows were enough to change every character. A fingerprint says THAT the data changed, never what changed.

The sketch shows how the data fingerprint behaves, with the real first 12 characters the lab computed. The same rows give the same fingerprint, a month later or a year later. Change 282 labels out of 30,000, under 1% of them, and the fingerprint is completely different. Read a different set of rows and it is different again. A fingerprint is a yes-or-no test: the data is the same, or it is not. It cannot tell you which rows changed or how much. And because it hashes the raw bytes, a harmless change also fires it: the same numbers stored as a different number type, or the columns in a different order. Fix the format and the column order before you fingerprint, and keep them fixed.

One honest note about the setting: the month is simulated. Every rebuild ran seconds after the original, in the same run of the program, on the same machine. The snapshot test turned the model into bytes in memory and back again in the same process; the file was never written to disk and read by a fresh program. So this lab tests what the record contains, not what a real month does to a file, a disk or a machine.

For each rebuild the lab stored its agreement with the reference answers, whether the hash of its answers matched the stored one, and whether the record's data fingerprint would have warned that the data had changed. One run each, 19 for the seed guesses. The design described the numbers and declared no significance test, a calculation of how likely a difference is to come from chance alone.

For this small model, both roads are fast. The real argument for the snapshot is not speed. It is that loading a file cannot silently produce a different model on the same library versions, and the answers check confirms it, while rebuilding can.

This is a real run in VS Code's terminal (python lineage_demo.py).

A real screenshot of VS Code's terminal after running python lineage_demo.py. It prints scikit-learn 1.9.1; data hash 7f391cd43a8a; answers hash b978599166b1; model file 369,066 bytes; full record, agree 1.0000, data ok, answers same; snapshot, agree 1.0000, data ok, answers same; backfilled, agree 0.9583, data CHANGED, answers differ; rows shifted, agree 0.7385, data CHANGED, answers differ; seed not recorded: 19 guesses; agree 0.9197 to 0.9702, exact 0.

When I ran it, it printed scikit-learn 1.9.1, the first 12 characters of both fingerprints, 7f391cd43a8a and b978599166b1, and a model file of 369,066 bytes. Then the five ways back: full record and snapshot at 1.0000 with the data fingerprint ok and the same answers; backfilled at 0.9583 and rows shifted at 0.7385, each with the data fingerprint CHANGED; and the 19 guessed seeds from 0.9197 to 0.9702 with 0 exact. All of it matches the lab's stored lineage.json, including both fingerprints and the file size, and the longest printed line was 56 characters. The report's demo mode checks all of it.

To see the quiet failure for yourself, change seed=7 to seed=3 in the full-record line, as if the seed had been lost and guessed. When I did, that line printed agree 0.9575, data ok, answers differ: the data check stayed quiet and only the answers fingerprint noticed. 0.9575 is the lab's stored agreement for seed 3.

Four brand cards in a two-by-two grid headed the tools, with their logos, titled what ran where. scikit-learn: the trees and the data download. Python: hashlib and pickle. NumPy: the rows and the hashes. pandas: the table of half-hours.

show
show
served

The five ways back. The full record passes the recorded rows and seed. The snapshot calls pickle.loads(blob) and uses the loaded model directly. The snapshot needs no training data, so its data check is not run: the script passes the record's own fingerprint, so that line always says ok. The backfill flips the labels of a random 1% of rows 0 to 30,000, chosen with seed 8 as in the lab, on a copy of the labels. The shifted rows use 6,000 to 36,000. The last loop trains one model for each seed from 0 to 19 except 7 and prints the lowest and highest agreement and how many were exact.

In the lab file, lineage_lab.py does the same, also records the numpy version, and stores the numbers in results/lineage.json. Its followup mode adds accuracy on two months, the thread counts and the times.