Packaging And Registry

Library Version Skew: Old Model Files Broke on Upgrade, and Retraining Made a Different Model

0 of 26 complete

0%

Contents

Back|Packaging And RegistryLibrary Version Skew: Old Model Files Broke on Upgrade, and Retraining Made a Different Model
1/26
59 min left
Prerequisites
Saving Formats Compared: The Scores Never Moved, the Bytes DidrequiredWhat Is in a Model File: It Is Mostly Numbers, and Loading It Runs CoderequiredWhat a Feature Is: A Better Model or a Better Feature?required
Related Topics
Same Code, Different Model: What Changes Between Two Identical Training RunsThe ML & AI LifecycleLineage and Rollback: Rebuilding Last Month's Model Exactly, and What Breaks When You CannotThe ML & AI LifecycleThe Lifecycle as One Loop: Nine Steps, Nine Measured Failures, and What Saw Each OneThe ML & AI LifecycleWrong Labels: How Many Can a Model Survive, and Can You Find Them?Data Engineering for MLRebalance, or Just Move the Threshold? Measured on Rare ClassesData Engineering for ML
1 of 26

A List of Steps That Names Every Part

Let me start with a long printed list of steps, like the one in the picture.

Imagine a workshop where a machine prints instructions for building something. Each step says exactly where to find a part: "take the small spring from drawer 14, shelf B". Another workshop, built the same way, can follow the list and build the same thing. But what if the second workshop is newer, and someone moved the small spring to drawer 15? The list still says drawer 14. The worker stops at that step, because the part is not where the list says.

A flat illustration of a man in a light room unrolling a very long sheet of paper from a small printer on a desk; the paper falls in loops across the floor. Below the scene: a saved model is a list of steps, and each step names a part by its exact place. If the new workshop moved one part, the list stops at that step.

Now imagine the opposite case. The list was printed in the new workshop, and an older workshop follows it. If every part it names is also in the old workshop, in the same place, the old one builds the same thing. Nobody would guess which direction is safer without trying both.

A saved machine learning model is a list like this. In this lesson I try both directions, for real, across four releases of one library over three years.

Where This Lesson Starts

This lesson uses the model from the features chapter. If it is new to you, please read what a feature is first. That lesson built six features for each customer of a real online shop. It asked one question: at the start of a month, will this customer buy something in the next 30 days? The model is a gradient boosted tree model from scikit-learn, and it scored a test AP of 0.5450.

Three other lessons sit close to this one. The first lesson of this chapter opens a pickle file and shows what is inside. Here I only use one fact from it: a pickle file holds names of Python things, and loading it looks those names up. Saving formats compared saved the model eleven ways and loaded every file back, always with the same library versions on both sides. It ended by saying the cross-version question belongs to this lesson.

And lineage and rollback, in the lifecycle chapter, has a slide called "What I Did Not Measure: Library Versions". This lesson measures exactly that.

The Words You Need First

Please read this slide slowly if any word is new. Every slide after it uses these words.

A hand-drawn grid of nine cards, three per row. Library: code other people wrote that my code uses, like scikit-learn. Version: a numbered release of a library, like 1.7.2. Venv: a folder with its own Python and its own libraries. Pip freeze: the list of every library and version in one venv. Pickle: Python's way to save an object as bytes in a file. Load: read those bytes and build the object again. Version warning: what scikit-learn says when the saved and running versions differ. Loading skew: a file saved with one version, loaded with another. Training skew: the same training run, two versions, two models. Below: model, feature, test AP and cutoff mean what they meant in the features chapter.

A library is code that other people wrote and that my code uses. scikit-learn is a library for machine learning; numpy is a library for arrays of numbers that scikit-learn uses underneath. A version is a numbered release of a library, such as 1.7.2. A bigger number is a newer release.

A virtual environment, or venv, is a folder that holds its own copy of Python and its own libraries. Two venvs on one laptop can hold two different versions of scikit-learn, and they never mix. pip freeze prints the exact list of libraries and versions inside one venv.

To pickle a model is to save it as bytes in a file with Python's built-in tool. To load it is to read the file and build the model again in memory.

Skew means two things that should match but do not. In this lesson, loading skew means a file saved with one version and loaded with another. Training skew means the same training code run under two versions gives two different models. They are separate problems, and I measure them separately.

Six Venvs, Four Releases, Two Controls

To test versions, I need several versions installed at once. So the lab made six separate venvs with uv, a fast tool for making venvs and installing libraries.

An isometric row of six blocks, taller for newer scikit-learn. Labels above: 1.3.2, 1.5.2, 1.5.2, 1.7.2, 1.7.2, 1.9.1. Below each: numpy 1.26, 1.26, 2.1, 2.3, 2.5, 2.5, and Python 3.12, 3.12, 3.12, 3.13, 3.13, 3.13. A note: released 2023-10-23, 2024-09-11, 2025-09-09 and 2026-09-10; 1.3.2, 1.5.2 and 1.7.2 got the numpy and scipy of about a month later, with uv's exclude-newer option; 1.9.1 copies my main venv. Two controls change one thing: 1.5.2 with numpy 1 or 2, and 1.7.2 with its own numpy or exactly 1.9.1's.

The four main venvs hold scikit-learn 1.3.2, 1.5.2, 1.7.2 and 1.9.1. Those four were released on 23 October 2023, 11 September 2024, 9 September 2025 and 10 September 2026, so they span almost three years. 1.9.1 is the version in my main venv, where the whole features chapter ran.

A real team that installed scikit-learn 1.3.2 in late 2023 got the numpy of late 2023 with it. To copy that, I used uv's --exclude-newer option. Its own help text says it will "Limit candidate packages to those that were uploaded prior to the given date". So the 1.3.2 venv got numpy 1.26.2, the 1.5.2 venv got numpy 2.1.2, and the 1.7.2 venv got numpy 2.3.4, each from about a month after its release.

But then two things change at once: scikit-learn and numpy. So I added two control venvs, each changing only one thing. "1.5.2 with numpy 1" is the 1.5.2 venv with numpy held below 2, so the pair differs only in numpy. "1.7.2 with numpy 2.5" has numpy, scipy and the other helpers at exactly the versions in my 1.9.1 venv, so that pair differs only in scikit-learn. The lab stored every venv's full pip freeze. Together the six venvs use about 950 MB of disk.

What the Documentation Promises

Before running anything, I read what the libraries themselves say. Every quote here was found word for word in the source file at a fixed version, and all of them are in results/skew-factcheck.json.

Four rows with a logo or a name. scikit-learn 1.9.1, model persistence page: skops.io, pickle, joblib and cloudpickle, none of them support loading a model trained with a different version; it might load, but that is entirely unsupported and inadvisable. The warning, in scikit-learn's own code: on load, it compares the version string stored in the file with the running version, and warns if they differ; that is all the warning checks. numpy 2.0 release notes: numpy.core was renamed to numpy._core; no backward compatibility guarantees for internals. uv pip install exclude-newer: only packages uploaded before a date. Caption: every quote was found in the source, at the version named, before I used it.

scikit-learn's page on saving models is direct. It lists four ways to save a model with Python: skops.io, pickle, joblib and cloudpickle. Of those four it says: "Note that none of these methods support loading a model trained with a different version of scikit-learn, and possibly different versions of other dependencies such as numpy and scipy." It goes further: models "might load in other versions, however, this is entirely unsupported and inadvisable." And it warns that results "could give different and unexpected results, or even crash your Python process."

It also names the harder direction. "It is not always possible to load a model trained with older versions of the scikit-learn library and its dependencies in an updated software environment. Instead, you might need to retrain the model with the new versions of all the libraries."

Then I read the code that gives the warning. When scikit-learn saves a model, it adds one extra field, the version string, such as "1.7.2". When it loads one, it reads that field and runs one line: if pickle_version != __version__:. If the two strings differ, it raises an InconsistentVersionWarning, whose message says this "might lead to breaking code or invalid results. Use at your own risk." The check ends there. The warning does not look at the model at all. Keep this in mind for the results.

How the Lab Was Built

I wrote the design into the docstring of scripts/labs/packaging/version_skew.py before it first ran. I had read the documentation and the warning's code. I had never trained this model under any scikit-learn other than 1.9.1, and I did not know whether any load would fail.

A page in four labelled zones. Fixed: lesson 1's six columns, computed once and saved: 44,521 train rows, 26,851 test rows; the same numbers reach every venv. Train, in each venv: the chapter model with seeds 0 to 19; seed 0 saved with pickle; its scores on every test row stored. Load, every file in every venv: a fresh process each: does it load, which warnings, which error, and are the scores equal to the last bit. Scored in one place: test AP for every result computed in the main venv, so the metric code never changes. Caption: guesses, written down first: some load fails; every load across versions warns; training differs.

The feature code is not a variable. The main venv computed lesson 1's six columns once, using lesson 1's own code, and wrote them to one file of plain numbers. Every venv trains on exactly those numbers. So pandas, which builds the features, never runs inside the old venvs, and it cannot be the reason for any difference.

Train, in each venv. Each venv trains the chapter model, the gradient boosted model with all its default settings, with seeds 0 to 19. A seed is the number that fixes the random choices inside training. The seed-0 model is saved with pickle, and its predicted chance of buying for all 26,851 test rows is stored.

Load, every file in every venv. Six files, six venvs: 36 loads. Each one runs in its own new process, so a crash would end only that one job. For each load the lab records whether it loaded, every warning, and the exact error if there is one. Then it checks whether the predictions equal the original ones to the last bit, with np.array_equal.

Scored in one place. Test AP, the score from the features chapter, is computed for every result back in the main venv. So the scoring code is the same for all of them.

A control for the check. The same model with seed 1 must come back NOT identical to seed 0. It did: all 26,851 rows differed, by up to 0.2729. So when the check says "identical", it means it.

What pickle.load Does With a Model File

Before the results, here is what happens inside a load, in the order it happens.

A sequence diagram with three lifelines: pickle.load, model.pkl and scikit-learn. Step 1, pickle.load reads the next step from the file. Step 2, the file answers with a name. Step 3, pickle.load asks scikit-learn to find that name. Step 4, missing: error, stop. Step 5, or found: build it. Step 6, versions differ: warn. Step 7, predict_proba. Step 8, scores, or an error. Caption: a load that fails at step 4 never reaches the warning at step 6.

A pickle file does not contain scikit-learn's code. It contains names, like sklearn.ensemble._hist_gradient_boosting.gradient_boosting.HistGradientBoostingClassifier, and the numbers that go inside each object. Loading reads one step at a time. When a step names a thing, Python imports that name from the libraries in the current venv.

If the name is not there, because the library renamed or moved it, the load stops with an error. If the name is there, Python builds the object. Only at that moment does scikit-learn's own code run, compare the two version strings, and warn.

So the order matters. A load that breaks on a missing name never reaches the warning. This is the drawer from the first slide: the worker stops at the missing part, before anyone checks which workshop printed the list.

The Lab's Report, Running

This is a real recording of the report script, skew_report.py, on the laptop where the lab ran.

A terminal recording of skew_report.py. It prints that features.npz matches its sha256, and that the main venv's test AP is 0.5450, equal to lesson 1's. A table of the six venvs with their scikit-learn, numpy and Python versions, pickle sizes, seed-0 AP, 20-seed mean and number of trees: five venvs at 0.5436 and 0.5446 with 58 trees, 1.9.1 at 0.5450 and 0.5452 with 54. Then: 1.3.2 to 1.7.2 every pair identical on all 20 seeds; 1.7.2 against 1.9.1, 0 of 20 seeds identical, seed 0 differs on 26,851 rows by up to 0.1909; the bootstrap interval minus 0.0010 to minus 0.0001. Then a six by six grid of loads marked same, no load or no pred, with warning counts. Then: 30 loads across venvs, 8 loaded and all identical, 20 load errors, none warned first, 2 errors at predict, 0 crashes. The last line: all 380 checks agree with the stored lab.

The report does not trust the lab's summary. It checks the fingerprint, or sha256, of every saved file. It recomputes every AP with its own loop, every "identical" with its own numpy calls, and the bootstrap from the stored predictions. Then it loads every file in every venv again, each in a new process. Each load must give the same outcome as the lab: the same error and message, the same warnings, and the same predictions. All 380 checks agreed.

One honest note about the lab's history. The first run named its output files in a way that let some files overwrite others on disk. Each result had already been read before that happened, so the results were right, but the report could not reread them. I fixed the names and ran the whole lab again. The new results file was equal to the first one, key for key. The docstring says so.

The Headline: Most Loads Across Versions Failed

Here is the main result: every saved file, loaded in every venv.

A six by six grid. Rows: the venv that saved the file. Columns: the venv that loaded it. Each venv is labelled by its scikit-learn version and numpy version: 1.3.2 with numpy 1.26, 1.5.2 with 1.26, 1.5.2 with 2.1, 1.7.2 with 2.3, 1.7.2 with 2.5, 1.9.1 with 2.5. The diagonal is all same. The 1.3.2 row: same, predict fails, predict fails, then load fails three times. The 1.5.2 with numpy 1 row: load fails, same, same, then three load fails. The 1.5.2 with numpy 2.1 row: two load fails, same, three load fails. The two 1.7.2 rows: two load fails, then same three times, then load fails. The 1.9.1 row: two load fails, then same four times. Below: across venvs, 8 loaded, 20 failed at load, 2 failed at predict, none crashed. Caption: every load into a newer scikit-learn failed, 13 of 13, covering the five files that had a newer venv to try; newer files worked in older ones only with numpy 2 on both sides.

Of the 30 loads across different venvs, 22 failed. 20 failed while loading, and 2 loaded but then failed at the first prediction. None of them crashed the process.

Every load into a newer scikit-learn failed: 13 of 13, covering the five files that had a newer venv to try. A file saved with 1.7.2 did not load in 1.9.1. A file saved with 1.5.2 did not load in 1.7.2 or 1.9.1. A file saved with 1.3.2 did not load in 1.7.2 or 1.9.1, and in 1.5.2 it loaded but could not predict. This is the common real situation: a team upgrades the server's libraries, and last month's model file stops working.

The other direction worked more often. The 1.9.1 file loaded in the 1.5.2 and 1.7.2 venvs. The 1.7.2 file loaded in 1.5.2. But every file saved with numpy 2 failed in a venv with numpy 1. So here, a newer file worked in an older scikit-learn only when both sides had numpy 2. I would have guessed the opposite, and this is why I tried both directions.

When a load worked, the predictions were identical. All 8 loads across venvs that worked gave back exactly the original predictions on all 26,851 test rows, to the last bit. The 6 loads on the diagonal, each file in the venv that made it, were identical too. I did not see one case where a model loaded, predicted, and gave quietly different numbers. I tested one model type, so I cannot say that never happens.

What the Errors Said

Here are four of the errors, as Python printed them.

Four errors in two columns. The 1.7.2 file in 1.9.1: ModuleNotFoundError: No module named '_loss'. The 1.9.1 file in 1.3.2: ModuleNotFoundError: No module named 'numpy._core.numeric'. The 1.3.2 file in 1.9.1: AttributeError: Can't get attribute '__pyx_unpickle_CyHalfBinomialLoss', cut before its file path. The 1.3.2 file in 1.5.2, loaded, then: AttributeError: 'HistGradientBoostingClassifier' object has no attribute '_preprocessor'. Below: every one of the 20 load errors was one of two kinds, 13 ModuleNotFoundError and 7 AttributeError. Caption: not one of them said the word version.

A ModuleNotFoundError means Python looked for a module, a file of Python code, by name, and there was none with that name. An AttributeError means it found the module or object, but not the thing inside it that was asked for. All 20 load errors were one of these two: 13 ModuleNotFoundError and 7 AttributeError.

Read the messages again. None of them says "this file was made by another version". No module named '_loss' sounds like a broken install, not an upgrade. If you saw only this message on a server at night, you could spend an hour checking the install before thinking about versions.

The fourth error is different. That file loaded with no error. It failed only when it was asked to predict, because the newer code expected a part, _preprocessor, that the older model never had. A health check that only loads the file would have passed it.

What the Warning Did and Did Not Catch

Now back to scikit-learn's own warning, InconsistentVersionWarning. The lab counted it on every load.

Three panels. Loaded: 7; across scikit-learn versions; warned 2 or 3 times; 2 then failed at predict. Load failed: 0 of 20; warned before the error. One version: 3; same scikit-learn, other numpy: no version warning. Below: it warns once per scikit-learn object in the file: the model, its bin mapper and, in the 1.7.2 and 1.9.1 files, a label encoder. Caption: a warning means the versions differ. Silence did not mean the load was safe.

The warning came on all 7 loads that crossed a scikit-learn version and got far enough to build the model. It came 3 times for the 1.7.2 and 1.9.1 files, once for each scikit-learn object inside. Those are the model itself, its bin mapper (the part that sorts each feature's values into buckets), and a label encoder (the part that remembers the class names). It came twice for the 1.3.2 file, which has no label encoder inside.

It never came on the 20 loads that failed at load time. Those stopped at a missing name, before scikit-learn's code ran, exactly as the sequence slide showed.

And it did not come on the loads between venvs with the same scikit-learn and a different numpy. There were 4 such loads, and the 3 of those 4 that loaded gave no version warning. That is correct: the warning only compares scikit-learn's version string. But it means silence tells you nothing about numpy.

So the warning is useful in one way: if you see it, the versions differ. scikit-learn's documentation shows how to turn it into an error so you cannot miss it. But in this lab, a quiet load was never a sign that all was well.

The Load That Passed and the Prediction That Did Not

Two of the 22 failures were quieter than the rest. Here is one of them, step by step.

Three hand-drawn boxes, top to bottom, for the 1.3.2 file loaded in 1.5.2. 1, pickle.load: no error. 2, two version warnings: _BinMapper, HistGradientBoostingClassifier. 3, predict_proba: AttributeError: no attribute '_preprocessor'. Caption: the 1.5.2 venv with numpy 1 did the same.

The file saved with scikit-learn 1.3.2 loaded in 1.5.2 with no error at all. It gave two version warnings, one for the bin mapper and one for the model. In the venv with numpy 2 it also gave a numpy warning, which I explain two slides on. Then, on the first call to predict_proba, it failed: the 1.5.2 code looked for a part called _preprocessor, which the 1.3.2 model did not have.

Here is one possible reason, which I did not trace in scikit-learn's history. 1.5.2's prediction code expects every model to have that part, and has no fallback for models saved before it existed. The documentation never promised one.

Here is the practical point. A check that loads the model and stops would have said "fine". A check that loads it and predicts on a few stored rows would have caught it at once. And a check that compares those predictions with stored ones would catch even a model that predicts without an error but gives different numbers.

Why the Loads Failed: Names That Moved

I asked this question after seeing the results, and the lab labels it as an addition. Did each failure really come from a missing name?

A hand-drawn sketch, found after the results. On the left a box: 1.7.2's file asks for _loss.CyHalfBinomialLoss. Two arrows lead right: to a box saying 1.7.2: importing sklearn also makes _loss, and to a box saying 1.9.1: no module named _loss. Below: read from the files' bytes with pickletools, and checked in every venv: "a name it asks for is missing here" matched "did not load" in 36 of 36 cases. Caption: the file holds names, not code. A renamed name breaks the load.

The lab read every file's list of names straight from its bytes with pickletools, a module in Python's standard library that reads a pickle without loading it. Then, in each venv, it looked up every name the same way pickle does. "A name this file asks for is missing here" matched "this load failed" in all 36 cases.

One name tells the story well. scikit-learn keeps its loss functions, the code that measures how wrong each prediction is, in a compiled module whose full name is sklearn._loss._loss. In the 1.5.2 and 1.7.2 venvs, importing scikit-learn also registers that module under a short name, just _loss. So their files saved the short name. In 1.9.1 the short name does not exist, so those files fail with "No module named '_loss'". Nothing in my code changed. A detail of how the library was built changed.

I have to be honest about one thing here. The first version of this name check looked each name up without importing scikit-learn first. It said the 1.5.2 and 1.7.2 files were missing _loss even in their own venvs, where they load fine. That was a fault in my check, not a finding. pickle imports the model's class first, which imports scikit-learn, which creates the short name. I fixed the check to do the same, and the docstring records the fix.

The Control: numpy Alone

The control pair has scikit-learn 1.5.2 on both sides. Only numpy differs: 1.26.4 against 2.1.2.

A short page for scikit-learn 1.5.2 both times, only numpy differs. Saved with numpy 2, loaded with numpy 1: ModuleNotFoundError: No module named 'numpy._core.numeric'. Saved with numpy 1, loaded with numpy 2: loaded; scores identical; one DeprecationWarning: numpy.core is now numpy._core. Caption: no scikit-learn version warning either way: scikit-learn was the same.

A file saved with numpy 2 did not load with numpy 1. The error was No module named 'numpy._core.numeric'. numpy's 2.0 release notes explain the name: they renamed numpy.core to numpy._core. A numpy 1 venv has no numpy._core, so it cannot find the name.

The other direction loaded, with identical scores and one DeprecationWarning, a warning that an old name still works but will go away. numpy 2 kept the old name numpy.core working, so it could still read the old file.

This matters for two reasons. First, scikit-learn's warning stayed silent both times, because scikit-learn's version did not change. Second, it explains a row of the headline grid: every file saved with numpy 2 failed in both numpy 1 venvs, whatever the scikit-learn version. The other control pair, 1.7.2 with numpy 2.3 against 1.7.2 with numpy 2.5, loaded both ways with identical scores and no warning at all.

Training Again: A Different Model

Now the second question. Suppose a team does not load the old file at all. It trains again under the new version, same data, same settings, same seed. Does it get the same model?

A dot chart of test AP when trained in each venv, one dot per seed, 20 seeds, with the 20-seed mean marked. The five venvs from 1.3.2 to 1.7.2 have exactly the same dots, from 0.5340 to 0.5484. 1.9.1 has different dots, from 0.5385 to 0.5488. Below: 20-seed mean 0.5446 in each of the first five, 0.5452 in 1.9.1; seed 0, 0.5436 against 0.5450; the first five were equal to the last bit on all 20 seeds. Caption: 1.7.2 and 1.9.1, same numpy: 0 of 20 seeds alike.

From 1.3.2 to 1.7.2, across almost two years and two numpy major versions, every venv trained exactly the same model. Every seed gave predictions equal to the last bit, in every pair of those five venvs. The model stopped after 58 trees at seed 0 in all of them.

Then 1.9.1 trained a different model. Not one of the 20 seeds matched. With seed 0 it stopped after 54 trees instead of 58, and its predictions differed on all 26,851 test rows, by up to 0.1909. The control pair proves this is scikit-learn and not numpy. 1.7.2 with numpy 2.5 and 1.9.1 have the same numpy, scipy and Python, and they still disagree on every seed.

The scores moved a little. At seed 0, test AP was 0.5436 in the older versions and 0.5450 in 1.9.1. Over 20 seeds the means were 0.5446 and 0.5452.

A short page for customer 12346 at the cutoff 2011-11-01, seed 0. Two cards: trained in 1.3.2 to 1.7.2, 0.08280344823931944; trained in 1.9.1, 0.10317002256025683. Below: recency 286.6 days, frequency 12, money minus £64.68, return share 0.29, tenure 686.6 days, products 27: the same six numbers both times. The 1.9.1 file, loaded in 1.7.2, still gave 0.10317002256025683. Caption: loading kept the model. Training again made a different one.

Here is one real customer, 12346, who appears in every lesson of the features chapter. On 1 November 2011, with exactly the same six numbers going in, the older versions gave a chance of buying of 0.08280344823931944 and 1.9.1 gave 0.10317002256025683. That second number is the one lesson 2 of this chapter found in all eleven of its files. And when the 1.9.1 file was loaded in the 1.7.2 venv, it still gave 0.10317002256025683. Loading kept the model. Training again made a new one.

Why 1.9.1 Trained a Different Model

I asked this after the results too, and it is labelled as an addition. What changed between 1.7.2 and 1.9.1?

Hand-drawn bars, found after the results, of how many of each column's bin edges moved between 1.7.2 and 1.9.1, with the column's distinct training values in brackets. Recency days (35,806): 207. Frequency (126): 0. Money (14,542): 118. Return share (2,427): 96. Tenure days (39,817): 227. Products (548): 130. Below: bar, edges that differ, of the positions both models have; only frequency, with 126 values, under the 255-bin limit, kept every edge; 1.9.0's changelog says it changed how these edges are computed. Caption: consistent with that change; this does not prove it is the only cause.

This model does not look at a feature's exact values. Before training, it sorts each feature into at most 255 buckets, called bins, and the boundaries between buckets are called bin edges. Then every tree asks questions like "is this customer's bin above 40?".

scikit-learn 1.9.0's list of changes says it "Fixed the way" this model computes "their bin edges", a fix "to properly and consistently handle" sample weights. Sample weights are optional numbers that make some training rows count more than others. The entry goes on: with no weights and fewer distinct values than the number of buckets, "the edges are still set to midpoints between consecutive feature values". Otherwise, the edges are now "weight-aware quantiles computed using the averaged inverted CDF method". In plain words: they changed the formula that places the edges, for features with more distinct values than there are buckets.

This model was trained without sample weights, so the first rule applies to any column with few values.

So, before running, I wrote a guess into the docstring: the edges differ only on columns with more than 255 distinct values. The 1.7.2 venv could load both its own file and the 1.9.1 file, so I compared their edges there. The guess held.

Frequency has 126 distinct values, and kept all 123 of its edges. The five other columns all have more than 255 values, and edges moved on every one of them. The count went from 96 for return share to 227 for tenure. For two columns, 1.9.1 even made fewer edges. Return share kept only 97 edges instead of 254, and 96 of those 97 moved. Products kept 131 instead of 254, and 130 of those moved. I compared only the positions both models have.

Different edges mean different buckets, so different trees, so a different stopping point. That fits the result. But matching a changelog is not proof that it is the only cause, and I did not test that.

Is the Score Difference Real?

The 20-seed means differed by about six ten-thousandths. Is that a real difference, or customer luck?

A page with two large ranges. Minus 0.0010 to minus 0.0001: the 95% interval of the difference, over 1,000 resamples of customers. 0.5385 to 0.5488: 1.9.1's own range over the 20 seeds. Caption: the older versions scored a little lower on average, yet higher on 9 of the 20 seeds.

The features chapter set a rule for this question. A paired bootstrap draws the test customers again at random with replacement, 1,000 times, within each month, and recomputes both scores on each draw. Here each score is the mean over 20 seeds. The mean over 20 seeds damps training luck, and the resampling measures customer luck. The middle 95% of the differences is the interval.

The interval for the older versions minus 1.9.1 was minus 0.0010 to minus 0.0001. It does not cross zero, so by the chapter's rule the older versions scored measurably lower, on these customers, on average.

But look at the size. 1.9.1 alone ranged from 0.5385 to 0.5488 across its 20 seeds, a spread more than ten times bigger. Seed by seed, the older versions scored higher on 9 of the 20. So one training run, compared with one other run, can show either order.

The practical lesson is not "1.9.1 is better". It is this: a retrained model is a new model. Its score will move, maybe a little, maybe not in the direction you hope. Check it like any new model before it replaces the old one.

My Guesses Before the Run, Checked

I wrote six guesses into the lab before it ran. Here they are against the results.

  1. "the diagonal is bit-identical in every venv (a same-version load gives back the same model)." Right: 6 of 6.

  2. "every cross-version load emits at least one InconsistentVersionWarning; the numpy-only pair (1.5.2-np1 and 1.5.2, same scikit-learn) emits none." Wrong on the first part. 26 loads crossed a scikit-learn version; only 7 warned. The 19 others failed at load before any warning. Right on the second part: the numpy-only pair gave no scikit-learn warning.

  3. "at least one cross-version pair fails with an exception, at load or at predict." Right, but far too modest: 22 of 30.

  4. "a model saved under numpy 2 does not load under numpy 1.x (1.5.2 -> 1.5.2-np1 and newer -> 1.3.2)." Right. Every file saved with numpy 2 failed in both numpy 1 venvs.

  5. "when a cross-version load succeeds and predicts, the predictions are bit-identical to the original." Right: 8 of 8.

  6. "training under 1.3.2 and 1.9.1 with the same seed does NOT give bit-identical predictions, but their 20-seed mean test AP is within 0.005." Right on both parts: 0 of 20 seeds matched, and the means were 0.0006 apart. I did not guess that 1.3.2, 1.5.2 and 1.7.2 would all match each other exactly.

Try It Yourself

The full lab needs six venvs. I wrote a small demo that shows all three answers with just one extra venv: your normal Python, plus one venv with scikit-learn 1.7.2.

A page in three labelled zones, headed skew_demo.py, designed before it ran. Here, your Python: build lesson 1's columns, train, save with pickle. There, a venv with 1.7.2: train the same model; load the file made here. Here again: load the file made there; compare the two trained models. Caption: it printed 3 warnings and identical scores; ModuleNotFoundError; 26,851 rows differ.

I wrote the demo's design into its docstring after the lab had run and before the demo first ran. Your normal Python builds the features, trains, and saves. Then the script starts the second Python itself, with the path you give it, and that Python trains the same model and loads your file. Last, your Python loads the second Python's file.

A real screenshot of VS Code with skew_demo.py open at the top of the file, showing its docstring: the two Pythons it needs, the uv commands that make the second one, how to run it, and the design written before it first ran.

Before you run this lab. You need Python 3 with pandas, pyarrow and scikit-learn: pip install pandas pyarrow scikit-learn openpyxl. First run python fetch_data.py from the scripts/labs/features folder; it downloads the shop data once, about 46 MB. Then make the second venv. With uv: uv venv ~/skew-172 --python 3.13, then uv pip install --python ~/skew-172/bin/python scikit-learn==1.7.2 --exclude-newer 2025-10-15. Without uv, python3.13 -m venv ~/skew-172 and ~/skew-172/bin/pip install scikit-learn==1.7.2 work too, but may pick a newer numpy than mine.

Then, inside the examples folder, run . On Windows the second Python is at . I ran it on a Mac with scikit-learn 1.9.1 as "here", using the lab's own 1.7.2 venv, which was made with the same uv commands. These libraries run on Windows and Linux too, but I have not checked the results there. If your normal Python has a different scikit-learn than 1.9.1, your results can differ from mine; the lesson is about exactly this.

Look Up Any Pair Yourself

This box holds the real outcome of all 36 loads from the lab, and each venv's versions. It needs nothing but Python, so it runs in your browser. It loads no model.

Press Run. It shows what happened when the 1.7.2 file was loaded in 1.9.1, and then the whole grid. Change TRAINED_IN and LOADED_IN to any two of the six venv names and run again.

The report script writes this box from the lab's stored results, runs it for all 36 pairs, and checks every printed outcome against the lab. Try TRAINED_IN = "1.5.2" with LOADED_IN = "1.5.2-np1": the same scikit-learn, and still a failure, because of numpy alone.

The Lab's Code, Piece by Piece

The lab is one file, scripts/labs/packaging/version_skew.py. The same file runs in every venv, so its parts that run inside the old venvs use only numpy and scikit-learn, and nothing newer than Python 3.12.

setup makes the six venvs with uv and stores each pip freeze. build_features runs in the main venv only. It imports lesson 1's own feature code, writes the train and test numbers to one file, and stops if the chapter model's test AP is not 0.5450.

child_train runs inside one venv. It trains 20 seeds, saves the seed-0 model with pickle protocol 5, and stores every prediction. child_load also runs inside one venv. It loads one file while recording every warning, records the exact error if there is one, then predicts and stores the predictions.

main runs all of it: six trainings, then 36 loads, each in a new process. It compares every result with compare, which returns whether two arrays are bit-identical, how many rows differ, and the largest difference. bootstrap_mean20 does the paired bootstrap of the 20-seed mean.

after holds the two questions asked after the results. pickle_globals reads the names a file asks for with pickletools, without loading it, and looks those names up in each venv. compares the bin edges of the 1.7.2 and 1.9.1 models.

How to Upgrade Without Breaking a Model

Here is how I would handle a library upgrade, using only what this lab measured.

A flowchart. A new scikit-learn or numpy is out. Build a new venv from a pinned list. Does the old file load there? No leads to: retrain in the new venv, score it, ship both together. Yes leads to: identical scores on stored rows? No leads to retrain. Yes leads to: keep the file; still unsupported. Below: retraining gave a different model here, so score it like any new model. Caption: never upgrade the serving venv in place and hope.

The chart starts by building a new venv, not changing the one that serves. Here are the same steps in words.

  1. Pin every version. Save the exact version of every library, not only scikit-learn, next to the model file. numpy alone broke one direction here. Lesson 4 of this chapter is about doing this well.

  2. Build the new environment beside the old one. Never upgrade the libraries under a running model. Make a new venv or a new container with the new versions.

  3. Load last month's file there, and predict. Loading is not enough: two files here loaded and then failed at the first prediction.

  4. Compare predictions on stored rows. Keep a few hundred input rows and the model's predictions on them, next to the file. In the new venv, compare with np.array_equal. Every load here that predicted at all gave identical numbers, but one model type is not a promise.

  5. If anything fails, retrain in the new venv. Then treat the result as a new model: score it on the test months and compare it with the old one, because here it was a different model.

  6. Ship the model and its environment together. The model file and the venv it was made in are one thing. Lesson 5 of this chapter puts both in one container image.

When to Load Across Versions, and When Not To

Load across versions only as a test, never as a plan. scikit-learn says it is "entirely unsupported and inadvisable", and here 22 of 30 such loads failed. Say an old file does load in a new venv and gives identical predictions on stored rows. You can keep it for a while, but you are relying on luck the library does not promise.

Do not trust the warning to protect you. InconsistentVersionWarning came on only 7 of the 26 loads that crossed a scikit-learn version. Turning it into an error, as scikit-learn's page shows, is still worth doing: it stops a mixed-version load early. But it cannot see numpy, and it cannot speak for a load that fails before it runs.

Do not expect retraining to give the same model. Across 1.3.2 to 1.7.2 it did here, to the last bit. From 1.7.2 to 1.9.1 it did not, on any seed. A changelog entry was enough to change every prediction.

Do keep the original training environment until the new model has been checked. If the new venv cannot load the old model and the new model is not ready, the old venv is your way back. The lifecycle chapter's rollback idea applies to libraries too.

Do not read these numbers as a rule for every model. I tested one model type. A random forest, a pipeline with a text step, or a model from another library may break in other places, or not at all.

What This Lab Cannot Tell You

Two columns titled shows and cannot show. Shows: which loads worked and failed, for one model type, six venvs, one Mac; that loads which worked gave identical scores here; that 1.9.1 trains a different model from 1.3.2 to 1.7.2. Cannot show: other models, a forest or a pipeline may fail differently; other machines, Linux, Windows, another chip; that the bin-edge change is the only cause.

One model type. Every number here is for the chapter's gradient boosted model with default settings. Other models have other internal parts, and other names that can move.

Four releases, not every release. I tested 1.3.2, 1.5.2, 1.7.2 and 1.9.1. A pair I skipped, like 1.6 and 1.7, might behave differently. The two controls covered numpy, but not scipy or Python on their own.

One Mac. I did not try Linux, Windows or another processor. scikit-learn's page says: "Aside for a few exceptions, persisted models should be portable across operating systems and hardware architectures assuming the same versions of dependencies and Python are used". I did not test that.

No silent differences found does not mean none exist. All 8 working loads were bit-identical. A different model could load, predict and give slightly different numbers. That is exactly why the stored-rows check in the how-to is worth having.

No timings. This lab measured equality, not speed.

Labelled additions. Two questions were asked after the results: which names each file asks for, and where the 1.9.1 bin edges differ. Both are labelled in the lab, the report and here. The name check had a fault in its first run, which I fixed and described above.

What to Do on Monday

A hand-drawn grid of six cards, titled five steps. 1, pin everything: a lock file with exact versions, next to the model. 2, catch the warning: turn InconsistentVersionWarning into an error. 3, test loads in CI: load last month's file in the new venv. 4, compare scores: np.array_equal on stored rows, not just no error. 5, retrain on upgrade, and score it as a new model. The reason: 22 of 30 loads across venvs failed here. Caption: the file is only half the model. The venv is the other half.

If you take one thing to work on Monday, find where your model file is saved, and look at what is saved next to it. Is there a list of exact library versions? Are there some stored input rows and the predictions the model gave on them? If the answer is no, the file alone cannot tell you whether the next server can use it.

Then add one test to your build, the automatic checks that run on every change, often called CI. It loads the current model file in the environment you are about to deploy, predicts on the stored rows, and compares with np.array_equal. In this lab, that one test would have caught all 22 failures: 20 at load and 2 at predict.

A closing card titled a model file belongs to its venv. Three numbers in large type: 22 of 30, loads across venvs that failed, at load or at the first prediction; 8 of 8, loads that worked and gave identical scores; 0 of 20, seeds where 1.7.2 and 1.9.1 trained the same model.

The one idea to keep: a saved model is not only its file. It is the file plus the exact libraries it was made with. Here, moving the file to newer libraries broke every one of those loads, and training again under the newest version made a different model. Save the versions, test the load and the predictions before every upgrade, and treat a retrained model as a new one.

Knowledge Check

Knowledge Check

4 questions - Score 80% to pass

Q1

Of the 20 loads that failed while loading, how many gave an InconsistentVersionWarning first, and why?

Q2

The file saved with scikit-learn 1.3.2 loaded in 1.5.2 with no error. What happened next?

Q3

Why do the venvs 1.7.2 with numpy 2.5 and 1.9.1 matter for the training result?

Q4

The 20-seed mean AP was 0.5446 for the older versions and 0.5452 for 1.9.1. What does the lesson conclude?

python skew_demo.py ~/skew-172/bin/python
skew-172\Scripts\python.exe
r"""Train with one scikit-learn, load with another, and see what happens.

Lesson 3 of 'Packaging, Registry and Versioning'. It needs TWO Pythons:
  1. your normal one, with pandas, pyarrow and scikit-learn (I used 1.9.1),
     and the shop data from the features chapter: run
     scripts/labs/features/fetch_data.py once first;
  2. a second, small venv with an OLDER scikit-learn. With uv:
        uv venv ~/skew-172 --python 3.13
        uv pip install --python ~/skew-172/bin/python scikit-learn==1.7.2 \
            --exclude-newer 2025-10-15
     (or: python3.13 -m venv ~/skew-172, then
          ~/skew-172/bin/pip install scikit-learn==1.7.2)
Then, inside this folder, with your normal Python:
    python skew_demo.py ~/skew-172/bin/python            # print the results
    python skew_demo.py ~/skew-172/bin/python out.json   # and save them
It prints no timings.

Design, written 2026-10-01 after the lab (version_skew.py) had run and
before this file first ran:
  Here: lesson 1's six columns (lesson 1's own code), written once to a
  temporary .npz file, so both Pythons train on the same numbers. Then
  HistGradientBoostingClassifier(random_state=0) is trained here and
  saved with pickle.
  There (the second Python): the same model is trained from the same
  .npz and saved; then the file made HERE is loaded there.
  Here again: the file made THERE is loaded here.
  For each load it prints the scikit-learn warnings, the error if any,
  and whether predict_proba equals the original to the last bit. It
  also prints whether the two versions trained the same model.
  With 1.9.1 here and 1.7.2 there, the lab found: 1.9.1's file loads in
  1.7.2 with 3 warnings and identical scores; 1.7.2's file does not
  load in 1.9.1; the two trained models differ. skew_report.py checks
  this demo's saved run against the lab.

Author: Roni Das
Created: 2026-10-01
"""
import json
import pickle
import subprocess
import sys
import tempfile
import warnings
from pathlib import Path

import numpy as np


def try_load(path, x):
    """Load a pickle, catching every warning; return what happened."""
    out = {"warnings": [], "error": None, "proba": None}
    try:
        with warnings.catch_warnings(record=True) as ws:
            warnings.simplefilter("always")
            with open(path, "rb") as f:
                model = pickle.load(f)
            out["proba"] = model.predict_proba(x)
        out["warnings"] = [f"{w.category.__name__}: {str(w.message).splitlines()[0]}"
                           for w in ws]
    except Exception as e:  # the error IS the result here
        out["error"] = f"{type(e).__name__}: {e}"
    return out


def there(tmp):
    """Runs in the SECOND Python: train, save, and load the file made here."""
    import sklearn
    from sklearn.ensemble import HistGradientBoostingClassifier as HGB
    tmp = Path(tmp)
    d = np.load(tmp / "features.npz")
    model = HGB(random_state=0).fit(d["x_tr"], d["y_tr"])
    with open(tmp / "there.pkl", "wb") as f:
        pickle.dump(model, f, protocol=5)
    np.save(tmp / "there_trained.npy", model.predict_proba(d["x_te"]))
    got = try_load(tmp / "here.pkl", d["x_te"])
    if got["proba"] is not None:
        np.save(tmp / "here_loaded_there.npy", got.pop("proba"))
    got.pop("proba", None)
    got["sklearn"] = sklearn.__version__
    (tmp / "there.json").write_text(json.dumps(got))


if __name__ == "__main__" and sys.argv[1] == "--there":
    there(sys.argv[2])
    sys.exit(0)

import sklearn  # noqa: E402
from sklearn.ensemble import HistGradientBoostingClassifier as HGB  # noqa: E402

sys.path.insert(0, str(Path(__file__).resolve().parents[2] / "features"))
import task  # noqa: E402
from what_a_feature_is import HAND_COLS as COLS, joined  # noqa: E402

other = sys.argv[1]
ev = task.load_events()
lab_tr, _, lab_te = task.splits(ev)
tr = joined(ev, lab_tr, task.TRAIN_CUTOFFS)
te = joined(ev, lab_te, task.TEST_CUTOFFS)
x_te = te[COLS].to_numpy(float)

with tempfile.TemporaryDirectory() as tmp:
    tmp = Path(tmp)
    np.savez(tmp / "features.npz", x_tr=tr[COLS].to_numpy(float),
             y_tr=tr["label"].to_numpy(float), x_te=x_te)
    model = HGB(random_state=0).fit(tr[COLS].to_numpy(float), tr["label"])
    original = model.predict_proba(x_te)
    with open(tmp / "here.pkl", "wb") as f:
        pickle.dump(model, f, protocol=5)

    r = subprocess.run([other, __file__, "--there", str(tmp)])
    if r.returncode != 0:
        sys.exit(f"the second Python failed (return code {r.returncode})")
    t = json.loads((tmp / "there.json").read_text())
    there_trained = np.load(tmp / "there_trained.npy")
    back = tmp / "here_loaded_there.npy"
    here_in_there = np.load(back) if back.exists() else None
    there_in_here = try_load(tmp / "there.pkl", x_te)

here_v, there_v = sklearn.__version__, t["sklearn"]
out = {"here": here_v, "there": there_v, "test_rows": len(te)}
print(f"here: scikit-learn {here_v}   there: scikit-learn {there_v}")
print(f"\n1. a file saved with {here_v}, loaded with {there_v}")
for w in t["warnings"]:
    print("  ", w.split(" when using")[0])
print("   error:", t["error"])
same = None if here_in_there is None else bool(np.array_equal(here_in_there, original))
print(f"   identical on all {len(te):,} test rows: {same}")
out["new_in_old"] = {"warnings": len(t["warnings"]), "error": t["error"], "identical": same}

print(f"\n2. a file saved with {there_v}, loaded with {here_v}")
for w in there_in_here["warnings"]:
    print("  ", w.split(" when using")[0])
print("   error:", there_in_here["error"])
same2 = None
if there_in_here["proba"] is not None:
    same2 = bool(np.array_equal(there_in_here["proba"], there_trained))
print(f"   identical: {same2}")
out["old_in_new"] = {"warnings": len(there_in_here["warnings"]),
                     "error": there_in_here["error"], "identical": same2}

diff = np.abs(there_trained - original).max(axis=1)
print("\n3. trained in both, same data, same seed 0")
print(f"   identical: {bool(np.array_equal(there_trained, original))}; "
      f"rows that differ: {(diff > 0).sum():,}; largest difference: {diff.max():.4f}")
out["trained"] = {"identical": bool(np.array_equal(there_trained, original)),
                  "rows_differ": int((diff > 0).sum()), "max_abs": float(diff.max())}
if len(sys.argv) > 2:
    json.dump(out, open(sys.argv[2], "w"), indent=1)

This is a real run in VS Code's terminal, inside the examples folder.

A real screenshot of VS Code's terminal after running python skew_demo.py with the 1.7.2 venv's Python. It prints here: scikit-learn 1.9.1, there: 1.7.2. Part 1, a file saved with 1.9.1, loaded with 1.7.2: three InconsistentVersionWarning lines, for LabelEncoder, _BinMapper and HistGradientBoostingClassifier; error: None; identical on all 26,851 test rows: True. Part 2, a file saved with 1.7.2, loaded with 1.9.1: error: ModuleNotFoundError: No module named '_loss'. Part 3, trained in both: identical False; rows that differ 26,851; largest difference 0.1909.

When I ran it, it printed the same three answers as the lab. The 1.9.1 file loaded in 1.7.2 with three warnings and identical scores. The 1.7.2 file failed in 1.9.1 with No module named '_loss'. And the two trained models differed on all 26,851 rows, by up to 0.1909. The report script checks the demo's saved run against the lab.

child_names
child_bins

skew_report.py checks all of it from the outside, as the recording showed.

Two other tools come up in this discussion. ONNX is a model format that a separate program, called a runtime, can run. It is the documentation's option for a model that must run outside its training environment. I did not test it here, and not every model converts to it. skops, from lesson 2, is a safety choice about what a load may run, not a cure for version skew.