Packaging And Registry

Reproducible Artifacts: Same Scores, Different File: One Byte Followed the Thread Count

0 of 27 complete

0%

Contents

Back|Packaging And RegistryReproducible Artifacts: Same Scores, Different File: One Byte Followed the Thread Count
1/27
65 min left
Prerequisites
Saving Formats Compared: The Scores Never Moved, the Bytes DidrequiredContainer Images: Most of the Size Was the Base and the Libraries, and Slimming Kept Every ScorerequiredWhat a Feature Is: A Better Model or a Better Feature?required
1 of 27

Two Copies of the Same Book

Let me start with two copies of the same book.

A friend and I each buy the same novel. We read it, and the story is the same, word for word. Are the two books the same? For reading, yes. But if you look closely, one copy may have a small printing mark on the last page, or a different date of printing. The story is the same; the copies are not quite the same.

A flat illustration of a woman at a wooden desk holding a magnifying glass over an open book while she writes notes in a small notebook, a laptop beside her and plants on the shelves. Below the picture: two copies can tell the same story and still not be the same copy; to know they are the same, you compare them mark by mark, or you compare their fingerprints.

Now imagine your job is to check that the two copies are exactly the same, down to the last mark of ink. You could compare them page by page with a magnifying glass, like the woman in the picture. Or you could find a faster way: something like a fingerprint for a whole book, which changes if even one mark changes.

A trained model saved to a file has the same two questions. When I train the same model twice, I know from earlier lessons that the predictions can come out the same. But is the file the same, every byte? In this lesson I find out, and when the answer is no, I find the exact mark that differs.

Where This Lesson Starts

This lesson uses the model from the features chapter. If it is new to you, please read what a feature is first. That lesson built six features for each customer of a real online shop. It asked one question: at the start of a month, will this customer buy something in the next 30 days? The model is a gradient boosted tree model from scikit-learn, and it scored a test AP of 0.5450. AP, average precision, is a score from 0 to 1 for how well the model ranks the buyers above the others.

Two lessons from the ML lifecycle chapter sit close to this one, and I will not repeat them. Same code, different model measured how much the scores move when only the random seed changes. Lineage and rollback rebuilt last month's model from its written record and checked that its answers came back. Both ask about the model's behaviour.

This lesson asks about the file. In saving formats compared, lesson 2 of this chapter, one saved file came out one byte longer than another. The reason was a thread count stored inside the model. Here I follow that thread to the end.

The Words You Need First

Please read this slide slowly if any word is new. Every slide after it uses these words.

A hand-drawn grid of ten cards, two per row. Byte: the smallest piece of a file, one number from 0 to 255. Sha256: a fingerprint of a file's exact bytes, written as 64 letters and digits. Byte-identical: every byte the same, so the sha256 is the same too. Fresh process: a new run of Python that starts from nothing. Thread: one line of work that runs at the same time as others. OMP_NUM_THREADS: a setting that says how many threads the maths code may use. PYTHONHASHSEED: the seed Python mixes into its hashes of text. Compression level: how hard zlib squeezes a file, from 1 to 9. Reproducible: anyone can rebuild bit-by-bit identical copies. Hex: a way to write one byte as two characters, from 00 to ff. Below: model, test AP and the six columns mean what they meant in the features chapter.

A file is a long row of bytes. A byte is a small number from 0 to 255, and every file, a photo, a song or a model, is only a row of them. Two files are byte-identical when they have the same length and every byte in the same place is equal.

Comparing two big files byte by byte is slow if you have to keep both. So people use a hash. A hash function reads every byte of a file and gives back a short code. The one used here, sha256, gives 64 letters and digits. If even one byte changes, the code changes completely. That is why people call it a fingerprint: you can compare two fingerprints instead of two files.

A process is one running copy of a program. A fresh process is a new run of Python that starts from nothing, with nothing left over in memory from the run before. A thread is one line of work inside a process; a program can run several threads at once, one on each core of the processor. OMP_NUM_THREADS is a setting that tells the maths code how many threads it may use. PYTHONHASHSEED is a setting that controls a random number Python mixes into its hashes of text.

Finally, hex is a short way to write a byte as two characters: 00 is 0, 0a is 10, ff is 255.

What the Sources Say

Before measuring, I read what the primary sources say. I downloaded each source file at a fixed version and found every quote in it word for word. They are all in results/repro-factcheck.json.

A page titled six things the sources say, with six rows. Reproducible-builds.org, the definition: a build is reproducible if given the same source code, build environment and build instructions, any party can recreate bit-by-bit identical copies of all specified artifacts. Reproducible-builds.org, how it is checked: the reproducibility of artifacts is verified by bit-by-bit comparison; this is usually performed using cryptographically secure hash functions. Scikit-learn 1.9.1, parallelism, with its logo: by default, the implementations using OpenMP will use as many threads as possible, i.e. as many threads as logical cores. Scikit-learn 1.9.1, model persistence, with its logo: this should make it possible to check that the cross-validation score is in the same range as before. Python 3.13, PYTHONHASHSEED and sets, with its logo: if this variable is not set or set to random, a random value is used to seed the hashes of str and bytes objects; changing hash values affects the iteration order of sets. Python 3.13, the gzip module, with its logo: if mtime is omitted or None, the current time is used; use mtime = 0 to generate a compressed stream that does not depend on creation time. Caption: every quote was found in the source file, at the tag or commit named in the fact-check.

Reproducible builds. The Reproducible Builds project works on this question for software. I use its definition. "A build is reproducible if given the same source code, build environment and build instructions, any party can recreate bit-by-bit identical copies of all specified artifacts." An artifact is a file a build produces; here, the saved model. It also says how to check it: "The reproducibility of artifacts is verified by bit-by-bit comparison. This is usually performed using cryptographically secure hash functions." In short, it says, reproducible builds "provide certainty that software is genuine and has not been tampered with."

scikit-learn sets a softer goal for models. Its page on saving models asks you to keep notes. Then, it says, "This should make it possible to check that the cross-validation score is in the same range as before." So scikit-learn promises a similar model with a similar score, not the same file.

Its page on parallel work adds a fact that matters here: "By default, the implementations using OpenMP will use as many threads as possible, i.e. as many threads as logical cores." is the tool scikit-learn uses to run its C code on many threads. One detail: the 1.9.1 code itself counts physical cores by default, not logical ones. On this Mac both are 10.

How the Lab Was Built

I wrote the lab's design into the docstring of scripts/labs/packaging/reproducible_artifacts.py before it first ran. Before writing it, I knew one fact from lesson 2: the model stores a thread count. I had also listed the attribute names of a toy model, to see whether it held a set. I had not saved or hashed any file of the chapter model for this lesson.

A page in four labelled zones, titled one model, trained in 55 fresh processes. The model: lesson 1's model, test AP 0.5450; every process trains on the same 44,521 rows, read from one saved copy. Five files each time: pickle protocol 5; joblib with no compression, zlib 3, gzip 3 and lzma 3; the main process takes the sha256 of every file itself. One change at a time: the same process 5 times; 5 fresh processes; threads 1, 2, 4, 8 and 10, 5 each; hash seeds 0 to 4; 8 row orders; every compression level, 5 times each; and again at the end. When two files differ: which bytes, which pickle instruction, which part of the model; and the predictions on all 26,851 test rows. Caption: guesses, written first: the thread count changes the bytes and not the scores; the row order changes both.

One model, many trainings. The main program first trains lesson 1's model once and stops unless its test AP is exactly 0.5450. Then it saves the 44,521 training rows to one file. Every fresh process reads that same file and trains HistGradientBoostingClassifier(random_state=0) on it, so the data can never be the thing that differs.

Five files each time. Each trained model is saved five ways. One is pickle protocol 5. The other four are joblib: with no compression, and with zlib, gzip or lzma at level 3. These are the formats lesson 2 compared. The main program reads every file back and computes its sha256 itself. It does not trust the number the child process reports.

One change at a time. The default is a fresh process with nothing set: no thread limit and a random hash seed. Against it, I changed one thing at a time. The thread count went to 1, 2, 4, 8 and 10, five processes each. The hash seed went to 0, 1, 2, 3 and 4. The training rows came in 8 orders. The joblib compression ran at every level, five saves each. And the default ran again at the very end of the run. In all, that is 55 fresh processes.

When two files differ, find out where. The lab compares the two files byte by byte. It reads both as pickle instructions. And it loads both models and compares them part by part. It also asks both models for their predictions on all 26,851 test rows. That way a difference in the file can be told apart from a difference in the model.

How a Fingerprint Behaves

Before the results, here is what a sha256 fingerprint does with a small change. These are two real files from the lab.

A hand-drawn sketch of two rows. Top: a box, pickle, 4 threads, 202,700 bytes, with an arrow to a box, sha256 eade9ce28dc89798 and dots. Bottom: a box, pickle, 10 threads, 202,700 bytes, with an arrow to a box, sha256 bec4d11e50f58827 and dots. Below: the two files are the same length, and 1 of their 202,700 bytes differs; of the first 16 characters of the two fingerprints, 0 sit in the same place. Caption: a fingerprint says the files differ; it never says where, or whether it matters.

Both files are pickles of the same model, 202,700 bytes each. They differ in exactly one byte. Yet their fingerprints share nothing: the first 16 characters do not match in a single place.

This is what makes a hash useful and also what makes it blunt. It is very good at saying "these two files are not the same". One changed byte out of 202,700 is enough. But it cannot tell you which byte changed, or whether the change matters. For that, you have to open the files. Most of this lesson is about opening them.

The Lab's Report, Running

This is a real recording of the report script, repro_report.py, on the laptop where the lab ran.

A terminal recording of repro_report.py. It prints that lesson 1's model has 44,521 train rows and test AP 0.5450, equal to the lab, then lab ran 2026-10-02 08:09:25 and this report 2026-10-02 09:01:43. Then a table of fresh processes with the first 12 characters of the pickle's sha256, the stored threads, and whether all 5 files equal the lab's: default run 1 and run 2, bec4d11e50f5, 10; OMP_NUM_THREADS 1, 63834fc2797a, 1; 2, 639c80564377, 2; 4, eade9ce28dc8, 4; 8, d005fa97b05c, 8; 10, bec4d11e50f5, 10; PYTHONHASHSEED 0 and 3, bec4d11e50f5, 10; threadpool_limits(4), eade9ce28dc8, 4; each line says yes, all 5 formats. Then: 4 vs 10 threads, 1 of 202,700 bytes differ, offset 2,018, BININT1 4 vs 10, after n_threads. Then the joblib gzip header: MTIME 0, OS byte 19, and decompressed equals the uncompressed file. Then row order: sorted, test AP 0.5459, and shuffle 0, 0.5470, each with 26,851 of 26,851 rows changed, the same file as the lab, and with no early stopping the same file as the stored order. Then the labelled addition after the review: with the held-out rows fixed, 5 training orders gave the same trees and predictions, and the pickles differed only in train_score_, up to 1.7e-16; copied across, the same file. Then 0 sets in the model and 0 time-like values. Then the labelled addition: Python's gzip.compress 2 seconds apart gave different files, and with mtime=0 the same. The last line says all 265 checks agree with the stored lab.

The report does not trust the lab. It never imports the lab's code. It builds lesson 1's training rows with the features chapter's own code, and it starts its own fresh processes from a small script it writes itself. In each process it trains the model again under one setting, and saves all five files with its own calls. Then every file's sha256 must equal the one the lab stored. They all did.

So the report is also a second test of the main result. A different program, started at a different time, made the same files, byte for byte. It also locates the one differing byte with its own loop, and checks the student demo's saved run against the lab.

The Headline: One File per Setting

Here is the main result, for every setting in one table.

A table of nine settings with three yes or no columns: one file, equal to the default, and same scores. Same process 5 times: yes, yes, yes. Fresh process 5 times: yes, yes, yes. OMP_NUM_THREADS 1, 2, 4 and 8: yes, no, yes, each. OMP_NUM_THREADS 10: yes, yes, yes. PYTHONHASHSEED 0 to 4: yes, yes, yes. End of the run 5 times: yes, yes, yes. Below: one file means each of the 5 formats had a single sha256 across every run of that setting; equal to the default compares the pickle with the default fresh process, which used 10 threads, all 10 cores here. Caption: only the thread count gave a different file, and never a different score.

Inside one setting, the file never changed. Five trainings in one process gave five identical files, in all five formats. Five fresh processes did too, and they matched the first five. So did five more fresh processes at the very end of the run.

The hash seed did not matter. With PYTHONHASHSEED set to 0, 1, 2, 3 and 4, every file was the default's file.

The thread count did. With 1, 2, 4 or 8 threads, each setting made its own file, the same every time, but a different file from the default. With 10 threads it was the default's file again. That makes sense on this Mac. It has 10 cores, and with nothing set, scikit-learn uses all 10.

And not one prediction moved. For every setting, the model loaded from the file gave predictions equal to the last bit on all 26,851 test rows. Test AP was 0.5450 every time.

My first run of the lab crashed after all the training was done, in a later step that compares row orders. I fixed one line and ran everything again. Before deleting the first run's files, I kept the sha256 of 25 of them. All 25 matched the second run, which started about two minutes later. That check was added after the crash, and the lab says so.

The One Byte, Found

So which byte is it? Here are the pickle bytes of the 4-thread and the 10-thread files, around the place where they differ.

Two blocks of hex bytes, one for 4 threads and one for 10 threads, each in two rows of 16 starting at offsets 2000 and 2016. The rows are the same except one marked byte, the third in the second row: 04 in the 4-thread file and 0a in the 10-thread file. Below: pickletools reads the same bytes as instructions: at 2,005, SHORT_BINUNICODE n_threads; at 2,017, BININT1 4; the marked byte is the argument of BININT1: 04 is 4, 0a is 10. Caption: everything else in 202,700 bytes was the same.

The two files differ at byte 2,018, counting from 0. In the 4-thread file that byte is 04. In the 10-thread file it is 0a, which is 10 in hex.

To see what that byte means, the lab read both files with pickletools, a tool in Python's standard library. It turns a pickle into the list of instructions that loading would follow, the same tool lesson 1 used. Just before the byte there is the text n_threads, the name of a field. Then comes an instruction called BININT1, which means "here is a small whole number", and the differing byte is that number.

Then the lab loaded both models and compared every part of them, one by one. The only difference was model._bin_mapper.n_threads: 10 in one, 4 in the other. The same was true for 1, 2 and 8 threads: one byte, at the same place, holding the thread count.

This is the field lesson 2 found by accident. The bin mapper is the part of the model that sorts each column into buckets before training. It keeps a note of how many threads it was given.

How the Thread Count Gets Into the File

Why does a model remember the number of threads it was trained with? Here is the path, step by step.

A sequence diagram with four lifelines: my code, scikit-learn, OpenMP and the file. Step 1, my code calls fit with x and y. Step 2, scikit-learn asks OpenMP how many threads. Step 3, OpenMP answers 10, its limit. Step 4, scikit-learn returns the model to my code. Step 5, my code calls pickle.dump on the file. Below: in step 4 the model's bin mapper, the part that sorts each column into buckets, keeps the answer as n_threads = 10; step 5 writes it into the file; with OMP_NUM_THREADS set, step 3 answers with that number instead. Caption: the predictions were identical for 1, 2, 4, 8 and 10 threads.

When fit starts, scikit-learn asks how many threads it may use. If OMP_NUM_THREADS is set, the answer is that number. If not, it is the smaller of OpenMP's own limit and the number of physical cores. scikit-learn's source code says so, and adds that it takes "cgroups quotas into account", the limits a tool like sets on a container. The bin mapper is created with that number and keeps it as one of its own fields. When pickle saves the model, it saves every field, so it saves this one too.

The number changes how the work is split, not what is computed. For this model, 1, 2, 4, 8 and 10 threads gave predictions equal to the last bit on all 26,851 test rows. And the lab checked this in a second way, after training. It took the default model, trained with 10 threads, and changed only that one field to 4. Saved again, it gave exactly the 4-thread file, in all five formats. So the stored number is the only difference between those files.

That was a check, not a fix. Do not do this to make your files match: it edits a private field, one whose name starts with an underscore, which scikit-learn may change without warning. Here it was harmless, because the bin mapper is used only during training, but the fixes on a later slide set the thread count the proper way.

That is a fact about this model and this version of scikit-learn, 1.9.1. A model whose sums depend on how the rows are split between threads could also differ in its numbers. This one did not.

What One Byte Does to a Compressed File

The pickle was 202,700 bytes for every thread count. Compressed, the sizes moved a little.

A bar chart of how many bytes each joblib zlib 3 file has over 91,188, by thread count: 1 thread, 2; 2 threads, 0; 4 threads, 1; 8 threads, 2; 10 threads, 2. Below: zlib 3 bytes by thread count, 1: 91,190; 2: 91,188; 4: 91,189; 8: 91,190; 10: 91,190; the pickle was 202,700 bytes for every thread count. Caption: same length uncompressed; the stored number compresses a little differently.

With zlib level 3, the file was 91,188 bytes for 2 threads, 91,189 for 4, and 91,190 for 1, 8 and 10. This is the one-byte difference lesson 2 saw: its lab ran with 4 threads and its demo with 10. Here is one possible reason, which I did not test. The compressor finds repeated patterns, and a 1 or an 8 may happen to match nearby bytes a little differently than a 2 does.

The bigger surprise was what one changed byte does to the rest of a compressed file.

Hand-drawn bars of the share of byte positions that differ between the 4-thread and 10-thread files, by format: pickle, 1 byte; joblib with no compression, 1 byte; zlib 3, 99.6%; gzip 3, 99.6%; lzma 3, 97.9%. Below: zlib 3, 90,800 of 91,190 positions; lzma 3, 73,481 of 75,021; decompressed, every compressed file equalled its uncompressed joblib file, 168 of 168. Caption: to find a difference, unpack first, then compare.

Uncompressed, the 4-thread and 10-thread files differ in 1 byte. Compressed with zlib, they differ in 90,800 of 91,190 byte positions. A compressor writes each part of its output based on what came before, so a change near the start shifts everything after it. Part of that count is also because the zlib files are not the same length, so every byte after the change sits one place later.

So if you ever compare two compressed model files and find them "completely different", do not believe it yet. Unpack both first. The lab did that for every compressed file it wrote: unpacked, each one equalled its own uncompressed joblib file, 168 of 168.

The Hash Seed Changed Nothing, and Why

PYTHONHASHSEED changes the order in which a set keeps its items. If the model held a set, its file could change from run to run. So I tested it, with a control.

Three panels. Sets in the model: 0, and 61 dicts, which keep the order they were filled in. Model file: 1 sha256, for all 5 hash seeds, the same as the default. Toy set: 5 sha256, for 5 hash seeds: the check can see a set. Below: the toy is a set of lesson 1's six column names, pickled in every process; its order, and so its bytes, moved with the hash seed; the model's file did not. Caption: a control that changes shows the test could have failed.

The lab walked through the whole trained model, every part inside every part. It found 0 sets and 61 dicts, Python's collections of named values. Since Python 3.7, a dict keeps its items in the order they were put in, so its order does not depend on the hash seed. Then the five hash seeds gave five identical model files, the same as the default's.

A result of "no change" is only worth something if the test could have shown a change. So every process also pickled a small toy: a set holding lesson 1's six column names. Its sha256 was different for each of the five seeds. Here are two of them.

A hand-drawn sketch of two boxes. Left, PYTHONHASHSEED=0: return_share, money, recency_days, products, frequency, tenure_days. Right, PYTHONHASHSEED=1: frequency, recency_days, products, money, tenure_days, return_share. Below: sha256 of the pickled set, fe262ae0e657 for seed 0 and 195a765605c2 for seed 1; a set is stored in the order it walks its items, and that order follows the hash seed. Caption: any set inside a model would make its file depend on PYTHONHASHSEED.

The six names are the same, but the order is different, so the bytes are different. If a future version of a library put a set of strings inside a model, its file would start to depend on the hash seed. This model has none.

The Row Order Changed the Model Itself

The thread count changed a note inside the file. The row order did something much bigger: it changed the model.

A bar chart of the trees kept by early stopping when the same training rows come in 8 orders: stored 54, sorted 58, reversed 51, and five shuffles s0 to s4 with 42, 59, 40, 65 and 54. Below: test AP ran from 0.5435 to 0.5490; every test row's score moved, in every order other than the stored one. Caption: each order gave the same file in both processes; a different order was a different model.

The lab trained on the same 44,521 rows in 8 orders. One was the stored order, sorted by month. One was sorted by customer, then month. One was reversed. And five were random shuffles. It ran all 8 in two separate fresh processes. Each order gave the same file in both processes, so each order is repeatable.

But each order gave a different model. The number of trees ran from 40 to 65, and every one of the 26,851 test scores moved, by up to 0.3272. Test AP ran from 0.5435 to 0.5490. I make no claim that any order is better. Each order is one draw, like a seed, and lesson 2 of the lifecycle chapter shows how much one draw can move.

An isometric row of 8 blocks, the height of each the pickle's size in kB: stored 202.7, sorted 216.8, reversed 192.1, s0 160.2, s1 220.4, s2 153.2, s3 241.5, s4 202.7, with the trees under each: 54, 58, 51, 42, 59, 40, 65 and 54. Below: s0 to s4 are the five shuffles; the first block is the stored order, 202,700 bytes; every other file was a different size, from 153,205 to 241,547 bytes, and differed from it in most of its bytes. Caption: the file mostly holds trees, so the number of trees sets its size.

The file sizes follow the trees. Lesson 1 of this chapter found the pickle is mostly the arrays of tree nodes, so more trees make a bigger file. The 65-tree model was 241,547 bytes and the 40-tree model was 153,205.

The Trees Changed Only Through the Held-Out Rows

Why would the order of the rows change the model at all? The design named two possible reasons before the run. Either the order changes which rows are held back for early stopping, or it changes the order in which numbers are added up. So the lab trained every order a second time with early stopping turned off. That second model is a check, not the chapter's model.

Two columns. Left, early stopping on, the chapter's model: 8 files; a different file for each of the 8 orders; it holds back 10% of the rows, chosen by position, so the order picks them. Right, early stopping off, a check only: 1 file; the same sha256 for all 8 orders, in both processes; every test score identical to the stored order.

Early stopping means the model holds back some training rows, trains tree after tree, and stops when the held-back rows stop improving. With default settings and more than 10,000 rows, this model holds back 10%. It picks them with a fixed random seed, and it keeps the share of buyers the same in both parts. But the seed picks positions: row 17, row 3,402, and so on. Put the rows in a different order, and the same positions hold different customers.

With early stopping off, all 8 orders gave exactly the same file, and every test score was identical. So, for this model, the trees changed only through the choice of held-back rows.

That check could not see everything, though. An independent review pointed out that with early stopping off the model keeps no training scores, so an order effect there would be invisible. I added a test after the review, and the lab labels it.

I kept the held-back rows fixed, 4,453 rows, the same ones the model would pick from the stored order. Then I put only the other 40,068 rows in 5 random orders. All 5 models kept 54 trees, with the same bins, the same tree nodes and identical predictions. But all 5 files differed from the stored-order file.

The only differing part was train_score_, the score on the training rows that the model records after each tree. It comes from an average of the loss over all the training rows, added up in row order. 12 to 18 of its 55 values moved in the last digits, by at most 1.7e-16. Copying the stored-order train_score_ into each model made the files identical. So the order of the rows can change the file even when it changes nothing about the model's predictions.

The Compression Level Is Part of the Recipe

Lesson 2 already showed that each compression setting makes a different size. Here the question is narrower: is each setting repeatable?

A bar chart of the joblib file's bytes at zlib levels 1 to 9, falling from about 96,000 to about 85,000. Below: level 1, 96,225 bytes; level 9, 85,395; also none 207,688, gzip 3 91,202, bz2 3 78,421, lzma 3 75,021; 13 settings, 13 different sha256; saved 5 times, every setting gave the same sha256 each time; every file loaded to identical scores. Caption: the level is part of the recipe, just like the protocol.

The lab saved one trained model 5 times at each of 13 settings: zlib levels 1 to 9, no compression, gzip 3, bz2 3 and lzma 3. Each setting gave the same sha256 all 5 times. The 13 settings gave 13 different fingerprints. And every file loaded back to identical predictions.

So the compressors here are deterministic: the same input and the same setting always give the same output. But the setting itself is part of the recipe. If one machine saves with compress=3 and another with compress=("lzma", 3), they will never agree on a hash, even though both hold the same model. A team that compares hashes must fix the format, the protocol and the level in code, not leave them to defaults.

Does Anything Record the Time?

Many file formats write the date and time inside the file. If a model file did that, two saves a second apart would never match.

Three panels. Pickle: 0 found, time-like values in 1,144 numbers and 125 texts. Gzip header: MTIME 0, bytes 4 to 7; the OS byte was 19. Later runs: same; the end of the run, and the report, gave the lab's files. Below: added after the results, Python's own gzip module, which joblib does not use, wrote MTIME 1790911893 and then 1790911895, a couple of seconds apart, so the two files differed; with mtime=0 they were the same. Caption: this joblib wrote no time; a plain gzip wrapper writes the time unless you say mtime=0.

The pickle holds no time. The lab read every number and every piece of text in the pickle, 1,144 numbers and 125 texts. It looked for anything that could be a time. That means a number that reads as a date between 2001 and 2033, in seconds, milliseconds or nanoseconds. It also means text shaped like a date. It found none.

Joblib's gzip writes no time. A gzip file has a field called MTIME in bytes 4 to 7 of its header. The gzip standard, RFC 1952, says "MTIME = 0 means no time stamp is available." Every gzip file joblib wrote had MTIME 0. The default ran at the start and at the end of the run, and the report ran later again; every time, the files were the lab's.

The header has one more field, the OS byte, which says what kind of system wrote the file. Here it was 19. zlib's own source code, in version 1.2.12, which this Python uses, writes 19 on Apple systems and 3 on other Unix systems such as Linux. So I expect a joblib gzip file written on Linux to differ from this Mac's in that byte. I did not measure that; it is what the source says.

Python's own gzip module is different. I checked this after the results, and it is labelled as an addition. I compressed the same pickle twice with gzip.compress, a couple of seconds apart. The two files had different MTIME values and different fingerprints. With mtime=0, they were the same. Joblib does not use this module. But if you wrap a pickle in gzip yourself, pass mtime=0.

Making the File Repeatable, Proved

So on one machine, one thing changed the file: the thread count. Here are two fixes, both tested in fresh processes.

A tall card titled both fixes made the 4-thread file. 5 of 5: OMP_NUM_THREADS=4, each with another PYTHONHASHSEED, one sha256, the 4-thread file's. 5 of 5: threadpool_limits(4) in the code, nothing in the environment, the same file again. A check, not a fix: the default 10-thread model, with its private n_threads field changed to 4 after training, saved to that same file, byte for byte. Below: sha256 eade9ce28dc89798b1d2c00b.

Fix 1: set the threads in the environment. Five fresh processes ran with OMP_NUM_THREADS=4. To make the test harder, each also had a different hash seed, from 0 to 4. All five made one file, the same as the 4-thread file, in all five formats.

Fix 2: set the threads in the code. Five fresh processes ran with nothing set. But the training line was wrapped in threadpool_limits(limits=4, user_api="openmp") from the threadpoolctl library, which scikit-learn already installs. All five made that same file again. I prefer this fix. It lives in the training code, so nobody has to remember to set a variable before running it.

Both fixes make the file repeatable by giving every machine the same number to store. They do not make a model trained with 4 threads the same file as one trained with 10. They make everyone use 4.

There is one catch, which I found in scikit-learn's source and did not test. With the environment variable set, scikit-learn uses that number even if the machine has fewer cores. Its source calls this "making it possible" to go past the number of cpus. With threadpool_limits and nothing in the environment, it still takes the smaller of the limit and the number of physical cores. So on a machine with 2 physical cores, the code fix would store 2, and the file would differ. Pick a number that every machine you train on has.

My Guesses Before the Run, Checked

I wrote eight guesses into the lab before it ran. Here they are against the results.

  1. "same: the five files in one process are byte-identical, every format." Right.

  2. "base_t0: the five fresh processes give byte-identical files, every format, and equal to same." Right.

  3. "omp_N: files differ between thread counts (n_threads is stored); omp_10 equals the baseline (10 cores here). Predictions identical for every thread count. Within one thread count, the K=5 files are identical." Right on all four parts. The first run had already shown me this table before the second run, as the lab's log says.

  4. "hash_S: byte-identical to the baseline (no set in the model). The toy set's sha256 differs between seeds." Right. All 5 toy fingerprints differed.

  5. "order: sorted, reversed and shuffled rows each give a different model: different bytes in many places, and different predictions; each order is identical to itself across the two processes." Right. The reason, held-back rows, came from the check with early stopping off.

  6. "level: each level gives different bytes; one level gives the same bytes all 5 times; every file loads to identical predictions. Decompressed, every compressed file equals the j0 file." Right.

  7. "time: no timestamp in the pickle; gzip MTIME is 0; base_t1 equals base_t0." Right.

  8. "fix_env and fix_code each give one sha256 across their 5 processes, equal to omp_4's." Right.

Eight right out of eight sounds good, but it mostly means the questions were easy to guess once lesson 2 had found the stored thread count. The useful parts are the exact places: which byte, which field, and that the row order works through early stopping.

Try It Yourself

The full lab starts 55 fresh processes. I wrote a small demo that does the core of it in five trainings.

A page in three labelled zones, headed repro_demo.py, designed before it ran. Here: train lesson 1's model, pickle it twice, compare the sha256. Three fresh processes: OMP_NUM_THREADS=1, =4 and not set: sha256, stored thread count, same scores? Locate and fix: find the differing byte; then two processes with threadpool_limits(4). Caption: it printed 1 byte differs, at 2,018, BININT1 1 vs 4; both fixed runs equal the 4-thread file: true.

I wrote the demo's design into its docstring after the lab had run and before the demo first ran. It trains the model once in its own process and pickles it twice. Then it starts three fresh copies of itself, with 1 thread, 4 threads and nothing set, and compares their files. It finds the differing byte. Last, it tries the code fix twice.

A real screenshot of VS Code with repro_demo.py open at the top of the file, showing its docstring: what it needs, how to run it, and the design written before it first ran.

Before you run this lab. You need Python 3 with pandas, pyarrow and scikit-learn: pip install pandas pyarrow scikit-learn openpyxl. scikit-learn installs threadpoolctl too. First run python fetch_data.py from the scripts/labs/features folder. It downloads the shop data once, about 46 MB, and writes one cleaned file.

Then, inside the examples folder, run python repro_demo.py. It needs no GPU. I ran it with scikit-learn 1.9.1 and Python 3.13 on a Mac with 10 cores. On your computer the fingerprints will almost surely differ from mine. Your "not set" line will store your own number of cores, and other library builds can change other bytes. What should hold is the pattern: the same fingerprint for the same thread count, and the same scores for all of them. Give it a file name, python repro_demo.py out.json, and it also saves every number. That is how was made.

Look Up a Setting, and Make Your Own Byte

This box holds the lab's real fingerprints and results. It needs nothing but Python, so it runs in your browser. It loads no model.

Press Run. Part 1 shows the 4-thread file: its fingerprint, whether it equals the default, the stored thread count and the test AP. Change CONDITION to "omp1", "hash_seed" or "shuffle0" and run again. Part 2 runs here and now. It pickles a tiny stand-in object that stores a thread count, twice, and finds the byte that differs. Then it shows gzip writing the time into a file. Change THREADS to (4, 4) and the two pickles become one.

The report script writes this box from the lab's stored results, runs it for five settings, and checks what it prints against the lab. In part 2, notice that the stand-in's differing byte sits right after the BININT1 instruction, exactly as in the real model.

The Lab's Code, Piece by Piece

The lab is one file, scripts/labs/packaging/reproducible_artifacts.py. It imports the features chapter's task and lesson 1's feature code instead of copying them, and it stops if the test AP is not 0.5450.

main builds the data once and saves the training rows. Then it calls run_child for every setting. That function starts a fresh Python process with the right environment and waits for it. When it returns, the main program reads every file the child wrote and computes the sha256 itself.

child is what runs inside each fresh process. It trains the model, once or five times. It may train inside threadpool_limits, or reorder the rows, or set the stored thread count by hand, as the setting says. Then it saves the five files. It also pickles the toy set, the positive control.

byte_diff, op_diff and obj_diff do the locating. The first compares two files byte by byte. The second reads both with pickletools and compares instruction by instruction. The third loads both models and walks through them part by part. containers counts the dicts and sets. timestamp_scan looks for anything time-like, and reads MTIME and the OS byte.

Why the Same Bytes Matter, and When They Do Not

All of this would be a curiosity if nothing used the bytes. But several common tools do.

A page in four labelled zones. Matters, verifying a rebuild: if the rebuild's sha256 equals the stored one, it is the same file; reproducible-builds.org checks exactly this way. Matters, a store keyed by hash: a registry or cache that names files by sha256 stores a new copy for every new hash, even when the scores are equal. Matters, the supply chain: a rebuild anyone can repeat shows the file came from that code and data; lesson 10 measures it. Does not matter, serving: the server needs the same predictions; here every file, whatever its bytes, gave them. Caption: decide which promise you need: the same file, or the same scores.

Checking a rebuild. Say your team keeps the sha256 of every model it ships. Months later, someone rebuilds the model from the same code and data to check it. If the files match, the check takes one line. If the thread count differs between the two machines, the check fails, and someone loses an afternoon. In this lab, that afternoon ends at one byte.

Storage that uses the hash as a name. Some stores name each file by its content's fingerprint. That way, the same file is stored only once. A retrain on a machine with another core count gives a new fingerprint, so the store keeps a second copy of what is really the same model. A cache that asks "have I built this already?" by hash will answer no.

The supply chain. Say anyone can rebuild your model and get your exact file. Then someone you do not have to trust can check the file you ship against your code and data. The Reproducible Builds project gives this as its reason: reproducible builds "provide certainty that software is genuine and has not been tampered with." Lesson 10 of this chapter asks the related question: how do you know the file you load is the file you trained?

When it does not matter. A server that answers questions needs the same predictions, not the same bytes. Every file in this lab gave identical predictions, whatever its bytes. If all you ever compare is scores, the thread count is harmless.

How to Make a Model File Repeatable

Here is how I would make a model file repeatable, using only what this lab measured.

A flowchart. Train leads to a diamond, thread count fixed? No leads to threadpool_limits(4) or OMP_NUM_THREADS=4, then to save with a fixed protocol and fixed compression; yes leads straight there. Then train again in a fresh process, then a diamond, same sha256? Yes leads to record the sha256; no leads to unpack both and find the field. Below: on another machine, compare the predictions too: the same file is the strong check, the same scores the useful one. Caption: a hash you never recompute proves nothing.

The chart is a loop. Fix what you can, train again in a fresh process, and compare fingerprints. If they differ, unpack both files and find the field, the way this lab found n_threads. Here are the steps in words.

  1. Fix the threads in the training code. Wrap fit in threadpool_limits(limits=4, user_api="openmp"), or set OMP_NUM_THREADS. Pick a thread count that every machine you train on has as cores, because the code fix stores fewer threads on a smaller machine.

  2. Fix the recipe in code. Write the pickle protocol, the compression and its level. Lesson 2 showed that the default protocol is 4 on Python 3.13 and 5 from Python 3.14. Here, every compression setting made its own file.

  3. Fix the data and its order. Read the training rows in a fixed order. In this lab the order changed the model itself, through early stopping. And even with the held-out rows fixed, it changed the file, through the training scores the model records, without changing the model.

  4. Train twice, in two fresh processes, and compare the sha256. This is the test. If they differ, unpack both files and compare them part by part, as the lab's does. The difference will name itself.

When to Chase the Same Bytes, and When Not To

Chase byte-identical files when something compares hashes. One example is a registry or a cache that names files by their fingerprint. Others are an audit that rebuilds your model to check it, or a supply chain where others must verify your file. Here, fixing the thread count was enough on one machine, and it cost one line of code.

Do not chase them when all you need is the same predictions. Every file here, whatever its bytes, gave identical scores. Spending a day on one stored byte helps nobody if nothing compares the bytes.

Do not expect the same file across machines without testing. This lab ran on one Mac. Another operating system writes a different gzip OS byte, by zlib's own source. Other library builds may change other bytes. Lesson 5 found Linux and macOS differed in the last digit of 33 predictions. Compare predictions with a tolerance, and treat a matching file hash as a bonus.

Do not trust a compressed-file diff. Unpack first. One changed byte became 90,800 different positions after zlib.

Do not leave the row order to chance if your model uses early stopping or any random split by position. Here a different order was a different model, with up to 0.3272 difference in one customer's score.

What This Lab Cannot Tell You

Two columns titled shows and cannot show. Shows: one machine, the same file every time, whatever the hash seed or the time; the thread count is the one field that moved, by one byte; row order changed the model through early stopping, and the file through train_score_. Cannot show: another computer, other cores, other libraries, other wheels; other models, a different class may store other fields; other formats, skops, ONNX, MLflow wrappers.

One machine. Every run was on one Mac with 10 cores, scikit-learn 1.9.1, numpy 2.5.3, joblib 1.6.0 and Python 3.13.15. I did not build the same file on Linux or Windows. The gzip OS byte is a difference I expect from zlib's source, not one I measured.

One model. HistGradientBoostingClassifier stores its thread count. Other scikit-learn models, or models from other libraries, may store other things that change between runs, or nothing at all. The method carries over: train twice, compare, and unpack to find the field.

One morning. The first run started at 08:07 and the recorded report ran at 09:01, all on 2 October 2026. The scan found no time-like value in the file, which is stronger evidence than the gap. But I did not save a file on one day and again a year later.

Five formats. Pickle and joblib only. Lesson 2 found that zipped skops files moved by tens to a few hundred bytes from run to run. I did not repeat that here, and I did not test MLflow's or ONNX's files.

Labelled additions. Two things were added after the results, and both are labelled in the lab, the report and here. One is the check against the first, crashed run's 25 files. The other is Python's gzip module with and without mtime=0. After an independent review I added a third: the held-back rows fixed and only the training rows reordered, which found the train_score_ difference.

What to Do on Monday

A hand-drawn grid of six cards, titled five steps. 1, fix the threads: threadpool_limits around fit, or OMP_NUM_THREADS. 2, fix the recipe: the pickle protocol, the compression and its level, in code. 3, keep the row order: sort or seed it, and write it down. 4, train twice: in two fresh processes, and compare the sha256. 5, store both fingerprints: the file's sha256, and a hash of the predictions. The reason: one stored number, n_threads, gave 5 different files for the same model here. Caption: the same scores do not promise the same file.

If you take one thing to work on Monday, run your training script twice, in two fresh terminals, and compare the sha256 of the two model files. On a Mac, shasum -a 256 model.pkl prints it. On any computer with Python, this line does the same: python -c "import hashlib,sys;print(hashlib.sha256(open(sys.argv[1],'rb').read()).hexdigest())" model.pkl. I ran both on one of the lab's files and they printed the same 64 characters. If the two match, write the fingerprint down next to the model. If they do not, load both files and compare them part by part until the difference names itself.

Then look for anything in your pipeline that compares model files by hash: a registry, a cache, a deployment check. That is where a stored thread count can quietly cost you.

A closing card titled same scores, not always the same file. Three numbers in large type: 1 byte, of 202,700 changed with the thread count, and changed the sha256; 0 scores moved with it, on 26,851 test rows, for 1, 2, 4, 8 or 10 threads; 5 of 5 fresh processes made one file once the threads were fixed in code.

The one idea to keep: the same scores do not promise the same file. Here, on one machine, the file depended on exactly one thing, the thread count the model stored. That changed one byte and never a score. The row order went further and changed the model itself. Fix the threads in code, fix the recipe, train twice, and let the sha256 tell you whether you did it.

Knowledge Check

Knowledge Check

4 questions - Score 80% to pass

Q1

The 4-thread and 10-thread pickles of the same model had different sha256 values. What did the lab find was different?

Q2

Changing PYTHONHASHSEED from 0 to 4 did not change the model file. Why could the lab trust that result?

Q3

Training on the same rows in a shuffled order gave a different model. What did the check with early stopping off show?

Q4

Which change made five fresh processes, with nothing set in the environment, save the same file as the OMP_NUM_THREADS=4 processes?

OpenMP

Python says that when PYTHONHASHSEED is not set, "a random value is used to seed the hashes of str and bytes objects". It also says: "Changing hash values affects the iteration order of sets." A set is a Python collection with no fixed order. And its gzip module writes the current time into a file unless you pass mtime=0. Each of these is something that could change a file's bytes. So each one is a thing to test.

I read scikit-learn 1.9.1's code to check how it picks the held-back rows. It splits them off with train_test_split, stratified by label, with a fixed random_state, and that works on positions. The early-stopping result agrees with that.

results/repro-demo.json
"""Train the same model in fresh processes and compare the FILES, not just the scores.

Lesson 6 of 'Packaging, Registry and Versioning'. It needs Python 3 with
pandas, pyarrow, scikit-learn and threadpoolctl (scikit-learn installs it),
and the shop data from the features chapter: run
scripts/labs/features/fetch_data.py once first. Then, inside this folder:
    python repro_demo.py            # print the results
    python repro_demo.py out.json   # and save them
It prints no timings. It starts a few fresh Python processes of itself.

Design, written 2026-10-02 after the lab (reproducible_artifacts.py) had
run and before this file first ran:
  1. Train lesson 1's model here (it must score test AP 0.5450) and save it
     with pickle protocol 5 twice: are the two files' sha256 equal?
  2. Train it again in three FRESH processes: OMP_NUM_THREADS=1, =4, and
     not set. Each saves its pickle. Print each file's sha256 (first 12
     characters), the thread count stored inside the model, and whether its
     predictions on all 26,851 test rows equal step 1's to the last bit.
  3. Locate the difference between the 1-thread and the 4-thread file:
     how many bytes differ, where, and which pickle instruction is there.
  4. The fix: two more fresh processes, thread count not set, but the
     training wrapped in threadpool_limits(4). Are both files the same as
     the OMP_NUM_THREADS=4 file?
  The sha256 values depend on your machine (step 2's "not set" uses all
  your cores); repro_report.py checks this run against the lab.

Author: Roni Das
Created: 2026-10-02
"""
import hashlib
import json
import os
import pickle
import pickletools
import subprocess
import sys
import tempfile
from pathlib import Path

import numpy as np

HERE = Path(__file__).resolve()


def sha(data):
    return hashlib.sha256(data).hexdigest()


def child(folder, out_name, limit):
    """Runs in a fresh process: train from the saved arrays, save a pickle."""
    from sklearn.ensemble import HistGradientBoostingClassifier as HGB
    from threadpoolctl import threadpool_limits
    folder = Path(folder)
    x, y = np.load(folder / "x_tr.npy"), np.load(folder / "y_tr.npy")
    model = HGB(random_state=0)
    if limit:
        with threadpool_limits(limits=int(limit), user_api="openmp"):
            model.fit(x, y)
    else:
        model.fit(x, y)
    with open(folder / out_name, "wb") as f:
        pickle.dump(model, f, protocol=5)


if len(sys.argv) > 1 and sys.argv[1] == "--child":
    child(*sys.argv[2:5])
    sys.exit()

from sklearn.ensemble import HistGradientBoostingClassifier as HGB  # noqa: E402
from sklearn.metrics import average_precision_score  # noqa: E402

sys.path.insert(0, str(HERE.parents[2] / "features"))
import task  # noqa: E402
from what_a_feature_is import HAND_COLS as COLS, joined  # noqa: E402

ev = task.load_events()
lab_tr, _, lab_te = task.splits(ev)
tr = joined(ev, lab_tr, task.TRAIN_CUTOFFS)
te = joined(ev, lab_te, task.TEST_CUTOFFS)
x_tr, y_tr, x_te = tr[COLS].to_numpy(float), tr["label"].to_numpy(), te[COLS].to_numpy(float)
y, cut = te["label"].to_numpy(), te["cutoff"].to_numpy()


def test_ap(p):
    """Average precision in each test month, then the plain mean."""
    return float(np.mean([average_precision_score(y[cut == c], p[cut == c]) for c in np.unique(cut)]))


out = {}
with tempfile.TemporaryDirectory() as tmp:
    tmp = Path(tmp)
    np.save(tmp / "x_tr.npy", x_tr)
    np.save(tmp / "y_tr.npy", y_tr)

    # 1. here, twice
    model = HGB(random_state=0).fit(x_tr, y_tr)
    here = model.predict_proba(x_te)
    a, b = pickle.dumps(model, protocol=5), pickle.dumps(model, protocol=5)
    out["ap"] = test_ap(here[:, 1])
    out["here_same"] = sha(a) == sha(b)
    print(f"test AP {out['ap']:.4f}; saved twice in this process: same sha256: {out['here_same']}")

    # 2. fresh processes
    def fresh(name, env_threads=None, limit=""):
        env = {k: v for k, v in os.environ.items() if k != "OMP_NUM_THREADS"}
        if env_threads:
            env["OMP_NUM_THREADS"] = env_threads
        subprocess.run([sys.executable, str(HERE), "--child", str(tmp), name, limit], env=env, check=True)
        data = (tmp / name).read_bytes()
        m = pickle.loads(data)
        return data, int(m._bin_mapper.n_threads), bool(np.array_equal(m.predict_proba(x_te), here))

    files = {}
    print(f"\n{'fresh process':26s} {'sha256':12s} {'stored':>6s} {'same scores':>11s}")
    for label, name, env_threads in [("OMP_NUM_THREADS=1", "t1.pkl", "1"), ("OMP_NUM_THREADS=4", "t4.pkl", "4"),
                                     ("not set", "tx.pkl", None)]:
        data, nt, same = fresh(name, env_threads)
        files[name] = data
        out[label] = {"sha": sha(data), "n_threads": nt, "same_scores": same, "bytes": len(data)}
        print(f"{label:26s} {sha(data)[:12]} {nt:6d} {str(same):>11s}")

    # 3. where do the 1-thread and 4-thread files differ?
    f1, f4 = files["t1.pkl"], files["t4.pkl"]
    pos = [i for i in range(min(len(f1), len(f4))) if f1[i] != f4[i]]
    ops = list(pickletools.genops(f4))
    at = [k for k, (op, arg, p) in enumerate(ops) if p is not None and p < pos[0] + 1][-1]
    key = next(ops[j][1] for j in range(at - 1, 0, -1) if "UNICODE" in ops[j][0].name)
    out["diff"] = {"bytes_differ": len(pos), "offset": pos[0], "op": ops[at][0].name,
                   "arg_1": pickle.loads(f1)._bin_mapper.n_threads, "arg_4": ops[at][1], "key": key}
    print(f"\n1 vs 4 threads: {len(pos)} of {len(f4):,} bytes differ, at byte {pos[0]:,}: "
          f"{ops[at][0].name} {out['diff']['arg_1']} vs {ops[at][1]}, just after the key '{key}'")

    # 4. the fix: limit the threads in code
    shas = []
    for i in range(2):
        data, nt, same = fresh(f"fix{i}.pkl", None, "4")
        shas.append(sha(data))
        print(f"threadpool_limits(4), run {i + 1}:  {sha(data)[:12]} {nt:6d} {str(same):>11s}")
    out["fix"] = {"shas": shas, "equal_omp4": all(s == out["OMP_NUM_THREADS=4"]["sha"] for s in shas)}
    print(f"both equal the OMP_NUM_THREADS=4 file: {out['fix']['equal_omp4']}")

if len(sys.argv) > 1:
    json.dump(out, open(sys.argv[1], "w"), indent=1)

This is a real run in VS Code's terminal: python repro_demo.py, run inside the examples folder.

A real screenshot of VS Code's terminal after running python repro_demo.py inside the examples folder. It prints test AP 0.5450 and that saving twice in this process gave the same sha256: True. Then three fresh processes with sha256, stored threads and same scores: OMP_NUM_THREADS=1, 63834fc2797a, 1, True; OMP_NUM_THREADS=4, eade9ce28dc8, 4, True; not set, bec4d11e50f5, 10, True. Then 1 vs 4 threads: 1 of 202,700 bytes differ, at byte 2,018, BININT1 1 vs 4, just after the key n_threads. Then threadpool_limits(4), runs 1 and 2, both eade9ce28dc8, 4, True, and both equal the OMP_NUM_THREADS=4 file: True.

When I ran it, all three fresh-process fingerprints were exactly the lab's files for 1 thread, 4 threads and the default. It found the same byte, 2,018, and both fixed runs made the 4-thread file. The report script checks the demo's saved run against the lab.

gzip_header

repro_report.py checks all of it from the outside, with its own training code, as the recording showed.

obj_diff
  • Store two fingerprints. Keep the file's sha256, and also a hash of the model's predictions on a fixed set of rows. The lifecycle chapter's lineage lesson calls this the answers fingerprint. On another machine, the file's hash may differ for harmless reasons, and the answers' hash tells you whether it matters.

  • Keep the versions next to the file. A different scikit-learn trains a different model, as lesson 3 of this chapter showed. No fix here survives a version change.