Packaging And Registry

Container Images: Most of the Size Was the Base and the Libraries, and Slimming Kept Every Score

0 of 28 complete

0%

Contents

Back|Packaging And RegistryContainer Images: Most of the Size Was the Base and the Libraries, and Slimming Kept Every Score
1/28
65 min left
Prerequisites
Pinning Dependencies: The Same Requirements File Installed Something Else Every SeasonrequiredLibrary Version Skew: Old Model Files Broke on Upgrade, and Retraining Made a Different ModelrequiredWhat a Feature Is: A Better Model or a Better Feature?required
Related Topics
Same Code, Different Model: What Changes Between Two Identical Training RunsThe ML & AI LifecycleLineage and Rollback: Rebuilding Last Month's Model Exactly, and What Breaks When You CannotThe ML & AI LifecycleThe Lifecycle as One Loop: Nine Steps, Nine Measured Failures, and What Saw Each OneThe ML & AI LifecycleWrong Labels: How Many Can a Model Survive, and Can You Find Them?Data Engineering for MLRebalance, or Just Move the Threshold? Measured on Rare ClassesData Engineering for ML
1 of 28

Packing for a Trip

Let me start with a trip, like the one in the picture.

Imagine you are going away for a week. You could put your whole home in a truck: the bed, the kitchen, the tools in the garage, every book. You would certainly have everything you need. But the truck is slow and costly to move, and almost none of it will be used.

A flat illustration of a woman standing in a light room with a large tote bag over one shoulder. Below the picture: a container image is the bag a model travels in. A smaller bag is lighter to carry, until the one thing you need is not in it.

So you pack one bag instead. A small bag is quick to carry. But a small bag has a risk too. If you forget the charger for your phone, the trip goes badly, and a lighter bag does not fix that.

A trained model also has to travel: from the laptop where it was made, to a server where it answers questions. In this lesson I pack the same model into bags of different sizes, weigh each bag, and check that the model inside still gives exactly the same answers.

Where This Lesson Starts

This lesson uses the model from the features chapter. If it is new to you, please read what a feature is first. That lesson built six features for each customer of a real online shop. It asked one question: at the start of a month, will this customer buy something in the next 30 days? The model is a gradient boosted tree model from scikit-learn, and it scored a test AP of 0.5450. AP, average precision, is a score from 0 to 1 for how well the model ranks the buyers above the others.

Two lessons of this chapter come just before this one. Library version skew showed that a saved model often fails to load when the library versions change. Pinning dependencies showed how to write down the exact versions. Here I take the next step: put the model, the exact libraries and Python itself into one sealed package.

The survey lesson on model packaging and containerization explains Dockerfiles in words, with example sizes. I will not repeat it. Here I build real images and measure them.

The Words You Need First

Please read this slide slowly if any word is new. Every slide after it uses these words.

A hand-drawn grid of nine cards, three per row. Image: a packed, read-only copy of all the files a program needs: Linux, Python, libraries, model, code. Container: one running copy of an image. Dockerfile: the list of steps that builds an image. Layer: the files one step added; an image is a stack of them. Base image: the image a Dockerfile starts FROM. Build cache: layers Docker keeps and reuses when a step has not changed. Full, slim, alpine: three sizes of the official Python base image. Multi-stage: build in one image, copy only the result into another. Wheel: a package that arrives ready-built, so pip compiles nothing. Below: model, test AP and the six columns mean what they meant in the features chapter.

is a tool that packs a program with everything it needs. A container image, or just image, is that package. It is a read-only copy of all the files the program needs, from a small Linux system up to my own code. A container is one running copy of an image, like one cake baked from a recipe.

A Dockerfile is the recipe. It is a short text file of steps, such as "start from this image", "copy this file in", "run pip install". Each step adds a layer, the set of files that step added or changed. An image is a stack of layers.

The first step names a base image, the image my Dockerfile starts FROM. The Python project publishes official base images in several sizes. The default one I call full. Slim leaves out many system packages. Alpine is built on a much smaller Linux, called Alpine Linux.

When Docker builds an image again, it can reuse a layer it built before, if nothing that layer depends on has changed. The saved layers are the build cache.

A wheel is a Python package that arrives ready-built. When pip finds a wheel for your system, it copies files and compiles nothing. When it does not, it downloads the source and tries to compile it, which needs a compiler, a program that turns C code into machine code.

What the Documentation Says

Before building anything, I read what the tools themselves say. Every quote here was found word for word in the source file at a fixed commit, and all of them are in results/img-factcheck.json.

Five rows, each with a logo. Docker docs, multi-stage builds: you can selectively copy artifacts from one stage to another, leaving behind everything you don't want in the final image. Docker docs, the build cache: if a layer changes, all other layers that come after it are also affected. The python image, slim variant: this image does not contain the Debian packages required to compile extension modules written in other languages. The python image, alpine variant: it uses musl libc instead of glibc, so software will often run into issues depending on the depth of their libc requirements/assumptions. pip docs, caching: pip provides an on-by-default caching; pip's caching behaviour is disabled by passing the --no-cache-dir option. Caption: every quote was found in the source file, at the commit named in the fact-check.

's page on multi-stage builds says: "You can selectively copy artifacts from one stage to another, leaving behind everything you don't want in the final image." Its page on the build cache says: "If a layer changes, all other layers that come after it are also affected."

The python image's own description is careful about slim. It says the slim image "does not contain the Debian packages required to compile extension modules written in other languages". It also says "pip install may fail when installing a Python distribution package from a source distribution". Debian is the Linux system inside the full and slim images.

About alpine, it says the image uses musl libc instead of glibc. A libc is the basic C library that almost every program on Linux calls for things like memory and maths. glibc is the common one; musl is a smaller one. The description warns that "software will often run into issues depending on the depth of their libc requirements/assumptions."

And pip's documentation says pip keeps a cache of downloads by default, and that --no-cache-dir turns it off. Each of these is a promise I can test.

How the Lab Was Built

I wrote the design into the docstring of scripts/labs/packaging/container_images.py before it first ran. Before that, I had pulled the three base images and looked at their sizes once, and the docstring says so. I had not built any image with this model.

A page in four labelled zones. One model, made once: lesson 1's model, trained on the Mac, test AP 0.5450; saved with pickle; the 26,851 test rows kept apart, never put in an image. Five Dockerfiles: a full base, b slim, c slim with a hash lock and no pip cache, d multi-stage, e alpine; every base pinned by its digest. Measured: size three ways and layer by layer; what is inside; the predictions copied out and compared with the Mac's, to the bit. Rebuilt: after no change, a code change and a requirements change: which steps ran again. Caption: guesses, written first: the full base is most of image a; every image predicts exactly what the Mac does. Title: one model, five Dockerfiles, 21 builds.

One model, made once. My main Python on the Mac trains lesson 1's model and saves it with pickle, Python's built-in way to save an object as a file. The lab stops if its test AP is not 0.5450. The Mac also predicts the 26,851 test rows, and those predictions are the reference.

A tiny program to run inside. Every image holds the same short script, predict.py. It loads the model, reads the test rows, predicts, and writes the predictions to a file. The test rows are never baked into an image. The lab copies them in when the container starts, and copies the predictions out when it stops.

Five Dockerfiles. They differ only in how the libraries get in.

Five rows, one per Dockerfile, each with its base image, its install line and a note. a: python:3.13; pip install -r requirements.txt; unpinned, numpy and scikit-learn. b: python:3.13-slim; the same install; the same as a. c: python:3.13-slim; pip install --no-cache-dir --require-hashes -r requirements.lock; exact versions, each with its hash. d: python:3.13, then slim; the same install into /venv, then COPY --from=build /venv /venv; multi-stage. e: python:3.13-alpine; the same as c; musl libc. Below: then all five copy model.pkl and predict.py; each base is pinned by its sha256 digest. Caption: the test rows go in with docker cp when the container starts.

Images a and b install from a file that just says "numpy" and "scikit-learn", with no versions. Images c, d and e install from a , the kind lesson 4 built. It lists every package at the exact version my Mac used, with a of each file, a fingerprint pip checks before installing. The flag tells pip not to keep its download cache.

The Lab's Report, Running

This is a real recording of the report script, img_report.py, on the laptop where the lab ran.

A terminal recording of img_report.py. It prints that the model, the rows and the Mac predictions match their sha256, Mac test AP 0.5450 equal to lesson 1's, and that the lock equals the main venv: numpy 2.5.3, scikit-learn 1.9.1, scipy 1.18.1. A table of four images: a full, inspect 544.9, unpacked 1569.1, added 363.7, pip cache 59.7; b slim, 186.2, 521.9, 363.7, 59.7; c slim, 126.6, 460.8, 302.6, 0.0; d slim, 126.9, 464.2, 306.0, 0.0; each with same bits False, 33 rows, max diff 1.1e-16, AP 0.5450. Then e alpine did not build: pip got scikit_learn-1.9.1.tar.gz, and no C compiler. Then the layer cache, which steps ran for each image and each rebuild. Builds quiet False, highest load seen 24.01, no timings. Then the labelled post-results lines: the four images agree with each other to the bit; raw tree sums equal on all rows; expit of the same Mac raw scores differs on 33 rows; AP equal per month, 0 rows moved rank, 0 decisions flipped; scikit-learn has 0 musl files for aarch64, scipy 1, numpy 1; build tools in the full base: gcc, g++, make, git, curl, none in slim. Last line: all 97 checks agree with the stored lab.

The report needs no . The lab stored every raw output it read: each build log, each layer list, each image's manifest, each du listing and every container's predictions. The report checks the sha256 of every saved file. It recomputes every AP with its own loop and compares every prediction with its own numpy calls. It adds up every layer again, and it reads every build log with its own parser, written apart from the lab's. All 97 checks agreed.

One line matters for what follows: "builds quiet: False; highest load seen 24.01". The load average is how many programs were waiting to run on the laptop. Other work was running at the same time, so build times would mean nothing. This lesson gives no timings, only which steps ran.

The Headline: Slim Cut the Most

Here are the four images that built, each measured two ways.

A bar chart of the four images that built, two bars each, in MB. Download, from docker image inspect, and on disk, unpacked: a full, 544.9 and 1,569.1; b slim, 186.2 and 521.9; c lock, 126.6 and 460.8; d multi, 126.9 and 464.2. Below: alpine, image e, did not build. Caption: going from the full base to slim saved more than every later move put together.

Image a, the full base, was 1,569.1 MB on disk. Image b, the same Dockerfile on slim, was 521.9 MB. That one word in the first line saved 1,047.2 MB.

The other moves saved much less. Turning off pip's cache and using the lock, image c, took b from 521.9 to 460.8 MB. The multi-stage build, image d, was 464.2 MB, a little bigger than c, not smaller.

Alpine did not build at all, so it has no bar. I explain why on its own slide.

The download sizes tell the same story, at 27 to 36% of the unpacked size: 544.9 MB for a, and 126.6 MB for c. The next slide explains why each image has two sizes.

Which Size Do You Mean?

Before going on, I have to be honest about a confusing thing. Different tools gave different sizes for the same image.

Three panels for image c. Inspect: 126.6 MB; docker image inspect, Size: compressed, what a pull downloads. History: 460.8 MB; docker history, every layer added up: unpacked, on disk. Du: 445.7 MB; du inside the container: the files it sees. Below: for every image, inspect's Size equalled the compressed layers in the image's manifest plus a few kB of metadata; this Docker uses the containerd image store, Docker Desktop's default. Caption: when someone says an image is 126 MB, ask: downloaded, or on disk?

docker image inspect reported 126.6 MB for image c. Adding up every layer in docker history gave 460.8 MB. And du, a Linux tool that adds up file sizes, run inside a container, said 445.7 MB.

I checked what the first number is. A manifest is the small file that lists an image's layers. The lab saved each image with docker save and read its manifest. For every image, inspect's Size equalled the sum of the compressed layers in that manifest, plus a few kB of other metadata. Compressed means squeezed, like a zip file. So 126.6 MB is roughly what a server downloads when it pulls the image, and 460.8 MB is what the unpacked files take on its disk.

On disk, a server with this kind of pays for both. Docker's page on the containerd image store puts it this way.

"When you pull an image, containerd keeps the compressed layers (as received from the registry) and also extracts them to disk."

So image c costs about 126.6 + 460.8 MB on a server. I saw the same thing on the slim base: docker image ls said 204MB, and its compressed 43.3 MB plus unpacked 158.2 MB make 201.5 MB.

A pull also skips layers the server already has. In Docker's words, when the layers are already there, it "doesn't need to pull them again". So a server that already holds the slim base downloads only my layers, not the whole 126.6 MB. I did not measure a pull.

What the Size Is Made Of

Now the question of the lesson. Where do those megabytes come from?

An isometric drawing of four stacks, one per image, each split into a lower block, the base image, and an upper block, what my Dockerfile added. a: 1,569.1 MB, base 1,205.4, added 363.7. b: 521.9 MB, base 158.2, added 363.7. c: 460.8 MB, base 158.2, added 302.6. d: 464.2 MB, base 158.2, added 306.0. Below: almost all of what was added is pip install; the model is 0.2 MB and the code 12.3 kB, too thin to see. Caption: in image a the base is 77% of everything; in c it is 34%.

Every image is two parts: the base image, which I did not build, and the layers my Dockerfile added on top.

In image a, the base was 1,205.4 MB of the 1,569.1 MB, or 77%. Everything I added was 363.7 MB. So the biggest part of the full image was not my model and not my libraries. It was the starting point.

In images b, c and d the base was the slim one, 158.2 MB. There, what I added was the bigger part: 302.6 MB in c, about two thirds of the image.

The model itself was tiny. Its layer was 0.2 MB. My code's layer was 12.3 kB. These are the two things the image exists to carry, and together they were well under a thousandth of image a.

So there are two weights to look at: the base, and the libraries. The next two slides open each one.

What Is in the Full Base

What fills 1,205.4 MB in the full base, when slim needs 158.2 MB?

Hand-drawn bars of how many MB more the full base holds than slim, by folder: /usr/lib plus 490 MB, /usr/share plus 237 MB, /usr/libexec plus 92 MB, /usr/bin plus 92 MB, /usr/include plus 58 MB. Below: in the full base only: gcc, g++, make, git, curl; /usr/libexec/gcc 92 MB, the compiler; /usr/include 58 MB, C headers; /usr/share/locale 95 MB; the slim base has none of the five tools. Caption: build tools and development libraries, which a model that only predicts never uses.

The lab ran du inside both base images, two folder levels deep, and compared them. Five system folders made up most of the difference. /usr/lib, where shared libraries live, held 490 MB more. /usr/share held 237 MB more.

Then, after seeing those folders, I asked one more labelled question: which tools are there? The full base has gcc and g++, two compilers, plus make, git and curl. The slim base has none of the five. The compiler's own folder, /usr/libexec/gcc, is 92 MB. The C header files in /usr/include, which a compiler reads, are 58 MB. Translations of program messages into other languages, in /usr/share/locale, are 95 MB.

The python image's description says the default image "contains the most commonly required Debian packages", so that most pip install commands work, even ones that have to compile. So the extra is build tools and development libraries. That is useful on a laptop where you build things. But my model only loads and predicts, and it never calls a compiler.

What Is in the Libraries

Now the second weight. What did pip install?

Bars with logos of the biggest folders in image c's site-packages, 297.3 MB in all: scipy 119.4 MB, sklearn 57.6 MB, numpy 39.9 MB, scipy.libs 31.1 MB, numpy.libs 28.1 MB, pip 11.4 MB, narwhals 5.3 MB. Below: scipy.libs and numpy.libs hold the compiled maths libraries the two packages bring with them; scikit-learn needs scipy, so it cannot be left out; pip came with the base image, not from my install. Caption: the libraries are the weight; the model file is 0.2 MB.

site-packages is the folder where pip puts installed packages. Inside image c it held 297.3 MB.

The biggest was scipy, at 119.4 MB, plus 31.1 MB more in scipy.libs. scipy is a library of scientific maths. My predict script never imports it by name. But scikit-learn lists scipy as a requirement, so pip installs it, and it cannot simply be left out. scikit-learn itself, in the folder sklearn, was 57.6 MB. numpy was 39.9 MB, plus 28.1 MB in numpy.libs.

The .libs folders hold compiled maths code that the wheels bring with them, so the packages do not need those libraries installed on the system. That is part of why a wheel installs on slim with no compiler: everything it needs is inside it.

So for this model, the weight I chose is mostly scipy and numpy. A model that used a deep learning library would carry far more, but I did not measure that here.

The Cache That Travels With the Image

Images b and c both use the slim base. Why was c 61.1 MB smaller?

A hand-drawn sketch of two boxes. b, pip install layer: 363.5 MB, pip's cache 59.7 MB. c, --no-cache-dir: 302.4 MB, pip's cache 0.0 MB. Below: the cache sits in /root/.cache/pip, inside the layer; the two layers differ by 61.1 MB; pip's cache is on by default, and the python image does not turn it off. Caption: every copy of image b carries that cache, and predicting never reads it.

When pip installs a package, it keeps a copy of what it downloaded, in a cache folder, so the next install on the same computer is faster. Inside image b that folder, /root/.cache/pip, held 59.7 MB. It was written during the pip install step, so it became part of that step's layer, and part of the image.

In image c, --no-cache-dir kept pip from writing it. The pip layer was 302.4 MB instead of 363.5 MB, a difference of 61.1 MB, very close to the cache's size.

My guess before the run was that b minus c would be about the size of pip's cache. It was.

The cache helps nobody here. A container that only predicts never runs pip, and every server that pulls image b downloads and stores those megabytes anyway. In a Dockerfile, --no-cache-dir is close to free.

The Multi-Stage Build Saved Nothing Here

The survey lesson calls a multi-stage build "pure profit". So I built one.

A sequence diagram with three lifelines: full image, /venv and slim image. Step 1, the full image makes a venv. Step 2, pip install into it. Step 3, COPY --from, the venv into the slim image. Step 4, the slim image's PATH points to /venv. Below: only /venv crosses to the slim image; the full image's compilers stay behind; then the model and code are copied in, as in c. Image d: 464.2 MB unpacked, against 460.8 MB for c, which installs straight into slim: 3.3 MB more. Caption: every package came as a wheel, so there was nothing compiled to leave behind.

A multi-stage Dockerfile has two FROM lines. The first stage, called build, starts from the full image and installs the libraries into a venv, a folder that holds its own Python packages. The second stage starts from slim and copies only that folder across. Everything else in the first stage, the compilers and headers included, is left behind.

Image d was 464.2 MB, and image c was 460.8 MB. So the multi-stage image was 3.3 MB bigger, not smaller. Both had the same packages at the same versions, as pip freeze inside each one showed.

My guess was that d would be within 5% of c, and it was within 1%. The reason I gave before the run is the likely one. Every package here arrived as a wheel, so nothing was compiled, and there was nothing to leave behind. Image c never had the compilers either, because slim does not have them. The two pips are not the same pip.

In c, the 11.4 MB pip in site-packages is the slim base's own, in /usr/local. Image d still has that base pip, 6.8 MB, untouched, and carries a second pip of 12.6 MB inside /venv. Measured with du, that is 8.0 MB more pip in d than in c, yet d was only 3.8 MB bigger by du. So something else made d about 4.2 MB smaller, and I did not trace what.

A multi-stage build earns its keep when a package must be compiled from source. Here none had to be.

Rebuilding: Which Steps Ran Again

A real model image is built again and again: when the code changes, when the model is retrained, when a library is upgraded. The build cache decides how much work each rebuild does.

A grid of five images, a, b, c, d and c2, against three changes: no change, code change, requirements change; each cell lists the steps that ran again. No change: nothing, for every image. Code change: code for a, b, c and d; copy everything and pip install for c2. Requirements change: requirements, pip install, model and code for a, b and c; requirements, pip install, copy the venv, workdir, model and code for d; copy everything and pip install for c2. Below: c2 is c with COPY . . before pip install, so the code is copied in before the libraries; for d, the requirements change re-ran the install in the build stage and every step after it in the final stage. Caption: in c2, changing one line of code installed every library again.

The lab built each image a first time with --no-cache, then three more times, and read every step's line in the build log: CACHED, or run again.

No change: nothing ran. Every step came from the cache.

A code change, one extra line in predict.py: only the last step ran in a, b, c and d. That step copies predict.py, a 12.3 kB layer.

A requirements change: pip install and every step after it ran. I added one comment line to the requirements file. The installed packages did not change. But compares a file's bytes, not its meaning, so a comment counts as a change. This is the build-cache rule from Docker's page, seen for real: once one layer changes, every layer after it is built again.

And c2. I added it only for this test. It is image c with one difference: it copies the whole folder in with COPY . . before pip install, a very common way to write a first Dockerfile. In c2, the code change ran pip install again, every library, for a one-line edit.

Order the Lines by How Often They Change

Here is the difference between c and c2, line by line.

Two columns of hand-drawn boxes, one per Dockerfile line, with the lines that ran again after a code change marked. c, libraries first: FROM python:3.13-slim, WORKDIR /app, COPY requirements.lock, RUN pip install, COPY model.pkl, COPY predict.py, marked ran. c2, everything first: FROM python:3.13-slim, WORKDIR /app, COPY . ., marked ran, RUN pip install, marked ran. Below: Docker reuses a step only if every step above it was reused; so the line that changes most, the code, goes last, and the libraries, which change least, go first. Caption: the same files go in; only the order of the lines differs.

Think of the cache like a stack of plates. If you change a plate in the middle, you have to lift off every plate above it and put them back. works the same way, from the top of the Dockerfile down.

In c, the lines go from what changes least to what changes most. The base comes first, then the requirements and the install, then the model, then the code. A code change touches only the last plate.

In c2, COPY . . puts the code and the requirements in one step, near the top. Any change to any file in the folder changes that step, and so pip install below it must run again.

The two Dockerfiles put the same files into the image. Only the order of the lines differs. On a busy team that rebuilds many times a day, that order decides whether a rebuild copies one small file or installs every library from the internet again.

The Surprise: Not Identical, in the Last Digit

Now the cost side. Did slimming change the model's answers? I guessed that every image would give exactly the Mac's predictions. I was wrong, but not in the way I worried about.

A tall card. 33 of 26,851: test rows whose chance of buying differs. 1.1e-16: the largest difference. 31 rows moved by one step of the last binary digit, 2 by two. The largest, at the cutoff 2011-08-01: Mac 0.49053056982242593, container 0.49053056982242604. Caption: the four images agreed with each other to the bit. Title: not identical, in the last binary digit.

Inside every image, the versions were exactly the Mac's: Python 3.13.15, numpy 2.5.3, scipy 1.18.1, scikit-learn 1.9.1. Yet 33 of the 26,851 predicted chances of buying were not equal to the Mac's. The largest difference was 1.1e-16, which means 0.00000000000000011.

A computer stores each of these numbers in binary, with a fixed number of digits. 31 of the 33 rows moved by one step of the very last binary digit, the smallest change such a number can make. 2 rows moved by two steps. The largest case was a customer at the cutoff of 1 August 2011: 0.49053056982242593 on the Mac, 0.49053056982242604 in the container.

The base image did not matter. All four images, full, slim, slim with the lock, and multi-stage, gave predictions equal to each other, to the bit. So this was not about slimming. It was about running on Linux instead of macOS, with the Linux builds of the same library versions.

Where the Difference Comes From

I asked this question after seeing the result, and the lab labels it as an addition. Which step of the prediction differs?

A hand-drawn sketch, found after the results. A box: raw score, sum of 54 trees, equal on every row; an arrow to a box: expit, turns the score into a chance; an arrow down to a box: same input, Linux vs Mac: 33 rows differ. Below: I fed the Mac's own raw scores to expit inside the container: 33 rows came back different, the same rows as before; one possible reason: expit calls the C library's exp function, and Linux and macOS ship different C libraries. Caption: not proven which line of code differs; consistent with that.

This model makes a prediction in two steps. First it adds up the outputs of its 54 trees into a raw score. Then a function called expit turns that raw score into a chance between 0 and 1.

So I ran the model inside image c again and saved the raw scores too. The raw scores were equal to the Mac's on every row, to the bit. Then I copied the Mac's own raw scores into the container and ran expit on them there. On the same input, Linux and the Mac disagreed on 33 rows, the same 33 rows as before.

So the difference enters at the last step. Here is one possible reason, which I did not prove line by line. expit uses the C library's exp function, and the container uses glibc while macOS uses Apple's own C library. Two correct libraries can round the very last digit differently. A second possible reason: the Mac and the container ran different compiled builds of the same versions. pip installed macOS arm64 wheels of numpy and scipy on the Mac and manylinux aarch64 wheels in the image, built with different compilers and settings. The result is consistent with either, and I did not separate them.

Does One Digit Matter?

A difference of 1.1e-16 sounds like nothing. But a test that checks "equal" cannot tell nothing from something. So I asked, again after the results: did it change anything the model is used for?

Three panels, found after the results. Test AP: equal; 0.5450 on both, to the bit, in all five months. Rank order: 0 moved, of 26,851 rows. Decisions: 0 flipped, at 0.5, and at the top 20% of each month. Below: the largest row moved from 0.49053056982242593 to 0.49053056982242604: close to 0.5, and still on the same side. Caption: a test that demands equal bits fails here; a test with a tiny tolerance passes.

Test AP was equal to the bit, in every one of the five test months. Not one row changed its place in the ranking from most likely to least likely buyer. And no decision flipped, neither at a cut-off of 0.5 nor at "the top 20% of each month". A shop might use a rule like that to pick who gets an offer.

The closest call was the largest difference itself, a chance of about 0.4905. It moved by two steps of the last digit and stayed below 0.5.

So for this model, on these rows, the Linux container and the Mac give the same results in every way that counts. But the practical lesson is about tests. If your deployment check says "the container's predictions must equal the training machine's predictions exactly", it will fail here, every time, for no real reason. Compare with a tiny tolerance, and check the ranking and the decisions too.

Alpine: The Smallest Base Did Not Build

Alpine's base image was the smallest of the three: 56.4 MB unpacked, against 158.2 MB for slim. Image e used it with exactly c's locked install.

Three rows with logos asking: a wheel for musl, Python 3.13, arm64? numpy 2.5.3: yes, pip downloaded it, in the log. scipy 1.18.1: yes on PyPI; the build stopped first. scikit-learn 1.9.1: no, pip took the source, scikit_learn-1.9.1.tar.gz. Then, compiling it: ERROR: Unknown compiler(s), a list that starts with cc, gcc and clang, cut short. Below: alpine uses musl libc instead of glibc, so a wheel built for glibc, a manylinux wheel, does not fit; with no wheel, pip compiles from source, and alpine has no C compiler; I did not try adding one. Caption: smallest base, and it did not build at all.

The build failed at pip install. The log shows what happened. pip found a musl wheel for numpy. For scikit-learn 1.9.1 it found none, so it downloaded the source, scikit_learn-1.9.1.tar.gz, and tried to compile it. The compile stopped at its first step: "Unknown compiler(s)", because alpine has no C compiler. The build never got as far as scipy. (Whether scipy has a musl wheel I checked on PyPI afterwards; it does.)

After the run I checked PyPI, the site pip downloads from, as a labelled addition. For scikit-learn 1.9.1 it lists 43 files, and none of them is a musl wheel, for any Python version or any processor. So an alpine image on an Intel or AMD server would hit the same wall. The ordinary Linux wheels it does list are manylinux wheels, built for glibc, which alpine does not have.

My guess before the run was half wrong. I expected numpy and scikit-learn to have musl wheels and was unsure about scipy. In fact PyPI has one for scipy, and scikit-learn has none.

You could add a compiler to alpine and compile scikit-learn. I did not try. It would make the image bigger, and the model would then run on a library I compiled myself, not the one the Mac used. This is the musl warning from the python image's description, in practice.

My Guesses Before the Run, Checked

I wrote seven guesses into the lab before it ran. Here they are against the results.

  1. "the full image (a) is the biggest by far, and the base, not my install, is most of it." Right: 1,569.1 MB, and the base was 77% of it.

  2. "the pip install layer is about the same size in a and b (the same wheels), and scipy is its biggest package." Right: the two layers were the same size, 363.5 MB, and scipy was the biggest folder.

  3. "b minus c is about the size of pip's cache, which is about the size of the downloaded wheels." Right on the first part: 61.1 MB against a 59.7 MB cache. I did not measure the downloads separately.

  4. "the multi-stage image d is within 5% of c." Right: d was 3.3 MB bigger than c.

  5. "every image that builds predicts exactly the Mac's numbers, to the bit." Wrong: 33 rows differed in the last binary digit, in all four images. AP was equal.

  6. "alpine (e): numpy and scikit-learn have musl wheels; I do not know about scipy. If one is missing the build fails." Half wrong: scikit-learn was the missing one, and the build failed.

  7. "cache: in a, b, c, d and e a code-only change reruns only the last COPY; a requirements change reruns pip install and every step after it; in c2 a code-only change reruns pip install." Right for a, b, c, d and c2. Image e never built, so it had no rebuilds.

Try It Yourself

The full lab builds five images and rebuilds them many times. I wrote a small demo that builds one, image c's kind, and checks it.

A page in three labelled zones, headed img_demo.py, designed before it ran. Here, your Python: train lesson 1's model, save it with pickle, predict the test rows. Build: a slim image, pinned to the versions your Python has, with --no-cache-dir. Run: copy the rows in, copy the predictions out; compare to the bit, then AP and rank. Caption: it printed 126.6 MB download, 460.8 MB unpacked, 33 rows differ, AP equal, 0 ranks moved.

I wrote the demo's design into its docstring after the lab had run and before the demo first ran. After its first run I changed two small things, and the docstring says so. It now prints the test AP and the rank check, and the layer names are no longer cut mid-word. The first run printed the same sizes and the same 33 rows.

The demo trains the model with your own Python. It writes a requirements file that pins numpy, scipy, scikit-learn and their helpers to exactly the versions your Python has. Then it builds one slim image with --no-cache-dir. Then it runs the image, copies the predictions out, and compares. At the end it removes the container and the image it built. The base image and 's build cache stay on your disk; docker builder prune frees the cache.

A real screenshot of VS Code with img_demo.py open at the top of the file, showing its docstring: what it needs, Docker running, the shop data, how to run it, and the design written before it first ran.

Before you run this lab. You need Docker, running: Docker Desktop on a Mac or Windows, or Docker Engine on Linux, so that docker info works in a terminal. You also need Python 3 with pandas, pyarrow and scikit-learn: pip install pandas pyarrow scikit-learn openpyxl. First run python fetch_data.py from the scripts/labs/features folder; it downloads the shop data once, about 46 MB. The first run of the demo also downloads the python slim base image and the libraries inside the build.

Look Up Any Image Yourself

This box holds the real sizes and rebuild results from the lab. It needs nothing but Python, so it runs in your browser. It builds no image.

Press Run. It shows image c's sizes, its biggest packages, and which steps ran again after a code change. Change IMAGE to "a", "b" or "d", and CHANGE to "again" or "reqs", and run again.

The report script writes this box from the lab's stored results, runs it for every image and every change, and checks what it prints against the lab. Try CHANGE = "code" and look at the line for c2: the only image where a one-line code change installed every library again.

The Lab's Code, Piece by Piece

The lab is one file, scripts/labs/packaging/container_images.py, and the Dockerfiles sit next to it in img_files/.

make_model runs in my main Python. It imports lesson 1's own feature code and trains the model, and it stops if the test AP is not 0.5450. Then it saves the model with pickle, the test rows and the Mac's predictions.

build runs one docker build with plain output, so every step prints a line, and stores the whole log. parse_steps reads that log and marks each step CACHED or ran. make_ctx makes the folder each build reads from; for the rebuild tests it adds one line to predict.py or one comment line to the requirements file.

measure collects everything about one image. It reads docker image inspect, the layer list from docker history, and the manifest from docker save. It runs du and pip freeze inside a container. creates a container, copies the test rows in with , starts it, and copies the predictions out.

How to Build a Model Image

Here is how I would build an image for a model like this one, using only what this lab measured.

A flowchart. A model to put in an image. Is there a wheel for every package on slim? Yes leads to: slim base, hash lock, --no-cache-dir. No leads to: multi-stage, compile in full, copy the venv to slim. Both lead to: order, requirements, install, model, code. Then: predict stored rows inside the image; compare. Below: compare with a tiny tolerance, not equal bits; Linux and macOS differed in the last binary digit here. Caption: alpine only after checking that every package has a musl wheel.

The chart starts from one question: does every package have a wheel for the slim base? Here are the same steps in words, each with the number from this lab behind it.

  1. Start from slim. Here it cut the image from 1,569.1 to 521.9 MB on its own. First check that every package you need has a wheel for it; pip prints "Downloading ... .whl" for a wheel and a .tar.gz name for source.

  2. Install from a lock with hashes, the exact versions that trained the model. Lesson 4 of this chapter shows how to make one.

  3. Add --no-cache-dir. It saved 61.1 MB here and costs nothing in an image that only predicts.

  4. Order the lines from what changes least to what changes most: requirements, install, model, code. Never COPY . . before pip install.

  5. Use a multi-stage build only when something must compile. Here every package was a wheel, and the multi-stage image was 3.3 MB bigger. Copying a venv across works only under two conditions. Both stages must have the same Python at the same path. And the compiled packages must need no system library that slim lacks. Here both stages used Python 3.13.15 in .

When to Slim, and When Not To

Use slim for a model that only predicts, when every package has a wheel. Here it gave the same predictions as the full image, to the bit, at a third of the size.

Use the full image while you are still building and testing, or when a package must compile and you do not want a multi-stage Dockerfile yet. The python image's own description recommends the default image when you are unsure. A smaller image that fails to build helps nobody.

Do not choose alpine for scientific Python without checking first. Here the smallest base did not build at all, because one package had no musl wheel. It may work for a different set of packages, or in another release.

Do not expect a multi-stage build to shrink every image. It removes what the build stage needed and the final stage does not. When nothing is compiled, there is little to remove.

Do not judge an image by one size number. The same image was 126.6 MB to download and 460.8 MB unpacked. The download decides how long a new server waits to start, minus any layers it already has. The disk a server needs is closer to both added together, because the containerd store keeps the compressed layers and the unpacked files.

Do not test containers for bit-equal predictions against a laptop. Here the base did not change a single bit, but the move from macOS to Linux changed the last digit of 33 numbers.

What This Lab Cannot Tell You

Two columns titled shows and cannot show. Shows: sizes for one scikit-learn model on Python 3.13, arm64 Linux images, one day's base images; that the four images predicted the same bits, and differed from the Mac at the very end of 33 numbers; which steps ran again after each change. Cannot show: times, the machine was busy, so no seconds are given; other models, a deep learning library is far bigger; other CPUs, an x86 server may differ in other bits.

One model, one set of libraries. Every size here is for scikit-learn 1.9.1 with numpy and scipy. A model that needs a deep learning library would be a different story, and I did not build one.

One day's base images, on one Mac. The base images were pinned by digest on 1 October 2026, and they were arm64 Linux images, because this Mac has an Apple processor. An x86 server would pull other images with other sizes. Its maths libraries could differ in other last digits, or in none.

No timings. The laptop's load average went as high as 24.01 during the builds, because other work was running. So I report which steps ran, never how long they took.

One kind of change each. The code change was one extra line, and the requirements change was one comment line. A real upgrade would change the installed packages too. The cache rule would be the same, but I did not measure the sizes after such a change.

Alpine was not pushed further. I did not add a compiler, so I cannot say how big a working alpine image would be.

Labelled additions. Five questions were asked after the main results. They are whether the images agree with each other, where the last-digit difference enters, and whether it moved the ranking or a decision. The other two are which wheels PyPI lists for musl, and which tools each base holds. All are labelled in the lab, the report and here.

What to Do on Monday

A hand-drawn grid of six cards, titled five steps. 1, start from slim: check every package has a wheel for it first. 2, lock with hashes: the exact versions that trained the model. 3, --no-cache-dir: no copy of the downloads left in the layer. 4, order the lines: requirements, install, model, code. 5, test inside: predict stored rows in the image; compare with a tiny tolerance. The reason: image a was 1,569.1 MB on disk; c was 460.8 MB, with the same predictions. Caption: most of an image is what you started FROM, and what pip brought.

If you take one thing to work on Monday, open the Dockerfile that serves your model and read its first line. If it says FROM python:3.x with no -slim, build a slim version beside it and compare the two with docker history. Look at which layer is the biggest. In my lab it was the base, then pip install, and the model itself was close to nothing.

Then look at the order of the lines. If COPY . . comes before pip install, move the code copy to the end. Every rebuild after a code change will then reuse the library layer instead of installing everything again.

A closing card titled the base and the libraries are the image. Three numbers in large type: 77%, of image a was its base, 1,205.4 of 1,569.1 MB; 460.8 MB, image c on disk, slim, locked, no pip cache; 33 rows, differed from the Mac in the last binary digit, and AP was equal to the bit.

The one idea to keep: a model's image is mostly not the model. Here it was the base image first and the libraries second, and the model was 0.2 MB. Slimming the base, turning off pip's cache and ordering the lines made the image smaller and the rebuilds cheaper, and none of it changed a single predicted bit. What did change the last digit was moving from macOS to Linux, so test inside the image, with a tolerance.

Knowledge Check

Knowledge Check

4 questions - Score 80% to pass

Q1

Image a used the full python:3.13 base and was 1,569.1 MB on disk. What was most of that size?

Q2

The multi-stage image d was 464.2 MB, and image c was 460.8 MB. Why did multi-stage not help here?

Q3

In image c2, the Dockerfile ran COPY . . before pip install. What happened after a one-line code change?

Q4

All four images gave predictions that differed from the Mac's on 33 rows. What did the lesson find?

lock file
hash
--no-cache-dir

Every base image is pinned by digest. A digest is a sha256 fingerprint of one exact image. A name like python:3.13-slim can point to a new image next week; a digest cannot.

Measured. For each image I measured its size three ways, and the size of every layer. I looked inside the biggest layers. And I checked whether its predictions equal the Mac's, to the last bit. Then I rebuilt each image three more times, after no change, after a code change, and after a requirements change, and read which steps ran again. That is 21 builds in all.

This Docker uses the containerd image store, which Docker Desktop's own page says is enabled by default. I did not find a Docker page that states what Size means in this mode, so I measured it. An older Docker setup may print a different number for the same image.

du came out a little under the history sum. I did not trace why. From here on, "size" means unpacked, on disk, unless I say download.

One small oddity, for honesty. In image c's first build, which used --no-cache, the log still marked the empty WORKDIR /app step as CACHED. I did not trace why; image b had built the same step on the same base just before. It changes nothing above.

Then, inside the examples folder, run python img_demo.py. I ran it on a Mac with an Apple processor, where Docker builds Linux images for arm64. On a Windows or Linux computer with an Intel or AMD processor, Docker builds x86 images. The sizes will differ, and the last-digit differences may be on other rows or none, because the maths libraries differ. If your Python is not 3.13, the demo uses the slim image for your version, without a pinned digest.

r"""Put the chapter's model in a slim container image, and check it predicts the same numbers.

Lesson 5 of 'Packaging, Registry and Versioning'. It needs DOCKER: Docker Desktop
(Mac, Windows) or Docker Engine (Linux), running, so that `docker info` works.
It also needs your normal Python with pandas, pyarrow and scikit-learn, and the
shop data from the features chapter: run scripts/labs/features/fetch_data.py
once first. Then, inside this folder:
    python img_demo.py              # print the results
    python img_demo.py out.json     # and save them
The first run downloads the python slim base image (about 43 MB) and the
libraries inside the build. It prints no timings.

Design, written 2026-10-01 after the lab (container_images.py) had run and
before this file first ran:
  1. Train lesson 1's model here, with your Python, and save it with pickle.
     Predict the 26,851 test rows here too.
  2. Write a requirements file that pins numpy, scipy, scikit-learn and their
     helpers to EXACTLY the versions your Python has, so the image gets the
     same libraries (the lab's image c does the same with a hash-locked file).
  3. Build ONE image on python:<your version>-slim with pip's --no-cache-dir.
     On Python 3.13 the base is pinned to the digest the lab used.
  4. Print the image's size two ways: `docker image inspect` (Size) and the
     sum of `docker history` (the unpacked layers), and the biggest layers.
  5. Run it: copy the test rows in, start it, copy its predictions out, and
     compare them with step 1's, to the last bit; then the test AP and the
     rank order, which is what the score depends on.
  6. Remove the container and the image. The base image and Docker's build
     cache stay; `docker builder prune` frees the cache if you want.

Changed after its first run (2026-10-01): the layer names lost a trailing
"# buildkit" that was cut mid-word, and the test AP and rank lines were
added. The first run printed the same sizes, 33 rows and 1.11e-16.

Author: Roni Das
Created: 2026-10-01
"""
import json
import pickle
import platform
import shutil
import subprocess
import sys
import tempfile
from importlib import metadata
from pathlib import Path

import numpy as np
from sklearn.ensemble import HistGradientBoostingClassifier as HGB
from sklearn.metrics import average_precision_score

HERE = Path(__file__).resolve().parent
sys.path.insert(0, str(HERE.parents[1] / "features"))
import task  # noqa: E402
from what_a_feature_is import HAND_COLS as COLS, joined  # noqa: E402

LAB_313_SLIM = ("python:3.13-slim@sha256:"
                "7c61056e61ac89e852de05f3dc6fa51a6dd2181797bceed46aa725dd7cb2cd3b")
TAG = "img-demo:latest"
PINNED = ["numpy", "scipy", "scikit-learn", "joblib", "threadpoolctl", "narwhals", "cloudpickle"]


def docker(*args, check=True):
    r = subprocess.run(["docker", *args], capture_output=True, text=True)
    if check and r.returncode != 0:
        sys.exit(f"docker {args[0]} failed:\n{r.stderr.strip()[-1500:]}")
    return r


if shutil.which("docker") is None or docker("info", check=False).returncode != 0:
    sys.exit("Docker is not running. Start Docker Desktop (or the Docker service) and try again.")

# 1. train and predict here
ev = task.load_events()
lab_tr, _, lab_te = task.splits(ev)
tr = joined(ev, lab_tr, task.TRAIN_CUTOFFS)
te = joined(ev, lab_te, task.TEST_CUTOFFS)
x_te = te[COLS].to_numpy(float)
model = HGB(random_state=0).fit(tr[COLS].to_numpy(float), tr["label"])
here = model.predict_proba(x_te)

py = ".".join(platform.python_version_tuple()[:2])
base = LAB_313_SLIM if py == "3.13" else f"python:{py}-slim"
with tempfile.TemporaryDirectory() as tmp:
    tmp = Path(tmp)
    ctx = tmp / "ctx"
    ctx.mkdir()
    with open(ctx / "model.pkl", "wb") as f:
        pickle.dump(model, f, protocol=5)
    shutil.copy(HERE.parent / "img_files" / "predict.py", ctx / "predict.py")
    # 2. pin the libraries to the versions this Python has
    pins = []
    for name in PINNED:
        try:
            pins.append(f"{name}=={metadata.version(name)}")
        except metadata.PackageNotFoundError:
            pass
    (ctx / "requirements.txt").write_text("\n".join(pins) + "\n")
    (ctx / "Dockerfile").write_text(
        f"FROM {base}\n"
        "WORKDIR /app\n"
        "COPY requirements.txt .\n"
        "RUN pip install --no-cache-dir -r requirements.txt\n"
        "COPY model.pkl .\n"
        "COPY predict.py .\n"
        'CMD ["python", "predict.py"]\n')
    print(f"base image: {base}")
    print("pinned:", ", ".join(pins))

    # 3. build
    docker("build", "--provenance=false", "--sbom=false", "-q", "-t", TAG, str(ctx))

    # 4. size, two ways, and the biggest layers
    size = int(docker("image", "inspect", TAG, "--format", "{{.Size}}").stdout)
    hist = [line.split("\t", 1) for line in docker(
        "history", "--no-trunc", "--human=false", "--format", "{{.Size}}\t{{.CreatedBy}}",
        TAG).stdout.splitlines()]
    layers = sorted(((int(s), c) for s, c in hist), reverse=True)
    print(f"\nimage {TAG}")
    print(f"  docker image inspect Size: {size / 1e6:,.1f} MB")
    print(f"  sum of docker history:     {sum(s for s, _ in layers) / 1e6:,.1f} MB "
          f"in {len(layers)} history rows")
    print("  the three biggest layers:")
    for s, c in layers[:3]:
        c = " ".join(c.replace("/bin/sh -c ", "").replace("# buildkit", "").split())
        print(f"    {s / 1e6:8.1f} MB  {c[:60]}")

    # 5. run it: rows in, predictions out
    docker("rm", "-f", "img-demo-run", check=False)
    docker("create", "--name", "img-demo-run", TAG)
    np.save(tmp / "x_te.npy", x_te)
    docker("cp", str(tmp / "x_te.npy"), "img-demo-run:/app/x_te.npy")
    print("\ncontainer says:", docker("start", "-a", "img-demo-run").stdout.strip())
    docker("cp", "img-demo-run:/app/pred.npy", str(tmp / "pred.npy"))
    docker("cp", "img-demo-run:/app/versions.json", str(tmp / "versions.json"))
    there = np.load(tmp / "pred.npy")
    inside = json.loads((tmp / "versions.json").read_text())

    # 6. clean up the container and the image
    docker("rm", "-f", "img-demo-run")
    docker("image", "rm", TAG)

diff = np.abs(there - here).max(axis=1)
same = bool(np.array_equal(there, here))
y, cut = te["label"].to_numpy(), te["cutoff"].to_numpy()


def ap(p):
    """Test AP: the plain mean of the five per-month APs, as in the features chapter."""
    return float(np.mean([average_precision_score(y[cut == c], p[cut == c]) for c in np.unique(cut)]))


def ranks(p):
    r = np.empty(len(p), int)
    r[np.argsort(-p, kind="stable")] = np.arange(len(p))
    return r


moved = int((ranks(there[:, 1]) != ranks(here[:, 1])).sum())
print(f"  inside: Python {inside['python']}, {inside['machine']}, libc {' '.join(inside['libc']) or 'unknown'}")
print(f"  here:   Python {platform.python_version()}, {platform.machine()}, {platform.system()}")
print(f"\nidentical on all {len(te):,} test rows: {same}; rows that differ: {(diff > 0).sum():,}; "
      f"largest difference: {diff.max():.3g}")
print(f"test AP here {ap(here[:, 1])!r}, inside {ap(there[:, 1])!r}; rows whose rank moved: {moved}")
print(f"removed the container and {TAG}; the base image stays")
out = {"base": base, "pins": pins, "inspect_size": size, "history_sum": sum(s for s, _ in layers),
       "history_rows": len(layers), "biggest_layer": layers[0][0], "identical": same,
       "rows_differ": int((diff > 0).sum()), "max_abs": float(diff.max()), "test_rows": len(te),
       "ap_here": ap(here[:, 1]), "ap_inside": ap(there[:, 1]), "rank_moved": moved,
       "inside": inside}
if len(sys.argv) > 1:
    json.dump(out, open(sys.argv[1], "w"), indent=1)

This is a real run in VS Code's terminal, inside the examples folder.

A real screenshot of VS Code's terminal after running python img_demo.py. It prints the base image, python:3.13-slim pinned by digest, and the seven pinned packages. Then image img-demo:latest: docker image inspect Size 126.6 MB; sum of docker history 460.8 MB in 15 history rows; the three biggest layers, pip install 302.4 MB, the Debian layer 109.5 MB, and a Python layer 43.7 MB. Then the container's line, scored 26851 rows with scikit-learn 1.9.1, numpy 2.5.3; inside Python 3.13.15, aarch64, glibc 2.41; here Python 3.13.15, arm64, Darwin. Then identical on all 26,851 test rows: False; rows that differ: 33; largest difference 1.11e-16. Test AP here and inside both 0.5449816307508443; rows whose rank moved: 0.

When I ran it, it gave the lab's numbers: 126.6 MB to download and 460.8 MB unpacked. 33 rows differed by at most 1.11e-16, the test AP was the same to every digit, and no row moved in the ranking. The report script checks the demo's saved run against the lab's image c.

predict
docker cp

after holds the questions asked after the results. Do the images agree with each other? Where in the prediction does the difference enter? Did the ranking or any decision move? Which wheels does PyPI have for musl? after5 lists the tools in the two bases. clean removes every image and build-cache record the lab made.

img_report.py checks all of it from the stored outputs, as the recording showed.

/usr/local
  • Pin the base by digest, so the starting point cannot change under you.

  • Test inside the image. Keep some input rows and the predictions the model gave on them. After each build, predict those rows inside a container and compare: a tiny tolerance, the same ranking, the same decisions.

  • Two more habits, which this lab did not test. Never copy a secret, such as a password or an access key, into an image. Every layer travels with the image, and deleting the file in a later step does not remove it from the earlier layer.

    And add a USER line so the container does not run as root; my images ran as root, which is the default.