This idea carries a full system design question on its own. Each walks through the full answer.
A shop owner adds up his March sales on the first of April. The total goes into a report, and the report goes to the bank. In June, he adds up March again, from the same sales records. He gets a smaller number.
Neither total is wrong. In April and May, customers sent things back. Each time, the shop wrote the return against the sale it undid, and some of those sales were made in March. So March changed, weeks after it had ended.

Now put a model in the owner's place. In April it learned from the March sales. In June someone asks what it learned from. If the only records left are the ones that were kept up to date, the answer is March as it looks in June. The model never saw that March. Only a copy made in April can answer the question.
Keeping those copies is what this lesson is about, and keeping them cheaply enough that you really do it. I measured it on a real shop's sales: 1,047,877 sale lines over two years, changed 16,675 times by real returns. I kept 110 versions of that table in four different ways, and counted every byte.
This is the sixth lesson of the chapter on data engineering for machine learning. Lesson 4, on data lakes, warehouses and lakehouses, built a table with a log of its versions. Lesson 5, on data validation, asked which checks catch a bad batch before it lands.
Another lesson, in the lifecycle chapter, measured the other half of this story. Lineage and rollback rebuilt last month's model from its written record. When the old labels had been corrected upstream, the rebuild agreed with the original on only 0.9583 of its answers. Its advice was to keep a stored version of the data that nobody edits in place. This lesson is about that advice. It covers how the tools keep versions, what keeping them costs, and how long they last. It ends with what happens when a person asks to be erased.
What this rebuild corrected. An earlier version of this lesson made seven claims its sources do not support. Each is fixed where it comes up, and the record is in
results/pin-factcheck.json.
- It opened with a story about a fraud model at DoorDash. I found no source for it.
- It said Git LFS gives "no dataset branching and no time travel". It does, because its pointers live in Git.
- It said DVC catches a damaged file "on pull". Only when the remote's
verifysetting is on.- It said storage "grows with change, not with versions". It grows with changed files.
- It said a lakeFS branch of "a 50 TB lake takes under a second". I found no measurement behind that.
- It said writes on a lakeFS branch "only touch metadata". They store new objects.
- It said an audit "six months later" gets "byte-identical rows". Only if retention kept the files.

The code and the settings are small text, and Git, the tool that keeps every version of a project's code, keeps every version of them too. The data is the input that keeps moving. New lines arrive every week, and old lines change when returns arrive. Pinning a model's data means being able to read, later, the exact rows it learned from.
Here is the plan. First, why the data does not go in the system that keeps the code, and how four common tools keep versions instead, each checked against its own documentation. Then the lab: one real table, 110 versions, four ways to store them. Then what the results mean for naming old versions, deleting them, and erasing a person.

A version is the table exactly as it was at one moment. A commit saves a new version, with a note of what changed. A hash is a short code computed from a file's bytes. Change one byte and the hash changes. Two files with the same bytes always get the same hash.
A store that keeps each file by hash names the file by its hash, not by where it sits, so the same bytes are kept only once. A snapshot is a table format's word for a version: the list of data files the table was made of at that moment. A tag is a name that points at one version, such as the version a model trained on.
Time travel means reading a table as it was at an old version or an old time. Retention is the rule for how long old versions are kept before their files are deleted. To rewrite a file is to change one row by writing the whole file again. Erasure is deleting one person's data because they asked. The tables in this lesson are stored as Parquet files, a common file format that stores a table column by column, compressed.
The first idea is a sensible one. Git keeps every version of the code, so put the data in Git too. It fails for reasons that are worth knowing, because the fix follows from them.

Git keeps every version of every file you commit. A normal clone, the copy of a project that each person downloads, contains all of them. Git can store a new version as a difference from the last one. The Git book says it "stores just the deltas from one version of the file to the next". A Git clone can also be made shallow, with --depth, to fetch fewer commits. But neither helps much with a large compressed file. GitHub also sets hard limits. Its documentation says "GitHub blocks files larger than 100 MiB", and Git warns above 50 MiB.
A diff is the other problem. For code, a diff shows the lines that changed, and a reviewer can read it. A Parquet file is compressed binary, so its diff says nothing a person can use.
Git LFS, Git Large File Storage, keeps a small pointer in Git and the file itself on a server. The pointers live in Git, so branches and old commits work the way they do for code. What LFS does not do is store less. GitHub's own example: "If you make a 1 byte change and push the file again, you'll use another 500 MB of storage."
The fix that every tool in this lesson uses is the same. Keep a small pointer in Git, and keep the heavy bytes in a store built for them.
DVC, Data Version Control, is an open-source tool that works beside Git. Its trick is content addressing: a file is named by a hash of what it holds, not by where it lives. A path such as s3://bucket/train_final.parquet can be overwritten tomorrow and still have the same name. A hash cannot. Change one byte and the name changes.

You run dvc add on a file or a folder. DVC computes an MD5 hash of each file, which is the only hash its documentation lists: "only md5 is currently supported". dvc push copies each file to a remote, a store such as S3, under a path made from its hash. DVC writes a small .dvc file with the hash, and adds the data itself to .gitignore so Git never takes it. You commit the .dvc file. A teammate who checks out that commit runs dvc pull and gets the files named by those hashes.
Two things follow from naming by bytes. First, the same bytes are stored once. DVC's documentation says that two files "with different names but the same contents" are tracked, "but only one copy is stored in the cache". Second, a changed file gets a new hash, so nobody has to remember to bump a version number.
By default, DVC does not check a file's hash again when you pull it. It re-hashes files on only when a remote's setting is on, and the documentation says it is on by default only for Google Drive remotes. If you want that check, turn it on, and expect pulls to be slower.
Storing files by hash means storage grows with what changed. That helps only if you know what a tool counts as a change. None of these tools looks inside a file. Change one row, and the whole file is new.
I checked that with real DVC, version 3.67.1, on the lab's own files. Here is what it did.

I wrote the last version of the lab's table as one Parquet file of 1,047,820 rows, 12.1 MB, and pushed it with DVC. Then I added 1 to one row's returned column, a column of units sent back, and pushed again. The remote grew by another 12.1 MB. In 2021 someone asked on DVC's own forum whether a tiny change to a file makes DVC keep a second copy. A reply there said: "Yes, we just save every one of them independently."
Then I gave DVC all 110 versions of the table, cut into one file per month. After 110 dvc add and dvc push runs, its remote held 816 files and 431.9 MB of data, plus 104.8 KB of DVC's own folder listings. Those totals match the lab's month store to the byte: 816 files and 431,946,842 bytes. DVC was given the lab's own files, so this does not test how I wrote them. It tests my counting: storing whole files by hash, DVC kept exactly what my month store said it would.

None of this is abstract. Here is what DVC's setup looks like, next to lakeFS's commands.
A DVC project keeps its settings in .dvc/config, an INI file: plain text, with [sections] of name = value lines. This one names an S3 bucket as the default remote and turns on the hash check on pull:
# .dvc/config
[core]
remote = storage
['remote "storage"']
url = s3://ml-datasets/dvcstore
region = us-east-1
verify = true
The workflow, in the shell:
dvc add data/sales # hash each file; write data/sales.dvc; add data/sales to .gitignore
dvc push # copy each file to the remote under its hash
git add data/sales.dvc data/.gitignore
git commit -m "Sales table, week 51"
git tag train-week-51 # the Git tag names this data version too
The .dvc file is the pointer that Git stores. This is a real one. I made it with DVC 3.67.1 by adding a six-byte test file, and it is 82 bytes long:
outs:
- md5: b1946ac92492d2347c6235b4d2611184
size: 6
hash: md5
path: a.txt
It holds the file's MD5, its size and its path. For a folder, the MD5 names a small listing of every file's hash, which DVC calls a .dir entry. After dvc push, the file sat in the remote at files/md5/b1/946ac92492d2347c6235b4d2611184. The first two characters of the hash name the folder, and the rest name the file.
lakeFS versions a whole bucket instead, with Git-like commands. Its command-line tool is lakectl. lakeFS also has an S3 gateway: it answers the same requests Amazon S3 does, so ordinary S3 tools can read and write its branches. I have not run lakeFS for this lesson; these commands are copied from the syntax in 's reference:
Versions let you get data back. Branches let you change data without breaking what others are reading. This is the job lakeFS is built for.
Say you want to relabel 12,000 rows and see whether a model gets better. The labels live in a that production jobs read every hour. Change them in place and every reader sees your half-finished work.
In lakeFS you branch first. Its documentation is exact about what that costs. "Creating a branch is a zero-copy operation: instead of duplicating data, lakeFS creates a pointer to the source commit for the branch." The documentation's word is "pointer", and a pointer does not grow with the lake.
Writes on the branch are not free. lakeFS stores each object it is given at a new, random address, and keeps the old one, so the relabelled files are real new bytes. What costs nothing is the branch itself.
When the experiment wins, you merge. lakeFS says merges "atomically update one branch with the changes from another": readers of main see all of the change or none of it. When it loses, you delete the branch. And by default, "lakeFS keeps all your objects forever", until you set rules for its garbage collection, the job that deletes objects no branch still needs.
DVC and lakeFS version files. A lot of training data lives in tables instead, and two table formats build versioning in. Delta Lake and Apache Iceberg write every change as a new snapshot and leave the old data files where they were. Iceberg began at Netflix. It joined the Apache Software Foundation's incubator in November 2018, and its proposal says it was "under active development at Netflix".
Reading an old snapshot is one line of . These are the forms in each project's documentation for Spark, a system that runs one query across many machines. I did not run Spark for this lesson:
-- Delta Lake: the table as it was at version 51
SELECT * FROM sales VERSION AS OF 51;
-- Iceberg: the table as it was at a time
SELECT * FROM prod.db.sales FOR SYSTEM_TIME AS OF '2010-11-30 00:00:00';
-- Iceberg: name the snapshot a model trained on (7302 stands for its id).
-- With no RETAIN the tag is kept until you drop it; RETAIN 365 DAYS would drop it after a year.
ALTER TABLE prod.db.sales CREATE TAG `train-week-51` AS OF VERSION 7302;
SELECT * FROM prod.db.sales VERSION AS OF 'train-week-51';
An audit months later can rerun that query only if the old files still exist. Each format deletes them on purpose, and the defaults are short.

Delta Lake keeps two windows, as lesson 4 explained. delta.logRetentionDuration, 30 days by default, says how long the log that lists each version keeps its entries. delta.deletedFileRetentionDuration, 7 days by default, says how long a data file the table has stopped using survives VACUUM, Delta's command for deleting old files.
I wrote the lab's design into the docstring of its script, versioning_demo.py, on 30 September 2026, before it ran. The script is the same file you can copy from this lesson. Before writing the design, I looked at the data once, to check the plan could work. I counted the cancellations and how far back their sales were. Nothing else ran.

The data is UCI Online Retail II, the same shop as lesson 3. It holds every invoice line of a UK online gift shop from 1 December 2009 to 9 December 2011. The table is the shop's sales, 1,047,877 lines, each with an id, which is its position in the file. I added one column, returned, the units sent back, which starts at 0.
An invoice number that starts with C is a cancellation, 19,494 lines in all. A cancellation is not a new row in the table. It is a correction: it adds its units to returned on the most recent earlier sale line by the same customer for the same product. 16,675 cancellations found such a line. The other 2,819 had no customer id, or no earlier sale to match, and changed nothing.
The versions. One commit a week, starting on 1 December 2009: first that week's sale lines are added, then that week's cancellations are applied. The design fixed four more commits in advance.
After week 40, the table is written again with nothing changed, as when a job runs twice. After week 60, the same rows are written in another order, sorted by customer and product. That is what happens when a query loses its , the part of a query that sets the order of the rows. In week 70, the returns arrive late, as a second commit. After week 95, one customer asks to be erased. That makes 110 versions over 106 weeks.
Here is what the demo stored in results/pin-demo.json. The last version of the table, as month files, is 14.5 MB.

Keeping a whole copy of every version cost 754.5 MB, 52 times the last version's size. The copy line curves upward because each copy is bigger than the one before. Storing whole copies by hash, so that a version with no change at all costs nothing, saved only 12.0 MB. Just two versions had no change: the table written twice, and the week from 28 December 2010, when the shop had no lines and no returns.
The month store kept 431.9 MB, 43% less than the copies. The week store kept 301.6 MB. The changes store kept 27.8 MB. The one step in the changes line, just after commit 60, is the commit that wrote the rows in another order. Every file was new, even though no value changed.

The heights are to scale. Every store holds the same 110 versions. To check that, the demo read each of the 8 tagged versions back from all three stores kept by hash. It compared the rows with the rows recorded when each was tagged. All 24 matched. That is a check on my code, not a finding: a store that loses rows is a bug. The copies were not read back.
Storage is not the only cost, though. To read the last version, a reader opened 25 files in the month store, 105 in the week store and 210 in the changes store. The changes store also read 1,060,857 rows to return 1,047,820, because a changed line is stored twice and the reader keeps the newest copy. Every new change file makes every read a little slower. So formats that write change files also fold them back into the main files now and then, a job called . Delta's documentation lists among the commands that apply its deletion vectors this way.
A return is written against the sale it undoes, and the sale can be old. The demo measured how old: the median return reached back 10.1 days. 26.5% reached back more than 28 days, 9.4% more than 90, and the longest reached back 655.9 days.

After the results, I counted the same thing in weeks, since the week store cuts files by week. 3,289 returns landed in the same week as their sale, and 4,855 in the week before. 51.2% reached back two weeks or more. And 48.2% landed in an earlier month than the month of the return, so in the month store they rewrote an older month's file.

That is where the bytes went. Across the weekly commits, the month store wrote 43.6 MB for the months that got new lines, and 369.0 MB for months that got only returns. A week rewrote a median of 7 month files, and at most 20. In 96.2% of weeks it rewrote more than one.
Each of those rewrites stored a whole month again to record a few changed numbers. The week store's files are smaller, so each rewrite cost less. The changes store never rewrote an old file for a return at all. It wrote the changed lines into one small new file.
A table changes in more than one way. The four extra commits in the design each stand for something that really happens to training tables, so here is what each one cost.

Written again cost nothing in any store kept by hash. That depends on one fact I had to check. Writing the same table with the same library settings must give the same bytes. With pyarrow 25.0.1 it did: every file came back with the same hash, so the commit added 0 bytes. If a writer puts the time of writing into its files, this saving disappears.
Another order is the opposite case. Not one value changed, yet every file in every store was new: 8.2 MB in the month store and 10.1 MB in the changes store. A hash of the bytes cannot tell a new order from new data.
The late returns changed 134 lines. The month store wrote 5.7 MB to record them, and the changes store 9.6 KB, 595 times less. The erasure removed 57 lines, and cost the month store 5.1 MB and the changes store 2.8 KB. The changes store records a deleted line with its id only, and I come back to why that matters.
Before you trust a total, it helps to check it roughly with a pencil. Here are five of this lesson's numbers, worked out by hand.
The copies. Each version is a whole copy, and the table grew from nothing to 14.5 MB. If it had grown evenly, the average version would be about half the last one, 7.3 MB. The measured total, 754.5 MB over 110 versions, is an average of 6.9 MB. The two are close. After a review, the report averaged the line counts of all 110 versions: the average version held 47.2% of the final lines, not half. And 47.2% of 14.5 MB is 6.9 MB.
The late returns. The month store wrote 5,707,709 bytes for them, and the changes store 9,592 bytes. Divide one by the other and you get 595. Both stores recorded the same 134 changed lines.
The date filter. At week 51, 911 lines had a different returned and 42 were erased. 911 plus 42 is 953. Divided by the tag's 490,766 lines, that is 0.19%, which is where "under 0.2%" comes from.
The price of the tags. Keeping the last 4 versions left 37.8 MB in the month store. Keeping the tags as well left 82.0 MB. The difference, 44.2 MB, is what it cost to keep the 7 tags that were not already among the last 4.
The erasure marker. The changes store recorded the 57 erased lines in one file of 2,803 bytes. Each line is written as its id and a flag, with the rest of its row left empty. Every Parquet file also carries a description of its own columns. After a review, the report wrote an empty file with the same 11 columns: 1,799 bytes. So the 57 markers themselves took about 1 KB. Either way, the marker holds no one's data, which is why the lines' real data had to be removed separately.
A team usually learns that its data changed from a check. So for every commit, the demo compared the new version with the one before, in four ways. The row count: did the number of lines change? The schema: did the column names or types change? The file hashes: did any month file get a new hash? And a row hash, a hash of every row's values that ignores their order.

The late returns are the case that matters. 134 lines changed, and the row count stayed at 665,731. The schema did not move either. A pipeline that watches only those two would say the data was the same. The file hashes and the row hash both saw it.
Most of this follows from how each check works, not from the data. A schema check cannot see changed values, and a file hash cannot tell a new order from new data. The reorder shows the second: the file hashes fired, and the row hash stayed quiet, because the rows were the same rows.
So each check has its own job. The row count and the schema are cheap, and they belong in validation, which lesson 5 covered: they catch a batch that is missing or malformed. A file hash is free once files are stored by hash, and it says that something changed. A row hash costs one pass over the data, and it says whether the rows themselves changed, in any order. Store it with each version, and a later reader can tell "the same data, stored again" from "different data".

Here is the mistake the opening story is about. A team wants to know what the week-51 model trained on. The table is still there, so someone runs it with a date filter: every line dated before the end of week 51. Is that the week-51 data?

It is not. For each tag, the demo took the last version, kept the lines dated before the tag's week ended, and compared them with the tag. At week 51 the tag had 490,766 lines. In today's table, 911 of those lines had a different returned, because returns arrived after week 51 for sales made before it. 42 more lines were gone, erased with the customer. Units returned went from 185,138 on the tag's 490,766 lines to 215,129 on the 490,724 still in today's table, so the two sums cover slightly different lines.
953 of 490,766 is under 0.2% of the lines. That sounds small. But the lineage lesson measured what a small correction can do: correcting under 1% of its labels changed 250 of its model's 6,000 answers.

A date filter on a table that is corrected in place answers a different question. It says what the rows for those dates hold now. Only the stored version says what they held when the model trained.
The filter fails because the returns changed old rows in place. After a review, the report filtered both tables instead: the sales dated before the tag's week ended, with only the returns dated before it too. That gave every tag's returns back exactly, 0 lines different. Only the erased lines were still missing, 42 at week 51. In this lab that is true by construction, since each week's returns were applied in that week's commit. A table that overwrites old values keeps no dated record to filter.
Keeping all 110 versions forever is the expensive end. Real teams delete old versions on a rule. At the end of the run, the demo tried three rules. Keep everything; keep only the last 4 versions; or keep the last 4 and the 8 tags. Then it deleted every stored file that no kept version needed, the way expire_snapshots and VACUUM do.

Keeping the last 4 versions shrank the month store from 431.9 MB to 37.8 MB. Keeping the tags as well took it to 82.0 MB, 2.2 times as much. That extra 44.2 MB is the price of being able to read what the 8 models trained on.

With only the last 4 kept, just 1 of the 8 tags could still be read, the week-103 tag, which was one of the last 4. For the others, only a few scraps survived. In the month store, 4.4% of the week-77 tag's bytes were still there, and nothing of the tags before it.
The changes store kept 4 of 8, with no rule sparing them. Its files are only ever added to, so an old version's files are a subset of a newer one's. Tags from after the reorder commit were still whole. That is a side effect of the layout, not a plan. Compaction, which folds the change files back into the main files, would remove it.
Keeping every version has a cost that is not measured in bytes. Under the EU's data protection law, GDPR, a person may ask for their personal data to be erased. Article 17 says the company must then erase it "without undue delay" where one of the listed grounds applies. The same law says personal data should be kept "for no longer than is necessary". In the lab, one customer asks to be erased after week 95. The customer was drawn at random, with seed 0, from those with a line in the table by then.

Customer 17388 had 57 lines, in 9 month files. The erasure commit removed them from the latest version. But every version kept from before still held them: 98 of the whole copies, and 242 stored month files. Keeping the last 4 and the tags, 7 kept tags still held them, in 33 month files. To erase them there too, the demo purged: it rewrote each of those files without the lines. That wrote 18.1 MB for the month store, 5.4 MB for the week store and 2.6 MB for the changes store.
Rewriting files erases nobody by itself, and the lab's purge only counts the bytes it wrote. The old files still exist until something physically deletes them. That means VACUUM, expire_snapshots or lakeFS's garbage collection. For DVC it means dvc gc, with its --cloud option for the remote, run on every copy of DVC's cache as well. A tag or a .dvc hash cannot be changed in place either. It has to be moved to the rewritten version.
The changes store has a trap of its own. Its erasure wrote a small file that only marks the lines as deleted, 2.8 KB. The lines themselves stayed in the files underneath: 10 files of the latest version still held them. A reader never sees them, but they are still stored. Delta's documentation warns about the same thing with deletion vectors: "Modified data might still exist in the old files. You can run VACUUM to physically delete the old files."
The four stores are my own code, so I checked the one that copies a real tool against that tool. I ran DVC and MLflow, each in its own Python environment, on 30 September 2026.

The script dvc_check.py ran the demo's 110 commits again. For each version, it wrote that version's month files into a folder, then ran dvc add and dvc push to a folder used as the remote. At the end, the remote held 816 files and 431,946,842 bytes. The demo's month store held 816 files and 431,946,842 bytes. The only extra was DVC's own folder listings, 104.8 KB.
The script mlflow_digest_check.py matched the returns another way, with pandas' merge_asof, and found the same 16,675. Then it logged five versions of the sale lines with MLflow. The results are on the slide "Which Check Saw the Change". One more detail from the run: the digest is 8 characters long, for example 53d5ccc7.
Both checks ran once, on one version of each tool. A later version could change what they do.
A version only helps if the model's record points at it. Here is the record I would keep, and how to check it later.

MLflow already records the code: its documentation lists the tag mlflow.source.git.commit, set "when run from git repo". The rest you add. This snippet records the version the job read and a hash of every row, next to MLflow's own dataset record:
import hashlib
import mlflow
import mlflow.data
import pandas as pd
def row_hash(df):
# a hash of every row that ignores their order
h = pd.util.hash_pandas_object(df, index=False).sort_values()
return hashlib.sha256(h.to_numpy().tobytes()).hexdigest()
# train: the rows read from the tagged version
with mlflow.start_run():
mlflow.set_tag("data.version", "train-week-51") # the tag or snapshot id
mlflow.set_tag("data.row_hash", row_hash(train))
mlflow.log_input(mlflow.data.from_pandas(train), context="training")
I ran it with mlflow-skinny 3.16.1 on a three-row table, and it logged both tags and the dataset. With a local folder as MLflow's store, version 3.16 refuses to start unless you set MLFLOW_ALLOW_FILE_STORE=true, or run a tracking server. With no settings at all, mlflow-skinny 3.16.1 tried sqlite:///mlflow.db and stopped with an error, because the skinny package cannot use that address. row_hash gave the same hash for the rows in another order.
The pair of the Git commit and the data version lets you find the exact code and the exact rows later. It does not rebuild the same model by itself. The lineage lesson found that a rebuild also needs the seed and the library versions. With the seed missing, 19 guessed seeds agreed with the original on 0.9197 to 0.9702 of rows. None matched exactly.
This script is the lab. It downloads the shop's data, matches every cancellation to its sale, and builds all 110 versions in the four stores. It prints what each store kept, which check fired on which commit, how today's table differs from each tag, and what each retention rule keeps. It needs no GPU. On my laptop it took between 43 and 83 seconds, depending on what else was running. It used about 3 GB of memory at the peak, because every stored file stays in memory.

Before you run this lab. You need Python 3 and three libraries: pip install scikit-learn pandas pyarrow. scikit-learn holds the download, and brings NumPy with it. pandas holds the dates, and pyarrow writes and reads the Parquet files. The first run downloads Online Retail II from OpenML (about 15 MB), so it needs an internet connection once. After that, scikit-learn keeps a copy in a folder in your home directory (scikit_learn_data).
I ran it with scikit-learn 1.9.1, pandas 3.0.6 and pyarrow 25.0.1 on a Mac. These libraries run on Windows and Linux too, but I have not checked the numbers there. Another version of pyarrow may write slightly different bytes, so the first line printed is the version. Give it a file name, python versioning_demo.py out.json, and it also saves every number; that is how results/pin-demo.json was made.
"""Data versioning: what keeping every version of a real table costs.
Lesson 6 of 'Data Engineering for ML', made small. It needs Python 3 with
scikit-learn, pandas and pyarrow (pip install scikit-learn pandas pyarrow).
The first run downloads UCI Online Retail II from OpenML (about 15 MB) and
keeps a copy. It writes nothing to disk: every file it stores stays in
memory, about 3 GB at the peak. It took 43 to 83 seconds on my laptop.
python versioning_demo.py # print the results
python versioning_demo.py out.json # and save every number
Design, written 2026-09-30 before the first run. Before writing it I
counted the cancellations and how far back their sales were, to check
that the plan could work; nothing else was run.
Data: every invoice line of a UK online gift shop, 2009-12-01 to
2011-12-09. The table is the shop's sales: every line whose invoice
does not start with C, with its position in the file as its id, and a
column `returned`, units sent back, which starts at 0. A line whose
invoice starts with C is a cancellation. It is not a row; it is a
correction. It adds its units to `returned` on the most recent earlier
sale line by the same customer for the same product. A cancellation
with no such line changes nothing.
Versions: one commit a week from 2009-12-01: that week's sale lines
are appended, then that week's cancellations are applied, in time
order. Four more commits, fixed now:
after week 40 the same table, written again (a job ran twice)
after week 60 the same rows in another order, sorted by customer
then product (a query lost its ORDER BY); the table
keeps that order from then on
week 70 the returns arrive late: the week's lines are one
commit and its returns a second
after week 95 a customer asks to be erased: every line of theirs
is deleted. Who: drawn at random (numpy seed 0) from
the customers with a line in the table by then.
Files: Parquet, pyarrow's defaults, one fixed column order and types.
Four ways to keep the versions. All but the first store a file under
the SHA-256 of its bytes, so the same bytes are stored once:
copy every version is a whole copy of the month files
month one file per month of sale; a change rewrites that file
week one file per week of arrival; new lines always go in a new
file; a change rewrites the file that holds the line
changes like week, but changed lines are written, whole, into one
new small file per commit; a reader keeps the newest copy
of each line and drops deleted ones
Reported: bytes stored for all versions, and per kind of commit; also
a whole copy that is skipped when nothing changed at all; files read
for the last version. Four checks on every commit, against the one
before: row count, column names and types, the month files' hashes,
and a hash of the rows that ignores their order. Which ones fire.
Tags: every 13 weeks a model would be trained (weeks 12, 25, ... 103;
no model is trained here). Those versions are tagged, and each tag is
read back from every store and checked against its row hash.
"As of": for each tag, the last version's lines with an invoice date
before the end of that week, against the tag: lines that differ.
Retention, at the end: keep everything; keep the last 4 versions;
keep the last 4 and the tags. Stored files no kept version needs are
deleted. Reported: bytes kept, tags that still read back, and stored
files that still hold the erased customer's lines. Then a purge that
rewrites those files, and what it costs. One run; every count and
byte is exact.
Changed after the first run, 2026-09-30, and said so in the lesson:
In `changes`, a deleted line was written whole into the change file,
which stored the erased customer's data again. It now keeps only the
line's id, as Iceberg's delete files and Delta's deletion vectors
record which rows are gone rather than the rows. The first commit is
left out of the checks, since nothing comes before it. The memory
note above was "a few hundred MB"; the first run peaked at 2.8 GB.
The purge wrote each file with the types Parquet reads back (its time
in milliseconds); it now writes the table's own types. The note above
said it "takes under a minute"; the runs took 43 to 83 seconds.
Author: Roni Das
Created: 2026-09-30
"""
import hashlib
import io
import json
import sys
import numpy as np
import pandas as pd
import pyarrow as pa
import pyarrow.parquet as pq
import sklearn
from sklearn.datasets import fetch_openml
D0, WEEK = pd.Timestamp("2009-12-01"), pd.Timedelta(days=7)
AGAIN, REORDER, LATE, ERASE = 40, 60, 70, 95
SCHEMA = pa.schema([
("line", pa.int64()), ("Invoice", pa.string()),
("StockCode", pa.string()), ("Description", pa.string()),
("Quantity", pa.int64()), ("InvoiceDate", pa.timestamp("s")),
("Price", pa.float64()), ("Customer_ID", pa.float64()),
("Country", pa.string()), ("returned", pa.int64())])
CHANGE = SCHEMA.append(pa.field("deleted", pa.bool_()))
HASHED = ("month", "week", "changes")
EMPTY = np.array([], np.int64)
# ───────── the data ─────────
raw = fetch_openml(data_id=43368, as_frame=True, parser="auto").frame
when = pd.to_datetime(raw["InvoiceDate"])
is_c = raw["Invoice"].str.startswith("C").to_numpy()
COL = {"line": np.arange(len(raw), dtype=np.int64),
"Quantity": raw["Quantity"].to_numpy(np.int64),
"InvoiceDate": when.to_numpy().astype("datetime64[s]"),
"Price": raw["Price"].to_numpy(np.float64),
"Customer_ID": raw["Customer_ID"].to_numpy(np.float64),
"returned": np.zeros(len(raw), np.int64)}
for c in ("Invoice", "StockCode", "Description", "Country"):
v = raw[c].to_numpy(dtype=object)
COL[c] = np.where(pd.isna(v), None, v) # missing text: None
week_of = ((when - D0) // WEEK).to_numpy().astype(int)
month_of = when.dt.strftime("%Y-%m").to_numpy()
n_weeks = int(week_of.max()) + 1
TAGS = [w for w in range(n_weeks) if (w + 1) % 13 == 0]
cust, code = COL["Customer_ID"], COL["StockCode"]
# Each cancellation, matched to the most recent earlier sale line by the
# same customer for the same product. A sale comes first in a tie.
ev = pd.DataFrame({"t": when, "c": is_c.astype(int), "i": COL["line"]})
ev = ev.sort_values(["t", "c", "i"], kind="stable")
last, returns = {}, []
for i, c in zip(ev["i"].to_numpy(), ev["c"].to_numpy()):
if cust[i] != cust[i]: # NaN: no customer, nothing to match
continue
if not c:
last[(cust[i], code[i])] = i
elif (cust[i], code[i]) in last:
returns.append((int(i), int(last[(cust[i], code[i])])))
by_week = {}
for c, s in returns:
by_week.setdefault(int(week_of[c]), []).append(
(s, int(-COL["Quantity"][c])))
# ───────── files ─────────
def arrow(ids, ret=None, deleted=None):
"""The rows `ids`, in that order, as an Arrow table. In a change file
a deleted line keeps only its id; the rest of its row is left empty."""
cols = {}
for f in SCHEMA:
v = ret if (f.name == "returned" and ret is not None) \
else COL[f.name][ids]
blank = None if (deleted is None or f.name == "line") else deleted
cols[f.name] = pa.array(v, type=f.type, from_pandas=True, mask=blank)
if deleted is None:
return pa.table(cols, schema=SCHEMA)
cols["deleted"] = pa.array(deleted, type=pa.bool_())
return pa.table(cols, schema=CHANGE)
def parquet(table):
buf = io.BytesIO()
pq.write_table(table, buf)
return buf.getvalue()
def row_hash(table):
"""One 64-bit hash per row, made from its values alone."""
h = np.zeros(table.num_rows, np.uint64)
for f in SCHEMA:
c = table.column(f.name)
if pa.types.is_timestamp(f.type): # Parquet keeps it in ms
c = c.cast(pa.timestamp("s"))
if pa.types.is_integer(f.type) or pa.types.is_timestamp(f.type):
c = c.cast(pa.int64()).fill_null(0)
v = c.to_numpy(zero_copy_only=False)
h = (h * np.uint64(1000003)) ^ pd.util.hash_array(v)
return h
def content(hashes):
"""A hash of the rows that ignores their order."""
return hashlib.sha256(np.sort(hashes).tobytes()).hexdigest()
class Store:
"""Files kept under the SHA-256 of their bytes, each stored once."""
def __init__(self):
self.blob, self.new = {}, 0
def put(self, table):
data = parquet(table)
key = hashlib.sha256(data).hexdigest()
if key not in self.blob:
self.blob[key] = data
self.new += len(data)
return key
# ───────── the live table, and how each layout keeps it ─────────
present = np.zeros(len(raw), bool)
ROWH = np.zeros(len(raw), np.uint64)
sorted_now = False # set by the reorder commit
parts = {"month": {}, "week": {}} # file name -> its line ids, in order
chg = {} # changes: name -> (ids, returned then, deleted or None)
store = {k: Store() for k in HASHED}
man = {k: {} for k in HASHED} # the files of the current version
V = []
def perm(ids):
"""Rows are written as they came, or sorted by customer then product
once the reorder commit has happened."""
if not sorted_now or len(ids) == 0:
return np.arange(len(ids))
return np.lexsort((ids, code[ids].astype(str),
np.nan_to_num(cust[ids], nan=1e9)))
def write_part(layout, name):
ids = parts[layout].get(name, EMPTY)
ids = ids[present[ids]]
ids = ids[perm(ids)]
if len(ids):
parts[layout][name] = ids
man[layout][name] = store[layout].put(arrow(ids))
else:
parts[layout].pop(name, None)
man[layout].pop(name, None)
def write_change(name, ids, deleted=None):
p = perm(ids)
ids = ids[p]
ret = COL["returned"][ids].copy() if name not in chg else chg[name][1][p]
dele = None if deleted is None else deleted[p]
chg[name] = (ids, ret, dele)
man["changes"][name] = store["changes"].put(arrow(ids, ret, dele))
def commit(name, week, add=EMPTY, fix=(), erase=EMPTY, again=False):
"""One version: append `add`, apply the returns `fix`, delete `erase`."""
before = {k: store[k].new for k in HASHED}
rows_before = int(present.sum())
present[add] = True
new, old, skipped = set(add.tolist()), set(), 0
for s, q in fix:
if not present[s]: # its line was erased
skipped += 1
continue
COL["returned"][s] += q
if s not in new:
old.add(s)
present[erase] = False
touched = np.array(sorted(new | old | set(erase.tolist())), np.int64)
for m in np.unique(month_of[add]):
parts["month"][m] = np.concatenate(
[parts["month"].get(m, EMPTY), add[month_of[add] == m]])
if len(add):
parts["week"][week] = add
for layout, of in (("month", month_of), ("week", week_of)):
todo = parts[layout] if again else set(of[touched].tolist())
for n in sorted(todo):
write_part(layout, n)
# changes: new lines in a new file; changed old lines in a small one
if len(add):
write_change(f"w{week:03d}", add)
oc = np.array(sorted(old | set(erase.tolist())), np.int64)
if len(oc):
write_change(f"c{len(V):03d}", oc, np.isin(oc, erase))
if again: # rewrite every file as it stands
for n, (ids, ret, dele) in list(chg.items()):
write_change(n, ids, dele)
keep = touched[present[touched]]
if len(keep):
ROWH[keep] = row_hash(arrow(keep))
live = np.flatnonzero(present)
first = store["month"].blob[next(iter(man["month"].values()))]
V.append({
"name": name, "week": week, "rows": int(len(live)),
"rows_before": rows_before, "added": int(len(add)),
"returns": len(fix), "skipped": skipped,
"old_changed": len(old), "content": content(ROWH[live]),
"schema": str(pq.read_schema(io.BytesIO(first))),
"files": {k: dict(man[k]) for k in HASHED},
"new": {k: store[k].new - before[k] for k in HASHED}})
# ───────── the 110 commits ─────────
erased = EMPTY
for w in range(n_weeks):
add, fix = np.flatnonzero((week_of == w) & ~is_c), by_week.get(w, [])
if w == LATE:
commit(f"week {w} lines", w, add=add)
commit(f"week {w} returns", w, fix=fix)
else:
commit(f"week {w}", w, add=add, fix=fix)
if w == AGAIN:
commit("written again", w, again=True)
if w == REORDER:
sorted_now = True
commit("another order", w, again=True)
if w == ERASE:
live = np.flatnonzero(present)
who = np.unique(cust[live][~np.isnan(cust[live])])
erased_customer = who[np.random.default_rng(0).integers(len(who))]
erased = live[cust[live] == erased_customer]
commit("erasure", w, erase=erased)
size = {k: {h: len(b) for h, b in store[k].blob.items()} for k in HASHED}
copy = [sum(size["month"][h] for h in v["files"]["month"].values())
for v in V]
TAGI = [[v["name"] for v in V].index(f"week {t}") for t in TAGS]
def read(v, layout):
"""Version v read back from its files: {line: (row hash, returned)}.
Base files first, then change files in the order they were written."""
got = {}
for n in sorted(v["files"][layout], key=lambda n: (str(n)[0] == "c", str(n))):
t = pq.read_table(io.BytesIO(store[layout].blob[v["files"][layout][n]]))
dele = t.column("deleted").to_numpy() if "deleted" in t.column_names \
else np.zeros(t.num_rows, bool)
for i, h, r, d in zip(t.column("line").to_numpy().tolist(),
row_hash(t).tolist(),
t.column("returned").fill_null(0).to_numpy().tolist(),
dele.tolist()):
got[i] = None if d else (h, r)
return {i: x for i, x in got.items() if x is not None}
# ───────── what it cost, and what each check saw ─────────
age = np.array([(when[c] - when[s]).total_seconds() / 86400
for c, s in returns])
out = {"versions_of": {"scikit_learn": sklearn.__version__,
"pandas": pd.__version__, "pyarrow": pa.__version__},
"lines": len(raw), "sale_lines": int((~is_c).sum()),
"cancellations": int(is_c.sum()), "returns_matched": len(returns),
"returns_changed_nothing": int(is_c.sum()) - len(returns),
"sale_lines_returned": len({s for _, s in returns}),
"weeks": n_weeks, "commits": len(V), "tags": TAGS,
"return_age_days": {
"median": float(np.median(age)), "max": float(age.max()),
"over_7": float((age > 7).mean()),
"over_28": float((age > 28).mean()),
"over_90": float((age > 90).mean()),
"earlier_month_file": float(np.mean(
[month_of[c] > month_of[s] for c, s in returns]))}}
distinct = {tuple(sorted(v["files"]["month"].items())): b
for v, b in zip(V, copy)}
out["stored"] = {"copy": sum(copy), "copy_hashed": sum(distinct.values()),
**{k: store[k].new for k in HASHED}}
out["last_version"] = {
"rows": V[-1]["rows"], "bytes": copy[-1],
"files": {k: len(V[-1]["files"][k]) for k in HASHED},
"changes_rows_read": sum(len(chg[n][0]) for n in V[-1]["files"]["changes"])}
def kind(v):
n = v["name"]
return "weekly" if n.split()[-1].isdigit() else {
"written again": "again", "another order": "reorder",
f"week {LATE} lines": "late_lines",
f"week {LATE} returns": "late_returns"}.get(n, n)
checks, cost, rewrites = {}, {}, []
for k in range(1, len(V)): # the first has nothing before it
v, p = V[k], V[k - 1]
fired = {"rows": v["rows"] != p["rows"],
"schema": v["schema"] != p["schema"],
"files": v["files"]["month"] != p["files"]["month"],
"content": v["content"] != p["content"]}
c = checks.setdefault(kind(v), {"commits": 0, **dict.fromkeys(fired, 0)})
c["commits"] += 1
for f_, yes in fired.items():
c[f_] += int(yes)
b = cost.setdefault(kind(v), {x: [] for x in ("copy",) + HASHED})
b["copy"].append(copy[k])
for x in HASHED:
b[x].append(v["new"][x])
if kind(v) == "weekly":
rewrites.append(sum(p["files"]["month"].get(n) != h
for n, h in v["files"]["month"].items()))
out["checks"] = checks
out["cost"] = {k: {x: int(sum(b[x])) for x in b} for k, b in cost.items()}
out["weekly_median"] = {x: float(np.median(cost["weekly"][x])) for x in cost["weekly"]}
out["weekly_empty"] = sum(v["added"] == 0 and v["returns"] == 0
for v in V if kind(v) == "weekly")
lr = V[[v["name"] for v in V].index(f"week {LATE} returns")]
out["late_returns"] = {"lines_changed": lr["old_changed"], "rows": lr["rows"],
"rows_before": lr["rows_before"]}
out["month_files_rewritten_a_week"] = {
"median": float(np.median(rewrites)), "max": int(max(rewrites)),
"share_more_than_1": float(np.mean(np.array(rewrites) > 1))}
out["same_bytes_written_twice"] = all(
V[[v["name"] for v in V].index("written again")]["new"][x] == 0
for x in HASHED)
# every tag read back from every store; then "as of" against the last one
ok, asof = 0, []
for t, k in zip(TAGS, TAGI):
for x in HASHED:
got = read(V[k], x)
ok += content(np.array([g[0] for g in got.values()], np.uint64)) \
== V[k]["content"]
if x != "month":
continue
ids = np.array(sorted(got), np.int64)
h = np.array([got[i][0] for i in ids.tolist()], np.uint64)
r = np.array([got[i][1] for i in ids.tolist()], np.int64)
end = np.datetime64(D0 + (t + 1) * WEEK, "s")
live = np.flatnonzero(present & (COL["InvoiceDate"] < end))
both = np.intersect1d(ids, live)
pos = np.searchsorted(ids, both)
asof.append({
"week": t, "tag_lines": int(len(ids)), "live_lines": int(len(live)),
"only_in_tag": int(len(np.setdiff1d(ids, live))),
"only_live": int(len(np.setdiff1d(live, ids))),
"changed": int((h[pos] != ROWH[both]).sum()),
"returned_changed": int((r[pos] != COL["returned"][both]).sum()),
"units_returned_tag": int(r.sum()),
"units_returned_live": int(COL["returned"][live].sum())})
out["tags_read_back"], out["tags_read_back_of"] = int(ok), 3 * len(TAGS)
out["as_of"] = asof
# retention: which stored files each rule keeps, and who is still in them
n = len(V)
POLICY = {"all": list(range(n)), "last4": list(range(n - 4, n)),
"last4_tags": sorted(set(range(n - 4, n)) | set(TAGI))}
holds = {x: {} for x in HASHED}
for x in HASHED:
for key, blob in store[x].blob.items():
t = pq.read_table(io.BytesIO(blob))
ln = t.column("line").to_numpy()
if "deleted" in t.column_names: # a marker holds only an id
ln = ln[~t.column("deleted").to_numpy()]
holds[x][key] = bool(np.isin(ln, erased).any())
def purge(x, keep):
"""Rewrite every kept file that holds an erased line, without it."""
kept = set().union(*(V[k]["files"][x].values() for k in keep))
ps, files = Store(), 0
for h in sorted(kept):
if holds[x][h]:
t = pq.read_table(io.BytesIO(store[x].blob[h]))
m = ~np.isin(t.column("line").to_numpy(), erased)
if "deleted" in t.column_names: # markers can stay
m |= t.column("deleted").to_numpy()
files += 1
if m.any(): # back to the table's own types
t = t.filter(pa.array(m))
ps.put(t.cast(CHANGE if "deleted" in t.column_names else SCHEMA))
return {"files": files, "bytes": ps.new}
ret = {}
for pol, keep in POLICY.items():
r = {"versions": len(keep), "copy": {
"bytes": sum(copy[k] for k in keep),
"tags": sum(k in keep for k in TAGI),
"versions_holding": sum(any(holds["month"][h] for h in
V[k]["files"]["month"].values())
for k in keep)}}
for x in HASHED:
kept = set().union(*(V[k]["files"][x].values() for k in keep))
share = [sum(size[x][h] for h in V[k]["files"][x].values() if h in kept)
/ sum(size[x][h] for h in V[k]["files"][x].values())
for k in TAGI]
r[x] = {"bytes": sum(size[x][h] for h in kept), "files": len(kept),
"tags": sum(s == 1.0 for s in share), "tag_share_kept": share,
"files_holding": sum(holds[x][h] for h in kept),
"bytes_holding": sum(size[x][h] for h in kept if holds[x][h]),
"versions_holding": sum(any(holds[x][h] for h in
V[k]["files"][x].values())
for k in keep)}
ret[pol] = r
out["retention"] = ret
ek = [v["name"] for v in V].index("erasure")
out["erasure"] = {
"customer": float(erased_customer), "lines": int(len(erased)),
"month_files": int(len(set(month_of[erased]))),
"week_files": int(len(set(week_of[erased].tolist()))),
"new_bytes": V[ek]["new"],
"returns_skipped_after": sum(v["skipped"] for v in V),
"last_version_files_holding": {
x: sum(holds[x][h] for h in V[-1]["files"][x].values()) for x in HASHED},
"purge_last4_tags": {x: purge(x, POLICY["last4_tags"]) for x in HASHED}}
out["series"] = [{"name": v["name"], "rows": v["rows"], "copy": b,
"new": v["new"], "content": v["content"][:16],
"files": {x: len(v["files"][x]) for x in HASHED}}
for v, b in zip(V, copy)]
# ───────── print ─────────
mb = lambda b: f"{b / 1e6:,.1f}" # noqa: E731
vv, a, lv = out["versions_of"], out["return_age_days"], out["last_version"]
print(f"scikit-learn {vv['scikit_learn']}, pyarrow {vv['pyarrow']}")
print(f"sale lines {out['sale_lines']:,}; "
f"cancellations {out['cancellations']:,}")
print(f"returns matched {len(returns):,}; "
f"changed nothing {out['returns_changed_nothing']:,}")
print(f"a return reaches back: median {a['median']:.1f} days")
print(f" over 28 days {a['over_28']:.1%}, over 90 {a['over_90']:.1%}")
print(f" into an earlier month's file: {a['earlier_month_file']:.1%}")
print(f"commits {n}, over {n_weeks} weeks; tags {len(TAGS)}")
print(f"last version: {lv['rows']:,} lines, {mb(lv['bytes'])} MB")
print(f"stored for all {n} versions, MB:")
for x in ("copy", "copy_hashed") + HASHED:
print(f" {x:<12}{mb(out['stored'][x]):>10}")
print("files read for the last version:")
print(" " + ", ".join(f"{x} {lv['files'][x]}" for x in HASHED))
print("checks that fired, by kind of commit:")
print(f" {'kind':<13}{'n':>4}{'rows':>6}{'schema':>7}{'files':>6}{'content':>8}")
for k in ("weekly", "again", "reorder", "late_returns", "erasure"):
c = checks[k]
print(f" {k:<13}{c['commits']:>4}{c['rows']:>6}{c['schema']:>7}"
f"{c['files']:>6}{c['content']:>8}")
print(f"late returns: {lr['old_changed']:,} lines changed, "
f"rows {lr['rows']:,}")
print(f"written twice, same bytes: {out['same_bytes_written_twice']}")
print(f"tags read back exactly: {ok} of {3 * len(TAGS)}")
print("as of: the last version filtered by date, vs the tag")
for r in asof:
print(f" week {r['week']:>3}: {r['changed'] + r['only_in_tag']:>5,} of "
f"{r['tag_lines']:>9,} lines differ")
e = out["erasure"]
print(f"erased customer {e['customer']:.0f}: {e['lines']:,} lines")
print(f" in {e['month_files']} month and {e['week_files']} week files")
for pol in POLICY:
print(f"keep {pol}: MB kept, tags kept, holding erased")
r = ret[pol]
print(f" copy {mb(r['copy']['bytes']):>8}{r['copy']['tags']:>4}/8"
f"{r['copy']['versions_holding']:>5} versions")
for x in HASHED:
print(f" {x:<9}{mb(r[x]['bytes']):>8}{r[x]['tags']:>4}/8"
f"{r[x]['files_holding']:>5} files")
print("purge, keeping the last 4 and tags: MB written")
for x in HASHED:
p = e["purge_last4_tags"][x]
print(f" {x:<9}{p['files']:>5} files{mb(p['bytes']):>9}")
if len(sys.argv) > 1: # a file name was given: save every number too
json.dump(out, open(sys.argv[1], "w"), indent=1)

The report lives in scripts/labs/dataeng/versioning_report.py. It does not trust the demo's code. The demo keeps each column as one numpy array and matches returns in one pass through the events. The report keeps each line as a plain Python record and matches each cancellation with a binary search over that customer's sales of that product. It builds every file from Python lists and writes it again.
Then it compares. For all 110 versions and all three stores kept by hash, the new bytes and the file counts must come back exactly. It hashes rows its own way, and the versions where the rows changed must be the same versions. The as-of differences, the retention totals and the purge must come back too, and so must DVC's count. 928 checks agreed. It stops at the first one that does not. I tested that by changing one stored number, which made it stop.
Its json mode writes results/pin-report.json, which the figures read. Its demo mode checks the demo's printed run line by line. Its box mode writes the playground on the next slide and checks it.
What came when: the demo's design, in its docstring, came before any run. The 30 erasures, the reach in weeks and the split of the month store's bytes came after I saw the results. I wrote them into the report's docstring before the report first ran.
Four things changed after the demo's first run, and its docstring lists them. In the changes store, a deleted line had been written whole into the change file. That stored the erased customer's data again, so a deleted line now keeps only its id. Also, the first commit is left out of the checks, the memory note was corrected, and the purge now writes the table's own column types.
This box has no data in it, only sizes. It holds every file the three stores kept by hash, as a size in bytes. That is 816 files for the month store, 1,880 for the week store and 329 for the changes store. For each of the 110 versions, it lists which files changed. Each number is written in base 64, counting with 64 digits instead of 10, to keep the box small. It runs in your browser.
As it is, the box prints each store's total and then applies the lesson's three rules. It keeps all 110 versions, then the last 4, then the last 4 and the 8 tags. The report checked that every total it prints matches pin-report.json exactly. It prints 431.9 MB kept in 816 month files for all versions, 37.8 MB for the last 4, and 82.0 MB when the tags are kept too.
Then try your own rule. keep(range(VERSIONS - 13, VERSIONS)) keeps the last 13 versions, about 13 weeks, a quarter of a year. keep(TAGS) keeps only the tagged versions. keep(range(0, VERSIONS, 4)) keeps every fourth version. Ask how many tags each rule keeps readable, and what that costs in each store.
The data. fetch_openml(data_id=43368) downloads Online Retail II once and reads the local copy after that. Each column becomes one numpy array over all 1,067,371 lines, so a line's id is its position. week_of and month_of say which week and month file a line belongs to.
Matching the returns. The events are sorted by time, with a sale before a cancellation stamped with the same minute. The loop remembers, for each customer and product, the most recent sale line. A cancellation that finds one becomes a return: that line's id and the units sent back.
arrow, parquet, row_hash, Store. arrow turns a list of line ids into an Arrow table with one fixed schema, and parquet writes it with pyarrow's default settings. row_hash makes one 64-bit number per row from its values alone. Store keeps each file under the SHA-256 of its bytes, and counts only the bytes it has not seen before.
commit. One version: append the week's lines, apply its returns, delete any erased lines. The month and week stores rewrite every file that holds a touched line. The changes store writes the new lines into a new week file, and the changed old lines into one small change file. It records each version's files, its row hash and the bytes it added.

Cut the table into files the way it changes. New data should go into new files. If corrections land in old files, write them as small change files, or use your table format's merge-on-read setting. Here that choice decided most of the bill. Then compact the change files now and then: here the last version took 210 files to read in the changes store, against 25 as month files.
Store files by hash, and write them the same way each time. Then a job that runs twice costs nothing. Check that your writer gives the same bytes for the same table.
Tag the version each model trains on. A Git tag over a .dvc file, a lakeFS tag, an Iceberg tag. Never a date filter on a table corrected in place, never latest.
Record the tag, a hash of every row, the commit and the seed. Next to the model, in the tracking tool. Do not rely on MLflow's dataset digest alone.
Set retention that keeps the tags. An Iceberg tag with no RETAIN, or a RETAIN longer than the model's life. For Delta, a copy of the training rows, or both of its retention windows raised.
On an erasure, rewrite the kept files too, then delete the old ones, and note it. Include files that hold rows marked as deleted. Then physically delete the old files, and move each tag to its rewritten version. Record that the tag's hash changed, and why.

Version it when a model's decisions may be questioned later. The EU's AI Act asks providers of high-risk systems for technical documentation. Where relevant, it includes "datasheets describing the training methodologies and techniques and the training data sets used", with their "provenance". A description is easier to write when you can still read the data.
Version it when old rows get corrected. Here, a return reached back a median of 10.1 days, and today's table disagreed with the week-51 tag on 953 lines.
Version it when you compare models trained months apart, or when a person's data may have to be erased from the past as well as the present.
Skip it for a small file that never changes. Git holds it as it is.
Skip it when the data can be rebuilt exactly from pinned inputs and pinned code. Then pin the recipe instead. And if you only need to prove later that the data did not change, a hash of every row is much smaller than the rows.

One shop, two years. A shop whose returns reach further back, or less far, would get different numbers. So would a table that is corrected more often, or never.
Returns matched by one rule. The data has no link from a cancellation to its sale. I matched each one to the most recent earlier sale by the same customer for the same product. The shop's own records might say otherwise.
My stores, not the real formats. Delta, Iceberg and lakeFS add logs and metadata, and aim for files of many megabytes. My week files are small. With bigger files, each copy-on-write rewrite costs more. Only the month store was checked against a real tool, DVC.
What came after the results. The 30 erasures, the reach in weeks and the split of the month store's bytes came after I saw the results. So did the four changes listed in the demo's docstring.

Take the model your team runs today and ask the first question: can you read the exact rows it trained on? If the honest answer is "we would filter the table by date", you have the week-51 problem from this lesson.
Then look at where your corrections land. If they rewrite old files, measure what a month of versions costs you now. Here, the same 110 versions cost 431.9 MB as month files and 27.8 MB as week files with change files.

The card keeps the lesson's two numbers. Today's table, filtered by date, differed from the week-51 tag on 953 of 490,766 lines. And how the table was cut into files decided whether keeping every version cost 431.9 MB or 27.8 MB.
The next lesson in this chapter is about data labeling: where labels come from, and how they go wrong.
4 questions - Score 80% to pass
In the lab, why did the month store keep 431.9 MB when the change-file store kept 27.8 MB for the same 110 versions?
134 lines changed in the late-returns commit. Which checks fired on it?
A team filters today's table to the dates before week 51 to get the week-51 model's training data. What did the lab find?
With mlflow-skinny 3.16.1, what happened to MLflow's dataset digest after a real return was applied at row 11,006?
dvc pullverifyThe other tools work in whole files too, each in its own way. lakeFS says that "when you update an object lakeFS creates a new physical address for that version". Delta Lake's documentation says that by default "when a single row in a data file is deleted, the entire Parquet file containing the record must be rewritten". Rewriting a file to change one row is called copy-on-write. Delta can avoid it with deletion vectors, small files that mark rows as removed. Iceberg calls the same idea merge-on-read: the change is written into a small new file, and readers apply it. Iceberg's setting write.update.mode is copy-on-write by default.
So the question that decides the bill is not "how often does the data change?" It is "where in the files do the changes land?" The lab measures exactly that.
lakectllakectl branch create lakefs://ml-lake/relabel-42 --source lakefs://ml-lake/main
# write the relabelled files to the branch through lakeFS's S3 gateway
aws s3 --endpoint-url https://lakefs.internal cp labels.parquet s3://ml-lake/relabel-42/labels/labels.parquet
lakectl commit lakefs://ml-lake/relabel-42 -m "Relabel 12k fraud rows"
lakectl merge lakefs://ml-lake/relabel-42 lakefs://ml-lake/main
lakectl tag create lakefs://ml-lake/train-week-51 lakefs://ml-lake/main
Both leave a training run with something exact to record: a Git tag and a .dvc hash, or a lakeFS tag. The word to avoid in a training config is latest. It is a name that points somewhere different next week.
VACUUM "is not triggered automatically". But when it runs, it deletes files the table stopped using more than 7 days ago, so older versions can no longer be read. The documentation says: "The ability to time travel back to a version older than the retention period is lost after running vacuum." Its promise of 30 days holds only "unless you have: Run VACUUM on your Delta table". Iceberg's expire_snapshots removes snapshots older than 5 days when you run it, and "will never remove files which are still required by a non-expired snapshot".
For a model's training data, the fix differs by format. Iceberg has tags, and its documentation says "the default retention for branches and tags is forever". For Delta, I found no way to spare one version from VACUUM.
Delta's documentation suggests a shallow clone to archive a training version: a new table that points at another table's files. This is a different thing from Git's shallow clone above, which fetches fewer commits. The documentation also warns: "If you run vacuum on the source table, clients will no longer be able to read the referenced data files." So for Delta, copy the training rows into a table of their own. Or raise both delta.deletedFileRetentionDuration and delta.logRetentionDuration to cover the longest audit you expect.
ORDER BYThe tags. Every 13 weeks, at weeks 12, 25, 38, 51, 64, 77, 90 and 103, a model would be trained, so that version is tagged. No model is trained in this lab. The tags stand for the versions a real team would have to keep.

The four stores keep the same 110 versions. Copy is the habit of copying the whole table into a new folder for every version. Month keeps one Parquet file per month of sale, stored by hash, the way DVC keeps a folder of month files. Week keeps one file per week of arrival, so new lines always go into a new file, and a change rewrites the file that holds the line: copy-on-write. Changes keeps the week files, and writes each commit's changed lines, whole, into one small new file: merge-on-read.
These are my own stores, built to cut the table the way each tool does. They are not Delta, Iceberg or lakeFS, which add logs, metadata and much larger files. For the month store, real DVC checked my count exactly.
OPTIMIZESo the same history cost 27 times as much to store in one layout as in another, and less to read in the other. The next slide shows where the month store's bytes went.
Then I tried a real tool's check. MLflow, a common tool for tracking training runs, computes a digest for a dataset you log, "a unique hash/fingerprint for dataset identification". Its source code builds a pandas table's digest from df.head(10000), its row count and its column names. It keeps only number columns, and text columns with no missing value anywhere. I ran mlflow-skinny 3.16.1, in its own environment, on the sale lines as they were and with four kinds of change.
A return at row 11,006 left the digest the same. A return at row 258 changed it. Moving row 0's date by a day left it the same, because date columns are not in it. Sorting the rows, with no value changed, changed it. The row count and the columns stayed the same all four times. My row hash changed for both returns and the date, and not for the sort.
After a review, I added one more test: change row 0 of each column in turn, and watch the digest. It ignored Description, a text column with 4,474 missing values, and the date. It still saw Customer_ID, a number column, even with 242,257 missing values.
The lineage lesson found the same thing with row numbers: "the last 30,000 rows" meant different rows a month later. A position, or a date on a table corrected in place, can point at different rows later. A tag does not.
One draw is one customer, so after the results I drew 30, with seeds 0 to 29, and replayed only the erasure commit for each.

Over the 30 customers, the erasure wrote 2,147 KB on average in the month store, from 390 to 6,688 KB. The week store averaged 670 KB, and the changes store 3 KB. The customers held from 1 to 577 lines, and they were in 4.9 of the 7 earlier tags on average, at most all 7.
The number of lines did not decide the cost. The largest month bill, 6,688 KB, came from a customer with 64 lines spread over 12 months. The customer with 577 lines wrote 6,130 KB. What mattered was how many files the person was in.
A purge changes the tagged versions, so their row hashes no longer match the hash recorded with the model. Write down that the change was an erasure, and when. I am not a lawyer, and this is not legal advice. How erasure applies to kept training copies is a question for the person who answers for your data protection.
This is a real run in VS Code's terminal (python versioning_demo.py).

When I ran it, it printed the versions of scikit-learn and pyarrow, the counts, the five stored totals and the checks. All of it matches the stored pin-demo.json, and the longest printed line was 52 characters. The report's demo mode checks those lines against the file.
To try something I have not run, change the 95 in AGAIN, REORDER, LATE, ERASE = 40, 60, 70, 95 to 30, so the customer is erased in week 30 and fewer tags come before it. I cannot tell you what it prints, because I have not run it. The question to ask is how the purge's bytes change.

scikit-learn only downloads the data here; nothing is trained. pandas turns the dates into dates. pyarrow writes every Parquet file, so pyarrow's own settings decide the bytes. DVC and MLflow each ran once, in their own environments.
The commits. One loop over the 106 weeks, with the four extra commits placed by the constants AGAIN, REORDER, LATE and ERASE.
The results. read reads a version back from its files. The rest compares each commit with the one before, applies the three retention rules, runs the purge, and prints. With a file name on the command line, json.dump saves every number.