Data Engineering
Many companies keep their data in two places. A warehouse holds clean tables for reports. A lake holds cheap files that anyone can read. Jobs copy data from one to the other every night, and when a number is wrong, someone has to find out which copy is wrong. A lakehouse tries to give you one place: the cheap open files of a lake, with the safe tables of a warehouse.
This topic explains how that works, one idea at a time, and then tests each idea on real data. Every lesson uses real New York taxi trips, millions of rows, and real tools: Delta Lake, Apache Iceberg, Apache Hudi, DuckDB and Unity Catalog. You will see what a reader sees while a writer is busy, what each table format writes to disk, what happens when two writers change one table at the same time, and what a catalog does that a plain folder cannot. Each lesson explains new words as they come up, so you do not need a data engineering background to start.
4 lessons in this topic · start with any of them
This page is the free part.
The course goes deeper on Data Engineering
₹499 in India$49 everywhere elseonce, for the whole course
The System Design course covers Data Engineering across a run of lessons, not one page. These 4 alone are about 282 minutes of step-by-step reading, every one with code you run in the browser, all with a quiz.
- Lakehouse Architecture: What a Folder of Parquet Files Gets Wrong
What is a lakehouse, and what breaks if a table is just a folder of Parquet files? Warehouse, lake and lakehouse in plain words first. Then a lab on 7,052,769 real NYC taxi trips: while a writer added February as 8 files, a reader that listed the folder saw every half-written state, 9 different row counts; a Delta Lake reader saw only 3,475,226 rows or 7,052,769, never anything between. During an overwrite, the folder reader counted 6,806,334 trips for one month that held 3,331,108. In a real race, 5 runs per case, the folder reader saw a half-written table 7 to 9 times per run and failed 955 to 5,003 reads; the Delta reader saw 0 half-written tables and 0 failures.
- Open Table Formats: What Delta Lake, Iceberg and Hudi Actually Write Down
Delta Lake, Apache Iceberg and Apache Hudi: what is actually in the metadata? I wrote the same 11,198,026 NYC taxi trips into all three, three commits each, and opened every metadata file. Delta kept 3 JSON commits, Iceberg 10 files in a four-level tree, Hudi a timeline and a metadata table. All three returned January's exact 3,475,226 rows when asked for the first version. After 120 small appends, an Iceberg planner opened 122 metadata files and a Delta reader 22; Hudi packed the rows into 1 file group instead of 120 files. Compaction fixed the count, and cleanup ended time travel.
- ACID on a Lakehouse: Two Writers, One Table, Who Wins
Two writers update the same Delta Lake table at once: who wins, and what does the loser see? I ran real writer processes against 3,475,226 NYC taxi trips, 5 runs per case. Two appends both committed every time. Two updates of different rows clashed every time, because they shared the same 9 files: one committed, the other got CommitFailedError. Splitting the table by the filter column let both commit. A check-then-append job started twice stored 317,150 trips for a day of 158,575 in 5 of 5 runs. In delta-rs 1.6.6, a WriteSerializable setting took effect only when spelled writeSerializable.
- Data Catalogs and Unity Catalog: What a Folder Cannot Do
What does a catalog like Unity Catalog do that a folder cannot? I put 3,475,226 NYC taxi trips behind three real catalogs: an Iceberg SQL catalog, the Apache Iceberg REST catalog and Unity Catalog OSS 0.6.0. A rename sent one SQL UPDATE and moved 0 of 9 files, while moving the folder broke a reader after 1 of 10 files. A writer that crashed before its swap left a newest metadata file with 123,646 trips nobody committed. Two catalogs on one folder both accepted writes and disagreed. Unity Catalog refused a user with no grants, but the same user read every row straight from the folder.
all part of
System Design Masterclass
770 lessons · about 283 hours · 18 free to read
₹499in India, by UPI
$49everywhere else, by PayPal
The concepts were explained in a clear and structured way, with practical examples that made even complex system design topics easier to understand. I especially liked the focus on real-world architecture, scalability, trade-offs. The content was well designed, engaging, and highly useful for anyone looking to strengthen their system design skills. Highly recommended for software engineers preparing for system design interviews or wanting to build a stronger foundation in designing scalable systems.
pay once,
yours for life
Warehouse, Lake and Lakehouse
A data warehouse is a database built to answer questions about history, like total sales per city per month. It is strict. You must describe every column before data goes in, and it keeps the data in its own private format. A data lake is the opposite: cheap storage, such as Amazon S3, full of files in open formats like Parquet. Nobody checks the columns on the way in, and any program can read the files.
In a lake, a table is often just a folder of Parquet files. A reader lists the folder and reads every file it finds. That is the weak point. If a writer is halfway through adding files, the reader sees a half-written table. In the first lesson, a reader that listed the folder while a writer added a month of trips saw 9 different row counts. A Delta Lake reader on the same data saw only the old total or the new total, never anything in between.
The fix is one extra idea: a list of files. Next to the data, the table keeps a transaction log, a small folder of notes that says which files belong to the table right now. A writer first writes its new files, then adds one note that names them. That one note is the commit, the single moment a change becomes real. Readers read the notes, not the folder, so they never see a change that is only half done.
Table Formats: Delta Lake, Iceberg and Hudi
Delta Lake, Apache Iceberg and Apache Hudi are three table formats. A table format is a set of rules for a table's metadata, the facts that describe the data: which files belong to the table, how many rows each file holds, and which version is the newest. All three keep the rows in ordinary Parquet files. They differ only in the metadata they keep beside them.
The second lesson writes the same taxi trips into all three and opens every metadata file. Delta keeps one folder of numbered commits. Iceberg keeps a tree of files that starts at a catalog. Hudi keeps a timeline of jobs and a small metadata table of its own. All three could give back an old version of the table exactly.
The differences show up when a pipeline commits often, a little data each time. After 120 small commits, Delta and Iceberg each held 120 small data files, and a reader had more metadata to open before it could plan a query. Hudi packed the same rows into one growing file. Compaction rewrites many small files into a few big ones without changing the rows. Cleanup then deletes the old files, and that is the moment you can no longer go back to old versions.
Two Writers, One Table
A database usually makes the second writer wait. It puts a lock on the rows, a mark that says busy, wait. A lakehouse table works another way, called optimistic concurrency control: every writer does its work first and checks for a clash only at the end, just before it commits. If nothing clashes, the change goes in. If something clashes, the writer gets an error and must try again.
The third lesson runs real writer processes against one Delta Lake table, 5 runs per case. Two writers that only add new rows both committed every time. Two updates of different rows still clashed every time, because their rows sat in the same files: one committed and the other got CommitFailedError. The table was never left half changed. Splitting the table into folders by the column the updates filter on let both updates commit.
The lesson also shows a mistake no setting can catch. A daily job checked whether a day was already loaded, then added it. Started twice at almost the same moment, both copies saw no rows, both added the day, and the table stored every trip of that day twice. The check ran outside the commit, so the table never knew about the rule. The lesson shows how to write the job so the table can catch the clash.
Catalogs and Unity Catalog
A catalog is a lookup from table names to table locations. You ask for taxi.trips, and the catalog tells you where the newest metadata lives. A folder cannot do this, because its path is both its name and its address. Change one and you change the other.
The fourth lesson puts the same trips behind three real catalogs: an Iceberg SQL catalog, the Apache Iceberg REST catalog and Unity Catalog OSS. Renaming a table changed one row in the catalog and moved no files. Moving the folder instead broke a reader that was halfway through. Two catalogs pointed at one folder both accepted writes and then gave two different row counts, so the rule is one table, one catalog.
Unity Catalog also keeps rules about people, called grants. It refused a user with no grants. But the same user could still read every row by opening the folder directly, because a catalog can only refuse requests that come to it. On real cloud storage you lock the storage itself and let the catalog hand out short-lived keys, which is called credential vending.
Frequently asked questions
Learn Data Engineering the interactive way
All 4 lessons with step by step diagrams, runnable code, and quizzes, plus the rest of the 770 lessons in the course. One payment, lifetime access, no subscription.
course 1
System Design Masterclass
From absolute beginner to principal engineer, drawn step by step.
- 770 interactive lessons
- Step-by-step system design diagrams
- Live code editors
- Quizzes with instant feedback
- Progress tracking and streaks
- Lifetime access and all future lessons