Before lakehouses, the common answer was to run both a lake and a warehouse. Raw data landed in the lake. A copy job then moved the clean part into the warehouse for dashboards. The 2021 paper calls this the two-tier architecture, and says data is "first ETLed into lakes, and then again ELTed into warehouses".
This has real costs. There are two copies of the truth, and keeping them in step "is difficult and costly", in the paper's words. The warehouse copy goes stale. And machine learning tools such as TensorFlow, PyTorch and XGBoost read data with code, not SQL, so they cannot open the warehouse's internal format. A data scientist either exports from the warehouse, which is a third copy, or reads the raw lake, which has no checks.
A plain lake has its own problems. The 2020 Delta Lake paper lists them. A job that changes a table writes many files one by one, so readers can see half a change. If the job crashes halfway, the table is left broken. And nothing checks a file's columns, so a bad file lands quietly. People call a lake in this state a data swamp.
The lakehouse answer is to keep one copy, in open files, and fix those problems with the log. Read the paper with one caution: its authors sell lakehouses, so it is their case for the idea, not a neutral survey.