A data lakehouse stores data as open-format files in low-cost object storage, as a data lake does, and adds a table layer that provides transactions, schema enforcement and indexing. SQL engines can then query those files with the reliability expected of a data warehouse.
The three layers
Object storage holds the files, most often in a columnar format such as Parquet. An open table format keeps the metadata that says which files make up a table, with its schema and snapshots, and that is what gives plain files transactions, schema changes and time travel. Apache Iceberg is one such format. Engines on top, for SQL or other processing, read and write the tables.
Because the format is open and the files are yours, several engines can use the same tables. A data warehouse, by comparison, has traditionally stored data in its own managed format behind one engine.
Landing vendor files
In a lakehouse, vendor files land first in a raw zone of your own storage. Keep them unchanged there, with their manifest. Convert them to a columnar format, write them into tables in your table format, and partition daily tables by day so that queries read only the days they need. The guide to Apache Iceberg tables for licensed external data works through the steps.
Time travel is not point-in-time data. A table format lets you query a table as it stood at an earlier snapshot, which tells you what your copy held then. It does not tell you when the world first observed each fact. Point-in-time analysis needs the observation time in the data itself, and a source that does not revise it afterwards.
In Fokals data
Fokals is delivered direct, as bulk exports in CSV, JSON or JSON Lines, which you land in your own storage and write into tables in your table format with the lakehouse's own loader. Every record carries the time it was observed, and daily and weekly datasets are written once and never revised, so a closed period is loaded once and appended. For the exports see the delivery page, and for the grain of each table the data dictionary.
Related terms
- Data sharing: read access to a provider's tables without a copy.
- Bulk export: how the files reach your storage.
- JSON Lines: the line-delimited format in which exports can arrive.
Frequently asked questions
What is the difference between a data lakehouse and a data warehouse?
A warehouse has traditionally stored data in its own managed format behind one engine. A lakehouse keeps data in open files in object storage and adds the table layer on top, so several engines can read the same tables. Both support SQL and transactions. The difference is where the data lives and how open its format is.
What is an open table format?
An open table format is a published specification for tracking which files make up a table, with its schema, partitions and snapshots, so that different engines can read and write the same table safely. Apache Iceberg is an example. It is the layer that gives files in object storage their transactions and time travel.
Can I load JSON Lines or CSV files into a lakehouse?
Yes. Land the files in object storage, read them with an engine that supports the format, and write them out as a table in your chosen table format, usually over a columnar file format such as Parquet. Partition daily tables by day. Keep the original files and any manifest in a raw zone so that you can reload.