This comparison is for a team deciding where a vendor's company data should live when the choice is between two lakehouse platforms, Databricks and Microsoft Fabric, or when the organisation already runs both. It sets out how each builds a data lakehouse, how each shares data with another organisation and how each governs tables held under a licence, and it loads one file on each. Statements about the two platforms come from Databricks's and Microsoft's own documentation, read on 4 October 2026.
Fokals is delivered direct, by REST API and as bulk files, which you land in a Unity Catalog volume or in the Files folder of a lakehouse and load with the platform's own loader: Auto Loader on Databricks, a notebook on Fabric. The file used below is the daily export of Hiring Activity, the count of open, new and closed postings per company, which the data dictionary describes: CSV with one header row and one row per company and closed UTC day.
Two ways to build a lakehouse
Databricks describes its lakehouse as built on Apache Spark with two further parts: Delta Lake, a storage layer with ACID transactions and schema enforcement, and Unity Catalog, for governance. Delta Lake is open source. It extends Parquet data files with a file-based transaction log, and it is the default format for every table on Databricks. A table or a volume is either managed, where Unity Catalog also handles the lifecycle of the files, or external, where it governs data in storage you manage.
Fabric starts from the storage. OneLake is a single data lake that every Fabric tenant includes, built on Azure Data Lake Storage, with tables stored in Delta Parquet or Iceberg format. A tenant cannot create a second one, and there is no infrastructure to provision. Inside it, workspaces hold items. A lakehouse item has two top-level folders, Tables for managed Delta tables and Files for everything else. Spark notebooks write to it, and a SQL analytics endpoint reads its Delta tables with T-SQL, read-only. A warehouse is a separate item for teams that write with T-SQL.
Both therefore keep your data as Delta tables in open files. The difference is where the choices sit. On Databricks you decide, table by table, whether Unity Catalog manages the storage or governs storage you already have. On Fabric the lake is given, and the choices are which workspace and which item.
| Databricks | Microsoft Fabric | |
|---|---|---|
| Storage | Delta Lake tables, managed or external, under Unity Catalog | OneLake, one per tenant, on Azure Data Lake Storage |
| Table format | Delta Lake by default | Delta Parquet or Iceberg in OneLake, Delta by default in a lakehouse |
| Where a vendor's files land | A Unity Catalog volume | The Files folder of a lakehouse |
| File ingestion | Auto Loader, streaming tables, COPY INTO | File upload, pipelines, Dataflow Gen2, notebooks |
| SQL over the tables | Databricks SQL | A read-only SQL analytics endpoint on a lakehouse, T-SQL writes in a warehouse |
| Reading data held elsewhere | Lakehouse Federation, read-only, through foreign catalogs | Shortcuts and mirroring |
| Sharing with another organisation | OpenSharing, to Databricks or to other platforms | External data sharing, from one Fabric tenant to another |
| Catalogue and access | Unity Catalog: privileges, row and column filters, lineage, audit log, tags | OneLake catalog, workspace roles and OneLake security roles, with Microsoft Purview for labels and audit |
Loading one file on each
Databricks recommends the medallion pattern and Microsoft recommends it for Fabric: a bronze layer that keeps data as it arrived, then silver and gold layers of cleaned and curated tables. A vendor's file belongs in bronze first on either platform.
On Databricks the file goes into a Unity Catalog volume, and Auto Loader picks it up. Auto Loader records each file it has seen in a checkpoint, so a file is processed exactly once. In SQL it runs behind a streaming table, through the read_files function.
-- Databricks: a bronze table fed from a volume, with the source file kept on every row
-- (catalog, schema and volume names are illustrative)
create or refresh streaming table company_data.bronze.hiring_daily
schedule every 1 day
as select *, _metadata.file_name as source_file
from stream read_files(
'/Volumes/company_data/bronze/fokals/hiring_daily/',
format => 'csv'
);On Fabric the file goes into the Files folder of a lakehouse. Microsoft's page on ingestion options lists file upload for small files, pipelines with a copy activity, Dataflow Gen2 and notebooks, and describes a notebook as the most flexible of them, which suits an export fetched with an API key. Its notebook guide, dated 2024, shows the write used here.
# Fabric notebook: read the landed file from Files and append it to a Delta table
df = (
spark.read
.option("header", True)
.option("quote", '"')
.option("escape", '"') # match how your sample file escapes quotes inside a JSON cell
.csv("Files/fokals/company_hiring_daily/2026-10-03/")
)
df.write.mode("append").format("delta").saveAsTable("company_hiring_daily")The two loads differ in what they remember. The streaming table knows which files it has read. The notebook appends whatever folder you point it at, so give each period a folder of its own and load each folder once. On either platform the rule of the data does the rest: daily Fokals rows are written once, after the day closes, and never revised, so a silver table keyed on the company ID and the day can take the same period twice without harm if you insert only the keys you do not hold. The same property makes the silver table point-in-time by construction: a query as of any past day returns what was known that day.
The longer versions are in loading company data into Databricks with Auto Loader and Microsoft Fabric: bringing external company data into OneLake.
When you run both
If your organisation runs Azure Databricks and Fabric, Microsoft documents a way to load once. A mirrored catalog from Azure Databricks brings the structure of the catalog into Fabric and reads the underlying data through shortcuts, with no data moved or replicated. Fabric adds a SQL analytics endpoint over it, and Power BI can report on it in Direct Lake mode.
Two limits on that page matter for the load above. Streaming tables and materialized views are not displayed in the mirror, and neither are external tables that are not in Delta format. Changes can take from a few seconds to several minutes to appear. The bronze streaming table will therefore not show in Fabric. Publish a plain Delta table downstream of it, which is the silver table you would build anyway, and mirror that.
Sharing with another organisation
Sharing matters when other vendors deliver that way, and when a redistribution licence has you passing derived tables to your own customers.
Databricks documents OpenSharing, which it announced in June 2026 as the next evolution of Delta Sharing. A share is a read-only collection of tables. A recipient with a Unity Catalog workspace is reached by Databricks-to-Databricks sharing. Any other recipient is reached by open sharing, with a bearer token or OpenID Connect federation, and reads the data with tools such as Apache Spark, pandas or Power BI. The data is not replicated, a provider can revoke access at any time, and a cloud vendor may charge egress when a share crosses clouds or regions. Delta Sharing, explained for data licensing teams covers the protocol.
Microsoft documents external data sharing as sharing from one Fabric tenant to another. Administrators enable it on both sides. The recipient accepts a link and chooses a lakehouse, where a shortcut to the shared data appears: read-only, in place, with nothing copied. A share can be revoked. The same page sets out the limits. Governance controls of the provider's tenant do not cross the boundary, the sharer cannot control who has access inside the consumer's tenant, and the consumer can grant access to anyone. OneLake shortcuts and external data sharing follows that through for a licence.
The conclusion is the same for both. The mechanism ends access to the share and nothing more, so the terms on onward use belong in the contract.
Governing licensed tables
A licence names who may use the data and for what. Each platform gives you a tool for each limit.
| Licence limit | On Databricks | On Fabric |
|---|---|---|
| Only named teams may read | Privileges on a catalog or schema kept for the vendor | A workspace kept for the vendor, with workspace roles |
| Some users may see only part | Row filters and column masks | Row-level and column-level security in OneLake security roles |
| Know what was built from it | Lineage, captured automatically down to the column | The lineage view of a workspace, with impact analysis |
| Show who read it | The audit log system table | OneLake diagnostics, and Purview Audit |
| Mark the licence scope | Tags on catalogs, schemas, tables and volumes | Tags, and sensitivity labels from Microsoft Purview |
Unity Catalog places every table and volume in a three-level namespace of catalog, schema and object, so one catalog for a vendor fences both its files and its tables. Governing licensed data with Unity Catalog works the pattern through.
Fabric splits access into two planes, as its security overview explains. Workspace roles and item permissions decide what a person can do to an item. OneLake security roles decide which folders, tables, rows and columns a person can read. One detail decides the design: workspace Admins, Members and Contributors already read and write all data in the workspace, and OneLake security roles bind only Viewers and people with Read permission on an item. Licensed tables therefore belong in a workspace whose contributors are all licensed users. Microsoft's governance overview adds that sensitivity labels and data loss prevention come from Microsoft Purview and need additional licensing.
Which fits, by use
- Reports are built in Power BI and the organisation is on Microsoft's cloud. Fabric keeps the tables in the lake that Power BI reads directly.
- Engineers build pipelines and want files tracked for them. Databricks's Auto Loader and streaming tables do that, in SQL or Python.
- You will share derived tables with customers on other platforms. Databricks documents open recipients. Microsoft documents sharing between Fabric tenants.
- You already run Azure Databricks and Fabric. Load once in Databricks and mirror the catalog into Fabric.
- Your storage is spread across clouds. Databricks workspaces can be hosted on AWS, Azure and Google Cloud. Fabric reaches Amazon S3 and Google Cloud Storage through shortcuts.
Microsoft Fabric vs Snowflake sets Fabric beside a warehouse on the same questions.
Frequently asked questions
Can Microsoft Fabric read tables from Databricks?
Yes, for Azure Databricks. Microsoft documents a mirrored catalog from Azure Databricks: Fabric mirrors the structure of the catalog and reads the underlying data through shortcuts, with no data moved or replicated. Streaming tables, materialized views and external tables that are not in Delta format are not displayed, and changes can take from seconds to several minutes to appear.
Do Databricks and Microsoft Fabric both use Delta Lake?
Yes. Databricks documents Delta Lake as the default format for all its tables: Parquet data files with a transaction log. Microsoft documents Delta Lake as the default table format of a Fabric lakehouse, and says OneLake stores tables in Delta Parquet or Iceberg format. Both describe the format as open, so the tables are files that other engines can read.
What is the difference between Unity Catalog and OneLake security?
Unity Catalog is the governance layer of Databricks: privileges on catalogs, schemas, tables and volumes, row filters and column masks, lineage and an audit log. In Fabric, workspace roles and item permissions control actions on items, and OneLake security roles control which tables, rows and columns a person reads. Workspace Admins, Members and Contributors read all data in their workspace whatever those roles say.
Can I share data from Microsoft Fabric with an organisation that does not use Fabric?
Microsoft documents external data sharing as sharing between Fabric tenants: the recipient accepts the share into a lakehouse in its own tenant, where a read-only shortcut appears. For recipients on other platforms, Databricks documents open sharing through OpenSharing, in which a recipient authenticates with a bearer token or OpenID Connect federation and reads with tools such as Apache Spark, pandas or Power BI.
How do I load Fokals data into Databricks or Microsoft Fabric?
Fokals is delivered direct, by REST API and as bulk files in CSV, JSON or JSON Lines. On Databricks you land the files in a Unity Catalog volume and Auto Loader loads them into a Delta table. On Fabric you land them in the Files folder of a lakehouse and a notebook appends them to a Delta table, as this comparison shows. Daily rows are written once and never revised, so each period loads once, and the tables are your copy under your licence.
The queries and code on this page are examples to adapt. Test them in your own environment before you rely on them.
What this page says about the products it names was checked against their public documentation on 4 October 2026. Product and company names are trademarks of their owners. Fokals is not affiliated with them or endorsed by them.