A team that licenses company data receives the same thing every day: files, or pages from an API, that have to become tables it can join to its own. This comparison takes one such file through Snowflake and through Databricks and sets the two side by side on the five things that shape the work: loading, semi-structured data, sharing, governance and AI functions. Every statement about either platform comes from its own documentation, read on 4 October 2026 and linked where it is used.
Fokals is delivered direct, by REST API and as bulk files in CSV, JSON or JSON Lines, which you load with the platform's own loader: COPY INTO on Snowflake, Auto Loader on Databricks. The file used here is the daily export of Hiring Activity from the hiring dataset: the open, new and closed postings of each company. The data dictionary describes it: CSV in UTF-8 with one header row, one row per company and closed UTC day, and the open postings by job function held as a JSON object in a single cell.
The two side by side
| Snowflake | Databricks | |
|---|---|---|
| Where files wait | A stage: internal, inside Snowflake, or external, over your own Amazon S3, Google Cloud Storage or Microsoft Azure storage | A Unity Catalog volume, managed by Databricks or registered over your own cloud storage path |
| Batch load | COPY INTO an existing table, on a virtual warehouse you size | A streaming table that reads files with Auto Loader, or COPY INTO a Delta table |
| Files already loaded | Skipped by COPY INTO unless you set FORCE | Recorded in a checkpoint by Auto Loader, skipped by COPY INTO |
| JSON in a column | VARIANT, OBJECT and ARRAY types | VARIANT type, JSON strings, structs, maps and arrays |
| Reading into it | column:key, cast with :: | column:key, cast with :: |
| Sharing | Secure Data Sharing between Snowflake accounts, reader accounts for consumers with none, and in preview a route for Iceberg REST Catalog clients | OpenSharing, an open protocol with recipients on Databricks or elsewhere |
| Governance | Horizon Catalog: tags, classification, masking and row access policies, lineage, access history | Unity Catalog: privileges, row and column filters, lineage, audit log, classification, tags |
| AI from SQL | Cortex AI Functions such as AI_COMPLETE, AI_CLASSIFY and AI_AGG | AI Functions: task-specific ones, such as those that classify and summarise, and a general-purpose query function |
Loading the file
Snowflake's loading overview calls the place where a file waits a stage. An internal stage holds files inside Snowflake, uploaded with PUT. An external stage points at your own storage in Amazon S3, Google Cloud Storage or Microsoft Azure, regardless of the cloud that hosts the account. COPY INTO then loads staged files into an existing table and reads CSV, JSON, Avro, ORC, Parquet and XML. A bulk load runs on a virtual warehouse that you name and size. Snowpipe, the continuous route, is described as loading micro-batches within minutes on compute that Snowflake manages.
-- Snowflake: load the day's file by header name, then read one key of the JSON cell
-- (stage and table names are illustrative)
copy into hiring_daily_landing
from @landing/hiring_daily/2026-10-03/
file_format = (
type = csv
parse_header = true
field_optionally_enclosed_by = '"'
error_on_column_count_mismatch = false -- the table may name fewer columns than the file
)
match_by_column_name = case_insensitive;
select company_id, day, f:sales::number as open_sales_postings
from (
select company_id, day, parse_json(by_function) as f
from hiring_daily_landing
) as parsed;On Databricks a place for arriving files is a Unity Catalog volume, an object that governs files of any format, with landing areas for raw data among its listed uses. Auto Loader processes new files as they arrive in storage. It reads JSON, CSV, XML, Parquet, Avro, ORC, text and binary files, and keeps a record of each file in a checkpoint so that data is processed exactly once. Databricks also documents its own COPY INTO as retriable and idempotent, and recommends streaming tables to SQL users for file ingestion.
-- Databricks: a streaming table that picks up each new file in the volume
-- (catalog, schema and volume names are illustrative)
create or refresh streaming table hiring_daily_bronze
schedule every 1 day
as select *
from stream read_files(
'/Volumes/company_data/fokals/landing/hiring_daily/',
format => 'csv'
);
select company_id, day, by_function:sales::bigint as open_sales_postings
from hiring_daily_bronze;When no schema is given, Databricks's read_files function assumes that CSV files have a header row, so the columns take the names in the file. A JSON cell contains commas and quotes. On either platform, open a sample file, see how quotes inside the cell are escaped, and set the quote and escape options to match before you rely on a load.
The difference is in who keeps the pipeline. On Snowflake the load is a statement that you schedule, on a warehouse whose size and running time you choose. On Databricks the route recommended to SQL users is a table that refreshes itself on the schedule you give it: a streaming table is refreshed by a serverless pipeline that the platform creates for it. Both skip files they have already loaded. Your fetch job writes each bulk file to the stage or the volume, and the loader takes it from there. The rule of the data then keeps the history table simple. Daily Fokals rows are written once and never revised, so key the history table on the company ID and the day and insert only the keys you do not yet hold. The full pipelines are in loading company data into Snowflake and loading company data into Databricks with Auto Loader.
Semi-structured data
Fokals files hold lists and objects as JSON in a single cell: the open postings by job function, by seniority and by country in Hiring Activity, the evidence behind each score in Intent Scores, the labels of each role in Job Postings. Both platforms keep such a value in one column and read into it with a path.
Snowflake's semi-structured types are VARIANT, which holds a value of any other type, OBJECT and ARRAY. Loaded JSON is converted to an internal format built on those types. A query walks it with a colon and dots and casts with ::, as the query above does, and the FLATTEN function turns an array or an object into rows.
Databricks's guide to semi-structured data sets out four choices: the VARIANT type, JSON kept as a string, structs, and maps with arrays. It says a JSON string is parsed in full for every query, and the page on querying variant data recommends VARIANT, created by parsing the JSON text, over JSON strings. The path notation is the same colon and ::, and it reads a JSON string too, as the second query above does. The same page sets a limit that shapes the model: a VARIANT column cannot be a partition or a clustering key and cannot be compared, grouped or ordered.
The working rule is the same on both. Keep the cell as delivered in the landing table, parse it once into the typed table, and pull the two or three keys that people filter on every day into plain columns. The keys inside a labelled cell, such as the postings by job function, follow the label version, which the manifest of each export names, so store the manifest beside the load.
Sharing
Sharing matters to a receiving team twice: some vendors deliver through a share, so nothing is loaded, and a licence that permits redistribution may have you passing derived tables on.
Snowflake's Secure Data Sharing works between Snowflake accounts. Its documentation says that no data is copied or transferred, that shared objects are read-only and that the consumer pays only for the compute used to query them. A direct share reaches accounts in the same region, and a listing is the documented route across regions and clouds. A provider can create a reader account for a consumer that has no Snowflake account. A separate page on sharing with non-Snowflake consumers describes a preview feature that shares Iceberg tables with clients of the Iceberg REST Catalog API.
Databricks calls its sharing OpenSharing, which it announced in June 2026 as the next evolution of the Delta Sharing protocol. Its documentation describes an open protocol with two routes: Databricks-to-Databricks sharing between Unity Catalog workspaces, and open sharing with a recipient on any platform, who authenticates with a bearer token or through OpenID Connect federation. Its page on reading shared data names Snowflake among the Iceberg clients that can read such a share.
Fokals data takes the file route described above. Delivered direct, by REST API and as bulk files, it becomes tables of your own on either platform. Where your licence covers redistribution, either mechanism above is how derived tables are passed on. The two mechanisms are set against each other in Delta Sharing vs Snowflake Secure Data Sharing.
Governance
A licence is a set of limits: who may use the data, for what, and whether anything built from it may leave the company. Both catalogues give you the same kinds of tool for holding those limits.
- Snowflake groups them under Horizon Catalog: object tagging, sensitive data classification, masking and row access policies, column-level lineage and access history.
- Databricks groups them under Unity Catalog: privileges on a three-level namespace of catalog, schema and object, row and column filters, lineage, an audit log system table, classification and tags. Volumes sit in the same namespace, so the raw files are governed beside the tables.
For licensed data the pattern is the same on both. Give the vendor's tables a database or catalog of their own. Tag them with the licence scope that the manifest names. Grant read access only to the groups the licence covers. Before anything is shared outside the company, use lineage to list the tables derived from them, because that list is what a redistribution clause is about. Governing licensed data with Unity Catalog works the Databricks side through.
AI functions
Both platforms let SQL call a language model on a column, and on both some of the functions are still marked as preview or beta, so check the status of each before you build on it. Snowflake's Cortex AI Functions include AI_COMPLETE, AI_CLASSIFY, AI_FILTER, AI_EXTRACT, AI_AGG and AI_SUMMARIZE_AGG. A role needs the USE AI FUNCTIONS privilege and one of two database roles, and availability depends on the region. Databricks's AI Functions divide into task-specific functions, such as those that classify, extract, summarise and translate, and one general-purpose query function. They do not run on Classic SQL warehouses, and model inference may be billed in addition to the compute that runs the query.
Company data gives these functions text in one place: the excerpt of each announcement in Company News, up to 1,200 characters in the company's own words.
-- my_companies is your own list of companies
-- Snowflake: one summary per company across its announcements
select n.company_id,
ai_agg(n.excerpt, 'Summarise these company announcements in three sentences') as summary
from company_news n
join my_companies u on u.company_id = n.company_id
group by n.company_id;
-- Databricks: one summary per announcement, of about 40 words
select n.company_id, n.url, ai_summarize(n.excerpt, 40) as summary
from company_news n
join my_companies u on u.company_id = n.company_id;The reference pages for AI_AGG and ai_summarize give the signatures. Keep generated text for reading and count with the data. Each announcement already carries its event types, classified once into 13 types under a named, frozen version, so a count of funding rounds or leadership changes should come from that field. Showing a generated summary to your own customers is embedding in a product, which is a licence scope of its own. Using company signals with Snowflake Cortex AI functions goes further on the Snowflake side.
Which to choose, by use
- Your other data is already on one of them. Load there. The join to your CRM or security master is the reason to load at all.
- Your team writes SQL and schedules statements. Snowflake's documented path is a stage,
COPY INTOand a warehouse you control. - Your team builds pipelines and wants files tracked for it. Databricks recommends Auto Loader, in a pipeline or behind a streaming table.
- Vendors or customers share with you on one platform. That platform saves a load for those datasets. Databricks documents recipients on any platform, and Snowflake documents sharing between its accounts, reader accounts and, in preview, a route for Iceberg REST Catalog clients.
- Raw files must be governed beside the tables. Databricks puts both in Unity Catalog through volumes. Snowflake controls named stages with access privileges.
Frequently asked questions
Is Snowflake or Databricks better for loading vendor CSV files?
Both load CSV with a header row, and both skip files they have already loaded. Snowflake documents a stage and the COPY INTO command, run on a virtual warehouse you size. Databricks documents Auto Loader, which tracks new files in a checkpoint, and recommends streaming tables to SQL users. Choose by where your other data sits and by whether your team prefers scheduled SQL statements or managed pipelines.
Does Databricks have a VARIANT type like Snowflake?
Yes. Databricks documents a VARIANT type for semi-structured data, created with parse_json and read with a colon path and a double-colon cast, the notation Snowflake uses for its own VARIANT type. Databricks recommends it over JSON kept as strings. A VARIANT column on Databricks cannot be a partition or a clustering key and cannot be grouped or ordered, so extract the fields you filter on.
Is Delta Sharing the same as OpenSharing?
Databricks announced OpenSharing in June 2026 as the next evolution of the Delta Sharing protocol, and its documentation now uses the new name. It is an open protocol: a provider shares tables with recipients on Databricks or on other platforms, who authenticate with a bearer token or through OpenID Connect federation. Snowflake Secure Data Sharing is documented as working between Snowflake accounts, with a separate preview route for Iceberg REST Catalog clients.
Can Snowflake read data shared from Databricks?
Databricks's page on reading shared data lists Snowflake among the Iceberg clients that can read an OpenSharing share through the Apache Iceberg REST Catalog API, with a credential that the provider issues. Snowflake's own Secure Data Sharing is documented as working between Snowflake accounts, with reader accounts for consumers who have none.
How do I load Fokals data into Snowflake or Databricks?
Fokals is delivered direct, by REST API and as bulk files in CSV, JSON or JSON Lines, which you load with the platform's own loader: a stage and COPY INTO on Snowflake, a Unity Catalog volume and Auto Loader on Databricks, as this comparison shows. Daily and weekly datasets are written once, after the period closes, and never revised, so the history table only gains rows. The tables are your copy under your licence, to join, govern and query beside your own.
The queries and code on this page are examples to adapt. Test them in your own environment before you rely on them.
What this page says about the products it names was checked against their public documentation on 4 October 2026. Product and company names are trademarks of their owners. Fokals is not affiliated with them or endorsed by them.