Comparison

CSV vs JSON Lines vs Parquet for bulk data delivery

Three file formats, three sets of trade-offs. What each keeps and loses when a vendor hands you a dataset as files, and a tested way to turn a CSV export into typed Parquet.

Updated 5 October 20268 min read

When a vendor delivers a dataset as files, the format decides how much work stands between the download and the first query. This comparison sets CSV, JSON Lines and Apache Parquet side by side for that job. By the end you will be able to say which format suits a first look, a table with nested values or repeated scans over many daily files, what each leaves you to declare, and how to convert the first two into the third.

The examples use Fokals datasets. Fokals delivers a bulk export of any dataset for any period as JSON, JSON Lines or CSV, direct, page by page through the export endpoints of its API or as a package of files with a manifest. The worked section below takes one of those exports and turns it into typed Parquet for a query engine in a single statement.

The three formats side by side

Each format has a written definition, and the definitions differ in how much they settle. CSV is described by RFC 4180, an informational memo of 2005 that records common practice. JSON Lines is three rules on one page. Apache Parquet is a specification for a binary, column-oriented file.

CSVJSON LinesApache Parquet
LayoutText: one record per line, fields separated by commasText: one JSON value per lineBinary: values stored column by column, in row groups
Column namesAn optional header lineRepeated as keys on every lineHeld once, in the file's schema
Value typesNone defined: a field is charactersThose of JSON: string, number, boolean, null, array, objectDeclared for each column, with logical types such as date, timestamp and decimal
Nested valuesNo provision: a list or object is written as text in one fieldKept as arrays and objectsKept as lists, maps and nested groups
CompressionOf the whole file, from outsideOf the whole file, from outsideInside the file, page by page, with a choice of codecs
Reading two columns of manyEvery line is parsed in fullEvery line is parsed in fullOnly the chunks of those columns are read
Opening itA text editor or a spreadsheetA text editorA program that reads the format

CSV: read everywhere, defined by convention

RFC 4180 is short. Each record is a line. An optional header line names the fields. Fields are separated by commas, and a field that contains a comma, a double quote or a line break is enclosed in double quotes, with any quote inside it written twice. The memo also says that the format had never been formally documented and that implementations differ considerably.

What it leaves open is what a loader has to decide. It defines no data types and no way to tell empty text from a missing value. It names US-ASCII as common usage and leaves other character sets to a parameter. A CSV file is therefore read correctly only when the conventions are written beside it.

For Fokals files those conventions are in the data dictionary. Files are UTF-8 with one header row. Times are UTC in ISO 8601, a day is a closed UTC day and a week runs Monday to Sunday and is identified by its Monday. Lists and objects are JSON in a single cell. A row of Hiring Activity for Acme Robotics, an illustrative company, trimmed to five columns and quoted the way RFC 4180 describes, reads:

company,day,open_postings,new_postings,by_function
Acme Robotics,2026-10-01,42,3,"{""software_engineering"": 18, ""sales"": 9, ""operations"": 15}"

Open a sample file before you set a loader's quote and escape options, because the JSON cell is where they are tested. Three habits prevent most faults.

  • Declare every type, and declare identifiers as text. The ISIN, LEI, FIGI, ticker and MIC that Fokals carries on every listed company are labels to join on. A reader left to infer types can read a column of numeric tickers as integers and drop their leading zeros.
  • Read an empty field as missing. The dictionary specifies which columns are a value or empty, such as the work mode and the advertised pay of a posting. An empty field is not a zero and not an empty text.
  • Check the field count. RFC 4180 says each line should hold the same number of fields, so a loader that fails on a different count catches a broken quote at the row where it happens.

JSON Lines: types and nesting kept, names repeated

The JSON Lines page asks three things of a file: UTF-8 encoding, a valid JSON value on every line and \n as the line terminator. A record never spans lines, so a file can be processed one record at a time.

Against CSV, two problems go away. A list or an object is an array or an object, not text inside a quoted field, so nothing needs a second parse. And a value says what it is: "0700" in quotes is a string and stays one, while 42 is a number. RFC 8259 of 2017, which defines JSON, gives it four primitive types (strings, numbers, booleans and null) and two structured types (objects and arrays). Written as a line of JSON Lines, the Acme Robotics record keeps its object as an object:

{"company": "Acme Robotics", "day": "2026-10-01", "open_postings": 42, "new_postings": 3, "by_function": {"software_engineering": 18, "sales": 9, "operations": 15}}

Two limits remain. JSON has no date type, so a day or a timestamp is a string that you still cast. And RFC 8259 notes that implementations agree exactly on integers within the range a double-precision number holds, which ends just short of 2 to the power 53 in either direction, so a large numeric identifier belongs in a string.

The cost is size and a quiet failure. Every line repeats every key, so compress the file: .jsonl.gz is the convention the format's page names. A transfer cut between two records leaves a shorter file that is still valid, so check completeness yourself. Count the rows for each day or week in the period and compare the last period with the one before it.

Parquet: typed columns, read by column

Parquet's documentation describes a column-oriented file format designed for efficient storage and retrieval. Its concepts page divides a file into row groups. Each row group holds one column chunk for each column, and each chunk is divided into pages, the unit at which values are encoded and compressed. The file layout puts the file metadata after the data, and a reader starts there to find the column chunks it wants.

Three things follow for a delivered dataset.

  • Types travel with the file. Columns have primitive types and logical types such as DATE, TIMESTAMP, DECIMAL, LIST and MAP. A day column is a date and a count of postings by job function can be a map, with no option to set on the reader's side.
  • A query reads less. An engine that wants the day and the count of open postings reads the chunks of those two columns. DuckDB's Parquet page, for one, says only the columns a query needs are read and that filters are pushed into the scan.
  • Compression is part of the format. Pages can be compressed with one of several codecs, among them Snappy, GZIP, Brotli and Zstandard.

The costs are as plain. A Parquet file cannot be read by eye: it needs a program that understands the format. It is not extended a line at a time: the metadata is written once, after the data, so a dataset grows by adding files, and an engine such as DuckDB reads a series of them as one table. And the project's overview warns that not all implementations support the same features of the format, so test the writer you choose against the readers you use.

Which format for which job

JobFormat that fitsWhy
A first look at a sampleCSVIt opens in a text editor or a spreadsheet, and the header names the columns
A dataset with list or object columns, such as the evidence attached to every row of Intent ScoresJSON LinesArrays and objects arrive as arrays and objects
A daily file loaded once into a warehouseCSV or JSON LinesThe loader reads either, and the format matters only until the load ends
Repeated scans over months of files in object storageParquetTyped columns, reads limited to the columns needed, built-in compression
The copy you keep as the record of a deliveryThe original files and the manifestThey are what arrived, in the form it arrived

No one of the three suits every job. The choice follows the job, and one delivery often passes through two of them: the vendor's text file for the record, and Parquet for the queries.

Converting an export to Parquet

The conversion is one statement in any engine that reads CSV and writes Parquet. The example uses DuckDB on an export of Hiring Activity from the hiring dataset. It reads every column as text, casts each to the type the data dictionary gives it and turns the JSON cell into a map.

COPY (
  SELECT
    company_id,
    company,
    isin,
    CAST(day AS DATE)                 AS day,
    CAST(open_postings AS INTEGER)    AS open_postings,
    CAST(new_postings AS INTEGER)     AS new_postings,
    CAST(closed_postings AS INTEGER)  AS closed_postings,
    CAST(CAST(by_function AS JSON) AS MAP(VARCHAR, INTEGER)) AS by_function,
    CAST(median_salary_usd AS DOUBLE) AS median_salary_usd
  FROM read_csv('company_hiring_daily.csv', header = true, all_varchar = true)
) TO 'company_hiring_daily.parquet' (FORMAT parquet, COMPRESSION zstd);

all_varchar = true switches off type detection, so nothing is guessed and a value that does not fit its cast stops the statement. DuckDB's CSV reader reads an empty field as NULL, because the default of its nullstr option is the empty string. The options are documented on DuckDB's pages for CSV and Parquet, and its JSON type page says that JSON can be cast to any of DuckDB's types and shows the cast to a nested type. The result answers a question by reading three columns:

SELECT day,
       sum(open_postings)        AS open_postings,
       sum(by_function['sales']) AS open_sales_postings
FROM 'company_hiring_daily.parquet'
WHERE day >= DATE '2026-10-01'
GROUP BY day
ORDER BY day;

Keep three rules when you convert.

  1. Write one Parquet file for each export file, in a folder for each dataset. Daily and weekly Fokals datasets are written once, after the period closes, and are never revised, so a new period adds a file and replaces none.
  2. Keep the original files and the manifest that comes with an export. The manifest names the period, the label versions and the licence. The Parquet copy is derived from the delivery and sits beside it.
  3. Carry every column across, including the label version where a dataset has one, so that a row still says under which version its labels were produced.

What a format changes, and what it keeps

A format changes how data is stored, and the record stays what it was. A Fokals export converted to Parquet is the same point-in-time record: every observation dated, every period written once and never revised. A cast records your reading of a column, and the data dictionary is what makes that reading right.

Parquet is also a file format and not a table: nothing in a file knows about the other files in the folder. For querying a folder of converted files in place, see querying company data on Amazon S3 with Athena, and for profiling a sample before any conversion, exploring bulk company data with DuckDB. Whether to take files at all, or to read the API, is the subject of API vs bulk files for company data delivery.

Frequently asked questions

Is Parquet better than CSV?

For repeated analytical queries over large files, Parquet fits better: columns are typed, a query reads only the columns it needs and compression is built in. For a first look, a small file or a load into a warehouse that converts the data anyway, CSV is simpler, because any tool opens it. Neither is better in general, so choose by what you will do with the file.

What is the difference between CSV and JSON Lines?

CSV holds one record per line as comma-separated fields, with no types and no place for nested values. JSON Lines holds one JSON value per line, so strings, numbers, booleans, nulls, arrays and objects arrive as what they are. JSON Lines repeats the field names on every line and is larger before compression. CSV opens directly in a spreadsheet, while JSON Lines is usually read by a script or a query engine.

Which file formats does Fokals deliver?

Fokals delivers bulk exports of any dataset for any period as JSON, JSON Lines or CSV, page by page through the export endpoints of its API or as a package of files with a manifest naming the period, the label versions and the licence. The files are delivered direct, and you load them into the warehouse, the lake or the query engine you use. The DuckDB statement in this comparison turns an export into typed Parquet in one step.

How do I convert a CSV file to Parquet?

Use an engine that reads CSV and writes Parquet. In DuckDB, read the file with its CSV reader, cast each column to its type and wrap the query in COPY ... TO 'file.parquet' (FORMAT parquet). Read every column as text first, with all_varchar = true, so that identifiers keep their leading zeros and no type is guessed. Keep the original file beside the Parquet copy.

Can Parquet store nested data such as JSON objects?

Yes. Parquet has logical types for lists and maps and encodes nested columns with definition and repetition levels, so a JSON object in a CSV cell can become a typed map column. It also has a JSON logical type for a document kept whole as text. Which to use depends on whether you query inside the value or only carry it.

Why do leading zeros disappear when I load a CSV file?

CSV carries no types, so a loader that infers them can read a column of digits as numbers, and a number has no leading zeros. The value was a label, such as a ticker or a company identifier, and the fix is to declare the column as text before loading. JSON Lines avoids the problem when the producer writes the value in quotes, and Parquet stores the declared type in the file.

The queries and code on this page are examples to adapt. Test them in your own environment before you rely on them.

What this page says about the products it names was checked against their public documentation on 4 October 2026. Product and company names are trademarks of their owners. Fokals is not affiliated with them or endorsed by them.