Comparison

API vs bulk files for company data delivery

Most teams need both: files to build the copy, an API to keep it current and to answer look-ups. How to assign each job to a route, with the HTTP and file-format rules that decide it.

Updated 5 October 20267 min read

A vendor that offers company data through a REST API and as bulk files is offering two tools, not two versions of one. An API answers a question and keeps a copy current. A file hands over a period of a table whole. This comparison assigns the jobs of a data team to one route or the other, sets out what the standards behind each route promise and what they leave open, and works one split through from first load to daily run.

Fokals delivers both, direct: a REST API of 25 endpoints, and bulk exports of any dataset for any period as JSON, JSON Lines or CSV, as the delivery page describes. The data is refreshed daily, and you load it into the warehouse, the lake or the product where you need it. That makes it a complete worked example: every job in the tables below can be assigned to one of its two routes.

Two shapes of the same data

REST names a style of interface, and the rules a client can rely on are those of HTTP. You send a request, usually a GET, and receive a representation of a resource, usually JSON. A long list arrives in pages, and with cursor pagination the client holds the position between requests.

A bulk file is rows written out in a text format. This comparison covers two. CSV is one record to a line with fields separated by commas. JSON Lines is one JSON value to a line.

REST APIBulk files
UnitA request and a page of recordsA file that covers a period of one table
Where the state isOn the client, as a cursorIn the file and its manifest
A failure shows asA status code, at onceA short or malformed file, when you load it
TypesJSON numbers, strings, booleans and nullIn CSV, text only; in JSON Lines, as JSON
FitsLook-ups, recent changes, short listsFirst loads, backfills, whole-list studies, evidence
Asks of youA client that paces, retries and keeps a bookmarkA landing area, a parser and a key for every table

Which job goes to which route

JobRouteWhy
First load of historyBulk filesA period arrives as one artefact that you check against its manifest
Daily top-up of tables you holdAPI feedsA feed read oldest first resumes from the last cursor
Look-up of one company inside a productAPIOne request returns the latest record
A study across the whole listBulk filesThe input is complete and fixed, so the result can be reproduced
Reloading a period after a faultBulk filesA closed period that is never revised reloads the same
Evidence for an audit or a licenceBulk filesThe manifest states what was delivered and under which terms
Alerts shown to usersAPI feedsEach read of a feed returns the new events in order
Latest announcements on a pageAPIA list that runs newest first puts them at the top

The rows follow how Fokals is built. Its feeds of signals, website changes and postings run oldest first from a time you set, so the last cursor is the bookmark for incremental sync, and its list of announcements runs newest first. Its daily and weekly datasets are written once, after the period closes, and are never revised. A package of files carries a manifest that names the period, the label versions and the licence.

What HTTP gives an API client

Three passages of the HTTP standards decide how a client should behave.

A read can be repeated. RFC 9110, published in 2022, defines GET as a safe method and safe methods as idempotent, and says an idempotent request can be repeated automatically when a connection fails before the response is read. A page that times out is sent again with the same cursor, and nothing is harmed.

A bearer key works for whoever holds it. RFC 6750, published in 2012, defines a bearer token as one that any party in possession of it can use, and says it needs to be protected from disclosure in storage and in transport. Fokals issues one bearer key to each client, with scopes. Keep it in a secret store, call the API from a server and never place the key in a page a browser loads.

A refusal carries its own instruction. RFC 6585, published in 2012, defines status 429 for a client that has sent too many requests in a given time, and allows the response to include a Retry-After header. RFC 9110 gives that header two forms, a number of seconds or a date. Fokals limits each key by the minute and by the day, so a client needs both a pace and a daily budget.

A request is a few lines. This one asks for the catalogue of Market Series metrics. The host is a placeholder.

curl --silent --show-error --fail \
  --header "Authorization: Bearer $FOKALS_API_KEY" \
  "https://$FOKALS_HOST/api/v1/series/metrics"

The retry rule is not much longer. This sketch waits as long as the server asks when the header gives a number of seconds, and otherwise backs off with a growing pause.

import time
import requests

def get(url, key, params=None, attempts=5):
    for attempt in range(attempts):
        r = requests.get(
            url,
            headers={"Authorization": f"Bearer {key}"},
            params=params,
            timeout=60,
        )
        if r.status_code == 429:
            wait = r.headers.get("Retry-After", "")
            time.sleep(int(wait) if wait.isdigit() else 2 ** attempt)
            continue
        r.raise_for_status()
        return r
    raise RuntimeError("still refused: stop, keep the cursor, resume later")

What a file format settles, and what it leaves open

RFC 4180, published in 2005, is an informational document that says there is no formal specification of CSV and records the format most implementations follow. It makes the header line optional, defines no data types, and says that a field containing a comma, a double quote or a line break should be enclosed in double quotes, with each double quote inside it written twice.

That last rule matters for company data. Fokals files are CSV in UTF-8 with one header row, and a list or an object is held as JSON in a single cell. A JSON cell is full of commas and quotes, so written as RFC 4180 describes it looks like this. The row is illustrative and trimmed to five columns.

company,day,open_postings,new_postings,by_function
Acme Robotics,2026-10-02,44,2,"{""software_engineering"": 21, ""sales"": 9}"

A reader that splits lines on commas breaks on that row, and a CSV parser does not. Parse first, then decode the cell and cast the numbers.

import csv
import json

with open("company_hiring_daily.csv", encoding="utf-8", newline="") as f:
    for row in csv.DictReader(f):
        open_postings = int(row["open_postings"])      # CSV carries no types
        by_function = json.loads(row["by_function"] or "{}")

JSON Lines settles more. The specification has three requirements: UTF-8 encoding, a valid JSON value on every line, and a newline as the line terminator. Numbers stay numbers, a missing value is null, and a nested value can stay nested. A reader takes one line at a time, so a large file needs little memory and a failed load can resume at a line. The specification suggests the extension .jsonl, recommends a stream compressor such as gzip, and notes that no media type has yet been standardised.

Neither format settles what a column means, whether the file is complete or which time zone a date is in. Those answers come with the delivery: the data dictionary for meaning, the manifest for completeness, and a stated convention for time. Fokals times are UTC in ISO 8601, a day is a closed UTC day, and a week runs Monday to Sunday. The comparison of CSV, JSON Lines and Parquet takes the formats further.

A worked split

Illustrative: Acme Robotics licenses hiring and technology data for 3,000 companies. It wants the datasets in a warehouse and behind a feature of its own product.

  1. Start from a sample. Sample data for the companies you track is sent on request with the data dictionary and the methodology. Write the parsers and choose the keys against those files.
  2. Build the copy from bulk exports. Request each closed period of the datasets you license. Store each file beside its manifest, compare the manifest's period with the one you asked for, and count rows against distinct keys.
  3. Hand over to the feeds. Start each feed at the end of the last exported period and save the cursor after every page you store. The guide to incremental sync builds that loop and its upserts.
  4. Serve look-ups from your own server. The product calls Acme's back end, which calls the API with the key and caches the answer.
  5. Repair with files. When a check on your side finds a day missing from your copy, export that period again. Daily and weekly rows are written once and never revised, so the reload matches the first delivery exactly.

Before step 2, estimate the requests. Both routes are paged, so a history costs rows divided by rows to a page, and an export page holds up to 1,000 rows. On illustrative figures, 2,000,000 rows at 1,000 to a page is 2,000 requests. Set that against the limits on your key, spread a long history over several days where it does not fit, and keep the daily run on its own schedule while the first load runs. The API reference gives the parameters of each endpoint.

The failure modes, side by side

  • Interrupted transfer. On the API a page is lost and the request is repeated. A JSON Lines file that was cut fails on its last line. A CSV that was cut between two records looks complete, so compare row counts with the manifest or with a previous load.
  • Loaded twice. Both routes need the same defence: a key for every table and an upsert on it.
  • Wrong types. JSON carries numbers and null. CSV carries text, and a tool that guesses types can turn an identifier into a number.
  • Changed layout. Select columns by name and ignore keys you do not know, so that an added field does not break the load.
  • Silent lag. An API job that returns nothing still succeeds. Alert on the age of the newest row, on both routes.

Frequently asked questions

Should I use an API or bulk files for a first load of company data?

Use bulk files when the history is long or the list of companies is large. A file covers a whole period, arrives with a manifest you can check it against, and spares your request allowance. For a short list the API alone is enough. Either way, load through the same keys and upserts, then keep the copy current from the API feeds.

Is JSON Lines better than CSV for bulk data?

It depends on the data and the tool. JSON Lines keeps types and nested values and can be read a line at a time, which suits rows that hold lists or objects. CSV opens in every spreadsheet and loader but carries text only, and nested values have to be quoted inside a cell. Choose JSON Lines for pipelines and CSV for inspection by people.

What does HTTP 429 mean when I call a data API?

It means you have sent more requests than the server allows in the current period. RFC 6585 defines the status and lets the server add a Retry-After header that says how long to wait, in seconds or as a date. Stop, wait for that time, and continue at a lower rate. If the daily limit is spent, save the cursor and resume the next day.

How do I avoid a gap between a bulk load and an API feed?

Start the feed at the end of the last exported period and let the two overlap slightly. Load both through an upsert on the table's key, so a row that arrives twice changes nothing. Then check that the days you hold have no hole between the end of the files and the start of the feed.

How does Fokals deliver company data?

Direct, by REST API and as bulk files. The API has 25 endpoints and returns JSON under one bearer key per client, with scopes, rate limits per key and cursor pagination. Bulk exports cover any dataset for any period as JSON, JSON Lines or CSV, page by page through the export endpoints or as a package of files with a manifest naming the period, the label versions and the licence. You load the data into the warehouse, the lake or the product where you need it.

Why do some CSV cells contain JSON?

A table can hold a value that is itself a list or an object, such as open postings counted by job function. CSV has no nested type, so the value is written as JSON text in one cell and quoted. Parse the file with a CSV parser, then decode that cell as JSON, or take the same data as JSON Lines.

The queries and code on this page are examples to adapt. Test them in your own environment before you rely on them.

What this page says about the products it names was checked against their public documentation on 4 October 2026. Product and company names are trademarks of their owners. Fokals is not affiliated with them or endorsed by them.