Comparison

Fivetran vs Airbyte for ingesting a vendor API

A vendor's REST API reaches your warehouse through a connector you build. How the two tools differ, as each documents itself: where the code runs, what you write, where the bookmark lives and how rows are counted.

Updated 5 October 20269 min read

A data vendor's REST API reaches your warehouse through a connector you build, and the ingestion tool you run decides what building means. This comparison sets Fivetran and Airbyte side by side for that one job: a custom connector to a vendor's REST API. It covers who runs the connector, what you write, where the bookmark for incremental sync lives, how each product counts rows and what open source means in Airbyte's case. Every statement about either product comes from its own documentation, read on 4 October 2026.

The example is the Fokals REST API. Fokals is delivered direct, by REST API and as bulk files, and a connector you build in either tool loads the API into your destination. The API gives a connector what it needs: JSON over REST, one bearer key, cursor pagination, and feeds that run oldest first from a time you set, so the last cursor is the bookmark. Both tools run the sync, keep the bookmark and write the rows to your destination, which is what each is built for when you want that work operated for you. The build itself, setting by setting, is in the guide to ingesting a vendor API with Fivetran or Airbyte. This piece is about choosing between the two.

The two side by side

FivetranAirbyte
Who runs a custom connectorFivetran, in an environment it hostsYou, on Airbyte Core in your own infrastructure, or Airbyte on its cloud plans
Documented routeConnector SDKConnector Builder, low-code CDK or Python CDK
What you writePython: an update function and, usually, a schema functionA form that produces a YAML manifest, the YAML itself, or Python
Where the bookmark livesA state dictionary your code defines and saves with op.checkpointThe state of the connection: for a Builder stream, the latest value of a date or time field
How a row is writtenop.upsert, op.update or op.delete, by primary keyThe sync mode of the connection: append, or append with deduplication by primary key
A rate-limit responseHandled in your codeRetried by default: 429 and 5XX responses, five times, with exponential backoff
Usage measureMonthly Active RowsCredits on the Standard and Plus plans, data workers on Pro and Enterprise Flex

Managed and open source, as each documents it

Fivetran documents two deployment models. In SaaS Deployment, data processing happens entirely within the Fivetran cloud. In Hybrid Deployment, data is processed inside your own network while the Fivetran cloud orchestrates the movement, for the connectors and destinations that support it.

Airbyte's documentation lists five plans. Airbyte Core is described as the open-source version, deployed locally or in your own infrastructure. Standard, Plus and Pro are cloud plans, and Enterprise Flex is described as a hybrid one. Running Core is your work: the quickstart installs it with a command-line tool called abctl, says Airbyte runs on Kubernetes, and recommends a machine with four or more CPUs and at least 8 GB of memory.

What open source means is on Airbyte's licences page. It says Airbyte's connectors and the rest of its public repositories are available under the Elastic License 2.0, that the Airbyte Protocol is under the MIT licence, and that Airbyte Cloud and Airbyte Enterprise require a commercial licence. The page adds that you should be fine unless you host Airbyte yourself and sell it as an ELT or ETL tool, or sell a product that exposes Airbyte's interface or API directly.

What you build in each

The Fivetran Connector SDK builds custom connectors in Python. The code runs in an isolated environment that Fivetran hosts, and Fivetran runs the connection on your schedule and manages the compute. A connector declares a required update(configuration, state) function and usually a schema(configuration) function that names tables and primary keys, as the SDK methods page sets out. Inside update your code calls the API and hands each row to an operation such as op.upsert. Fivetran's page on function connectors, which run code in your own cloud account, says users who signed up on or after 22 July 2025 have no access to them and that all users can use the Connector SDK.

Airbyte's connector development page lists several routes and recommends the first of these three for an API source.

  • Connector Builder. A no-code tool in the Airbyte interface, laid over the low-code YAML format. Its overview says it builds source connectors only. Forms cover authentication, pagination, incremental sync, record processing and error handling.
  • Low-code CDK. The same declarative YAML, written by hand.
  • Python CDK. The page says it gives the most flexibility and needs the most code and maintenance.

A Builder connector can also hold custom components, which are Python classes. Their page marks them as unsafe and experimental, and an administrator has to switch them on.

One vendor API through both

A few properties of an API shape its connector. For Fokals they are JSON over REST, one bearer key per client with scopes, rate limits per key by minute and by day, and cursor pagination. The feeds of signals, website changes and postings run oldest first from a time you set, so the last cursor is the bookmark. Bulk exports come through export endpoints as JSON, JSON Lines or CSV. The API reference gives the endpoints and parameter names.

What the API doesFivetran Connector SDKAirbyte Connector Builder
One bearer keyRead from configuration and sent as a header by your codeBearer Token authentication, with the key as a user input
Cursor paginationA loop in update that sends the cursor back until none is returnedCursor Pagination: the next cursor is read from the response body or a header and injected into the next request
Feeds oldest first from a time you setState holds what you choose: the feed's last cursor, or the last time seenIncremental sync on a date or time field such as the observation time, whose latest value becomes the state
Rate limits per keyYour code waits, or checkpoints and ends the runDefault retries, or a backoff that takes the wait from a response header

The third row is the one to weigh. Fivetran's state management page describes state as a JSON dictionary that you define and that is saved only when you checkpoint, so the bookmark can be the feed's own cursor. Airbyte's incremental sync page describes a bookmark that is a time: the most recent date met in a cursor field becomes the state, and the next sync starts from it. With a time as the bookmark, the rows at the boundary can be read again, which a primary key and a deduplicating sync mode absorb.

The fourth row follows the two error handling pages. Airbyte's page on error handling says a connector retries 429 and 5XX responses five times by default with exponential backoff, and lists a constant backoff, an exponential one and two that read the wait from a response header. Fivetran's page on error handling advises retrying and honouring any Retry-After header, which in the Connector SDK is code you write.

Here is the Fivetran side for Hiring Activity, the daily count of open, new and closed postings per company in the hiring dataset. The data dictionary gives its grain as one row per company and closed UTC day, which is the primary key. Daily and weekly Fokals datasets are written once, after the period closes, and are never revised, so the bookmark is the last day loaded.

from fivetran_connector_sdk import Connector
from fivetran_connector_sdk import Operations as op

TABLE = "company_hiring_daily"


def schema(configuration: dict):
    return [{"table": TABLE, "primary_key": ["company_id", "day"]}]


def update(configuration: dict, state: dict):
    day = next_closed_day(state.get("last_day"), configuration["first_day"])
    while day is not None:
        for row in export_rows(configuration, TABLE, day):
            op.upsert(table=TABLE, data=row)
        op.checkpoint(state={"last_day": day})
        day = next_closed_day(day, configuration["first_day"])


connector = Connector(update=update, schema=schema)

The two helpers are yours to write against the API reference. The first returns the next UTC day that has closed, or nothing once every closed day is loaded. The second requests one dataset's export for one day with the bearer key, follows the pages and yields each row as a dictionary. On a long first sync, checkpoint on a timer instead: the state page advises roughly every ten minutes and not more than once a minute.

The same dataset in the Airbyte Builder is a set of choices, not code:

  1. Authentication: Bearer Token, the key entered when the source is created.
  2. Stream: the export request for the dataset, a record selector that points at the list of rows, and a primary key of the company ID with the day.
  3. Pagination: Cursor Pagination.
  4. Incremental sync: the day as the cursor field, a start date as a user input, the start and end of each interval injected into the request, a step of one day and a lookback window.
  5. Connection: the sync mode Incremental Append + Deduped.

One property of the data serves both. Daily and weekly rows carry a flag, described in the data dictionary, that marks a period written more than seven days after it closed, so a pipeline can pick up a late period and a model can set it apart. Read a trailing window again on each run, with a lookback window in Airbyte or date arithmetic in Fivetran, and a longer window now and then. With the key above, a second reading changes only what is new.

What each counts

Neither product's prices are given here, but the units differ and the units decide the bill.

Fivetran's pricing page defines Monthly Active Rows as the distinct rows synced to a destination in a calendar month, tracked by primary key. A row counts once in a month however often it is updated. The Connector SDK page says the SDK is measured in Monthly Active Rows like any other connector.

Airbyte's page on billing and credits says credits pay for syncs on the Standard and Plus plans and that API and custom sources are measured in rows. For an incremental sync those are typically the rows added, edited or deleted, and for a full refresh every row synced is charged. Pro and Enterprise Flex are described as capacity-based plans that use data workers.

Put two Fokals datasets through both. Hiring Activity holds one row per company and day, so an illustrative set of 2,000 companies adds at most 60,000 rows in a 30-day month, each new once. A row of Job Postings changes while the posting is open, as its last-seen date moves and its closing date is set. Under Fivetran's definition that posting counts once a month. Under Airbyte's it typically counts when it is added and again when it is edited. Ask either vendor how rows read a second time by a lookback are counted.

Which to choose, by use

  • You already run Fivetran and want nothing new to operate. The Connector SDK keeps the connector beside your other connections, hosted by Fivetran.
  • You want the connector held as configuration. The Airbyte Builder produces a YAML manifest for an API that fits its forms, as a cursor API with a bearer key does.
  • The pipeline must run inside your own network. Airbyte Core is deployed in your own infrastructure. Fivetran's Connector SDK page documents Hybrid Deployment for the SDK too, and Fivetran's own pages say which plans include it.
  • The API needs logic of its own. The list of Fokals announcements runs newest first, so a run is complete only when it reaches the bookmark. The Fivetran SDK is code from the first line. In Airbyte, test that stream in the Builder before you trust it, and expect the Python CDK where the forms end.
  • You need a file, not a pipeline. An export loaded by your warehouse's own loader may be enough, as API vs bulk files sets out.

After the rows land

Both tools land rows, not a model, and neither adds data: the joins on the company ID, ISIN or FIGI come after. Fokals data is shaped for those joins. Each company has one stable company ID across the datasets, and a listed company carries its ticker, MIC, ISIN, LEI and share-class FIGI, so the hiring rows a connector lands join to technographic, intent and announcement data on the company ID, and to your own security master on ISIN or FIGI. The use case on keeping a warehouse in step with incremental API sync takes bookmarks, upserts and backfills further.

Frequently asked questions

Is Fivetran or Airbyte better for a custom API connector?

It depends on who should run it and how you want to write it. With Fivetran you write Python against the Connector SDK and Fivetran hosts and schedules the code. With Airbyte you fill in the Connector Builder's forms, which produce a YAML manifest, and run the result on Airbyte's cloud plans or on Airbyte Core in your own infrastructure. Choose Fivetran to operate nothing new, and Airbyte for configuration over code or for running it yourself.

Is Airbyte open source?

Airbyte's documentation calls Airbyte Core its open-source version, deployed locally or in your own infrastructure. Its licences page says the connectors and the rest of its public repositories are under the Elastic License 2.0 and the Airbyte Protocol is under the MIT licence, while Airbyte Cloud and Airbyte Enterprise require a commercial licence. Read the Elastic License terms if you plan to offer Airbyte to others as a service.

What is the Fivetran Connector SDK?

It is Fivetran's documented way to build a custom connector. You write a Python module with an update function that fetches data and passes rows to operations such as op.upsert, and usually a schema function that names tables and primary keys. You test it locally with fivetran debug and release it with fivetran deploy, as the command reference describes. Fivetran then hosts the code, runs it on your schedule and hands it the state it last saved.

How do Fivetran and Airbyte count rows from a custom connector?

Fivetran measures Monthly Active Rows: distinct rows, by primary key, synced in a calendar month, each counted once however often it changes. Airbyte's Standard and Plus plans use credits, measured in rows for API and custom sources: typically the rows added, edited or deleted on an incremental sync, and every row on a full refresh. Its Pro and Enterprise Flex plans are capacity-based, and Airbyte Core is run on your own machines.

How do I ingest the Fokals API with Fivetran or Airbyte?

Build a custom connector against the API reference, with the Fivetran Connector SDK or the Airbyte Connector Builder. The Fokals REST API returns JSON with one bearer key per client and cursor pagination, and its feeds of signals, website changes and postings run oldest first from a time you set, so the last cursor is the bookmark for incremental sync. Daily and weekly datasets are written once and never revised, so a sync keyed on company ID and day loads each period once. Bulk exports in JSON, JSON Lines or CSV serve a first load.

The queries and code on this page are examples to adapt. Test them in your own environment before you rely on them.

What this page says about the products it names was checked against their public documentation on 4 October 2026. Product and company names are trademarks of their owners. Fokals is not affiliated with them or endorsed by them.